fix(ingest): read Spring AI's Anthropic and Bedrock Converse usage as cache-exclusive - #1176
Conversation
Maple review🟢 Confidence 4/5 · likely safe to merge Adds a
What was checked
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review. 📝 WalkthroughWalkthroughCache-exclusivity selection now checks span attributes for Spring AI. The previous model-name-based rules for Strands, Google ADK, Agno, and Microsoft Agent Framework are removed. Tests cover Spring AI Anthropic, Bedrock Converse, and OpenAI usage. ChangesAI usage cache accounting
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to This change fixes how cache tokens are counted for Spring AI Anthropic and Bedrock Converse spans. It is limited to ingest usage accounting and no merge-blocking risk was found. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Warning Some tools did not complete. Review the errors below. 🔧 Clippy (1.98.1)Clippy execution failed Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Problem
Spring AI's native Anthropic client (
AnthropicChatModel.getDefaultUsageand its streaming path) and its Bedrock Converse client (BedrockProxyChatModel) reportgen_ai.usage.input_tokenswithout the cache, as the Anthropic Messages and Converse APIs do. The cache comes separately asgen_ai.usage.cache_read.input_tokensandgen_ai.usage.cache_creation.input_tokens.input_excludes_cache()inapps/ingest/src/ai_session/usage.rshad nospring_aiarm, so ingest read these spans as cache-inclusive and subtracted the cache from the prompt. Thecache > promptguard only catches prompts smaller than their cache; once uncached input is at least as large as the cache, the session total came out too low.Evidence from EU prod, using synthetic spans built from the Spring AI v2.0.1 source (
gen_ai.system=anthropic, opchat, scopeorg.springframework.boot):verify3-b-edge-convverify3-b-bigprompt-convverify3-b-smallprompt-convFix
Add a
spring_aiarm: the prompt excludes the cache whengen_ai.systemisanthropicorbedrock_converse. In Spring AI this attribute is set perChatModelclass (AiProviderenum), so it names the client that made the call, not the model's vendor. The OpenAI client, which is also what Spring AI apps use for OpenRouter, reportsopenaiand stays inclusive.Spring AI writes only
gen_ai.system(AiObservationAttributes.AI_PROVIDER), notgen_ai.provider.name, in both v2.0.1 andmain, so only that key is read. The lookup runs only for Spring AI model-call spans.Tests
spring_ai_anthropic_and_converse_are_cache_exclusiveinusage.rscovers the three sessions above, a Bedrock Converse span, and a Spring AI OpenAI-client span that stays inclusive.cargo test --lib ai_session: 83 passed.Note
Spring AI 1.1.x's Anthropic client reports no cache attributes at all, so the cache is invisible on those spans and ingest cannot recover it.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by CodeRabbit