fix(ingest): keep TypeScript ADK's Gemini thinking whole - #1173
Conversation
TypeScript ADK records Gemini's candidatesTokenCount as gen_ai.usage.output_tokens and the Maple span processor adds thoughtsTokenCount as the reasoning figure, so the completion excludes the reasoning. Usage::read assumed the semconv convention and clamped reasoning to the completion: 100 output + 300 thoughts became 0 output, 100 reasoning. Every other Gemini emitter checked (Python ADK, the OTel and OpenInference Google GenAI instrumentations, langchain-google-genai) reports candidates plus thoughts and keeps the clamp.
Maple review🟢 Confidence 4/5 · likely safe to merge Adds an
What was checked
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review. 📝 WalkthroughWalkthroughUsage normalization now distinguishes Google ADK TypeScript ChangesUsage normalization
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~12 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The change preserves TypeScript ADK output and reasoning counts separately while retaining existing normalization for other emitters. No concrete merge-blocking risk is established; merge after normal checks pass. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The change is narrowly scoped to separating reported output and reasoning tokens. No introduced access-control or privilege issue was identified. Downstream use of these counts for billing or quota enforcement remains unverified. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Warning Some tools did not complete. Review the errors below. 🔧 Clippy (1.98.1)Clippy execution failed Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Usage::read(apps/ingest/src/ai_session/usage.rs) assumes the semconv convention, where the completion figure already contains the reasoning tokens. It clamps reasoning to the completion and subtracts it. An emitter whose completion is only the visible candidates loses its thinking. For example, output 100 + thoughts 300 is stamped as output 0, reasoning 100. EU prodverify2-adk-ts-syntheticshows this.This PR adds a per-emitter "output excludes reasoning" rule next to
input_excludes_cache. Like that rule, it is keyed on the emitter, never ongen_ai.provider.name. For an exclusive emitter, visible output = the completion and reasoning = the reasoning figure, with no clamp. Every other emitter keeps the clamp.Which emitters exclude reasoning
The source shows that only one Gemini path found reports a completion without the thoughts.
output_tokens@google/adk2.1.0, latest) + the setup guide'sAdkSpanProcessorcandidatesTokenCountonly; the processor addsthoughtsTokenCountasgen_ai.usage.reasoning.output_tokensand does not touch the outputcore/src/telemetry/tracing.ts(main):gen_ai.usage.output_tokens = usageMetadata.candidatesTokenCount; the guide's processor code; the prod synthetic span (output 100, reasoning 300)generate_contentspan)candidates_token_count + thoughts_token_countgoogle/adk/telemetry/_token_usage.py:TokenUsageoutput = candidates + thoughts (2.10 docstring: "candidate plus reasoning");reasoning.output_tokens= thoughtsopentelemetry-instrumentation-google-genai1.2b0generate_content.py~L432: "candidates_token_count excludes thoughts; output_tokens must be the full output count including reasoning tokens"openinference-instrumentation-google-genai1.4.8llm.token_count.completion= candidates + thoughts_utils.py_get_token_count_attributes_from_usage_metadatalangchain-google-genai(via LangChain instrumentations)chat_models.py~L2052@ai-sdk/google2.0.100outputTokens: candidatesTokenCount;reasoningTokensis not written to spans inai@5generate-text.ts@ai-sdk/googletextdetailconvert-google-usage.tsai.usage.outputTokenDetails.textTokensThe trace-capture recordings have no Gemini-native model call. All ADK captures route through OpenRouter (
openrouter/...ids), so the Python ADK and Google GenAI evidence comes from source. On EU prod, the only Gemini spans that carry a reasoning figure are the two synthetic verification spans.The rule
output_excludes_reasoning(vendor, span)returnsvendor == "google_adk" && span.name == "call_llm".call_llmis never the model call;generate_contentis (fix(ingest): count TypeScript ADK's call_llm as the model call #1170). Socall_llmcounting as the model call means TypeScript ADK.gcp.vertex_ai) goes through the same span, so the same rule covers it.Tests
google_adk_ts_call_llm_owns_usage: now carries the reasoning figure. Output 100 + thoughts 300 gives[600, 400, 0, 100, 300].google_adk_python_gemini_output_includes_thoughtsotel_google_genai_output_includes_thoughtsopeninference_google_genai_completion_includes_thoughtslangchain_js_reasoning_is_clamped_to_the_completionopenai_agents_ts_reasoning_is_clamped_to_the_completionvercel_v7_reasoning_past_the_completion_is_clampednumber_values_of_any_scalar_typecargo test --lib ai_session: 82 passed. Clippy reports nothing new inusage.rs.Merge order
This touches
usage.rsnext to #1170's change (is_model_call'scall_llmarm). #1170 is already onmainand this branch is cut from it, so there is nothing left to order.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by CodeRabbit
call_llmspans where completion counts already exclude reasoning tokens. Reasoning remains reported separately, and visible output totals are no longer reduced by subtracting those tokens again.