Next Python SDK major - #5005
sentrivana wants to merge 351 commits into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #5005 +/- ##
===========================================
+ Coverage 70.55% 83.76% +13.21%
===========================================
Files 180 180
Lines 18077 18080 +3
Branches 3008 3009 +1
===========================================
+ Hits 12754 15145 +2391
+ Misses 4432 1943 -2489
- Partials 891 992 +101
|
Codecov Results 📊✅ 57303 passed | ❌ 1 failed | ⏭️ 2727 skipped | Total: 60031 | Pass Rate: 95.46% | Execution Time: 163m 35s 📊 Comparison with Base Branch
➕ New Tests (1)View new tests
➖ Removed Tests (1)View removed tests
❌ Failed Tests
|
Semver Impact of This PR⚪ None (no version bump detected) 📋 Changelog PreviewThis is how your changes will appear in the changelog. New Features ✨
Bug Fixes 🐛Anthropic
Documentation 📚
Internal Changes 🔧
Other
🤖 This preview updates automatically when you update the PR. |
Fixes for things that the bots [surfaced](#5005) on the major branch: - some version checks were too late (after patching) - fix TrytondWSGI integration name/`_MIN_VERSIONS` entry mismatch Also, changed the warning of the `DidNotEnable` message from "X not installed" to "X not installed or incompatible".
…#7201) The sync request/response handler passed the _isolation_ scope to `_set_transaction_name_and_source`, but the transaction/segment span lives on the _current_ scope. As a result the route-resolved name never reached the span for sync endpoints, which were instead named by the raw URL from the ASGI middleware (`transaction_info.source` of `url` rather than `route`). Async handlers already used the current scope and were unaffected. Pass the current scope (already computed above) so sync and async handlers behave identically: - streaming: the segment name / `sentry.segment.name.source` are route-based - static: the transaction event name / source are route-based For parametrized routes this also removes high-cardinality URL transaction names for sync endpoints. Found while working on #7183.
The options were deprecated with 2fef9bc.
Inline the branches that set the route template as the segment name (i.e., `transaction_style="route_pattern"`). Closes #7761
Inline the branches that set the route template as the segment name (i.e., `transaction_style="uri_template"`). Closes #7757
Inline the branches that set the route template as the segment name (i.e., `transaction_style="url"`). Closes #7758
Inline the branches that set the route template as the segment name (i.e., `transaction_style="url"`). Closes #7763
- inline things to avoid mental context switches - use global API
Inline the branches that set the route template as the segment name (i.e., `transaction_style="method_and_path_pattern"`). Stop including the HTTP method in the name. Closes #7801
Inline the branches that set the route template as the segment name (i.e., `transaction_style="url"`). Closes #7811
Always filter the query string through the `data_collection` `url_query_params` setting. The legacy `send_default_pii` path is resolved into an equivalent `data_collection` default, so the raw query string is no longer passed through. Update the Django ASGI and FastAPI tests to configure `data_collection` instead of `send_default_pii`. Refs PY-2798 Refs #7566
…7772) `data_collection` is now the default way to control what data the SDK collects, so the always-on `EventScrubber` is no longer needed. The expectation is that, if people need to filter out request bodies, etc., that they use the `before_send` callback. Removes `sentry_sdk.scrubber` (`EventScrubber`, `DEFAULT_DENYLIST`, `DEFAULT_PII_DENYLIST`), the scrubbing step in `_prepare_event`, and the `event_scrubber` init option. Passing `event_scrubber=` to `init` now raises a `TypeError`. Refs PY-2798 Refs #7566
### Description Align span names with OTel conventions: > Span name MUST be of the format Service.Operation as per the AWS HTTP API, e.g., DynamoDB.GetItem, S3.ListBuckets. This is equivalent to concatenating rpc.service and rpc.method with . and consistent with the naming guidelines for RPC client spans. So we use modeled `ServiceID` and operation for span names: `aws.<hyphenized-service-id>.<Operation>` to `<ServiceID>.<Operation>`; e.g. `aws.s3.GetObject` => `S3.GetObject`. #### Issues Resolves #7498
| self.options["project_root"], | ||
| ) | ||
|
|
||
| if event is not None: | ||
| event_scrubber = self.options["event_scrubber"] | ||
| if event_scrubber: | ||
| event_scrubber.scrub_event(event) | ||
|
|
||
| if scope is not None and scope._gen_ai_original_message_count: | ||
| spans: "List[Dict[str, Any]] | AnnotatedValue" = event.get("spans", []) | ||
| if isinstance(spans, list): | ||
| for span in spans: | ||
| span_id = span.get("span_id", None) | ||
| span_data = span.get("data", {}) | ||
| if ( | ||
| span_id | ||
| and span_id in scope._gen_ai_original_message_count | ||
| and SPANDATA.GEN_AI_REQUEST_MESSAGES in span_data | ||
| ): | ||
| span_data[SPANDATA.GEN_AI_REQUEST_MESSAGES] = AnnotatedValue( | ||
| span_data[SPANDATA.GEN_AI_REQUEST_MESSAGES], | ||
| {"len": scope._gen_ai_original_message_count[span_id]}, | ||
| ) | ||
| if previous_total_spans is not None: | ||
| event["spans"] = AnnotatedValue( | ||
| event.get("spans", []), {"len": previous_total_spans} | ||
| ) | ||
| if previous_total_breadcrumbs is not None: |
There was a problem hiding this comment.
EventScrubber removal leaves user extras and contexts unredacted by default
With send_default_pii=False (the default), arbitrary values added to event extra or contexts are no longer passed through EventScrubber’s denylist before sending. data_collection filters specific collection categories but does not cover arbitrary custom fields. The removed event_scrubber option now raises during initialization, and the migration guide does not explain the change. Please document the removal and recommend before_send for equivalent scrubbing of custom event data.
Evidence
Scope._apply_extras_to_event()and_apply_contexts_to_event()copy custom values into the event without filtering them.Client._prepare_event()proceeds fromhandle_in_app()to serialization andbefore_send; there is no EventScrubber pass in that path._map_from_send_default_pii()configures filtering for specific categories, not arbitraryextraorcontextsfields.event_scrubberis no longer an accepted option, andMIGRATION_GUIDE.mddoes not document its removal or a replacement.
Also found at 2 additional locations
sentry_sdk/client.py:131-133sentry_sdk/consts.py:1373-1374
Identified by Warden · code-review, find-bugs · 5ZW-LDV
|
|
||
| ip = _get_ip(scope) | ||
| if ip: | ||
| client_options = sentry_sdk.get_client().options | ||
| if client_options["data_collection"]["user_info"]: |
There was a problem hiding this comment.
ASGI middleware dereferences unresolved data_collection for a non-recording client
When the SDK has not been initialized, get_client() can return a NonRecordingClient whose options retain data_collection=None. For requests with a client IP, the direct access to client_options["data_collection"]["user_info"] raises TypeError before the ASGI app is called. Guard this access with has_data_collection_enabled() and the legacy PII fallback, or handle None, as the WSGI middleware does.
Evidence
Scope.get_client()falls back toNonRecordingClientwhen no active client is available; its options come fromDEFAULT_OPTIONS, wheredata_collectiondefaults toNone._run_app()gets an IP from_get_ip(scope)and directly indexesclient_options["data_collection"]["user_info"]at line 228 when one is present.- That access raises
TypeErrorbefore the downstream ASGI app is called; the WSGI middleware instead checkshas_data_collection_enabled()and has a legacy PII fallback.
Identified by Warden · code-review, find-bugs · RBX-L42
| span = sentry_sdk.start_span( | ||
| name=f"invoke_agent {run_name}" if run_name else "invoke_agent", | ||
| attributes={ | ||
| "sentry.op": OP.GEN_AI_INVOKE_AGENT, | ||
| "sentry.origin": LangchainIntegration.origin, | ||
| SPANDATA.GEN_AI_OPERATION_NAME: "invoke_agent", | ||
| }, | ||
| ) |
There was a problem hiding this comment.
Stream agent spans drop gen_ai.response.streaming attribute
Include SPANDATA.GEN_AI_RESPONSE_STREAMING: True in the start_span attributes so stream invoke_agent spans stay distinguishable from invoke.
Evidence
- Prior code set
SPANDATA.GEN_AI_RESPONSE_STREAMING: Trueon both the span-streaming and legacy branches. - The consolidated
sentry_sdk.start_span(... attributes=...)block omits that attribute entirely. - Other streaming AI integrations (e.g. openai_agents agent_run, google_genai, openai) still set
GEN_AI_RESPONSE_STREAMINGfor stream paths.
Identified by Warden · code-review, find-bugs · NDG-GCQ
| if "incoming_request" in data_collection["http_bodies"]: | ||
| if "body" in aws_event: | ||
| request["data"] = aws_event.get("body", "") |
There was a problem hiding this comment.
Request body attached without max_request_body_size check
When attaching aws_event body data, call request_body_within_bounds() like aiohttp/WSGI so max_request_body_size still limits payload size.
Evidence
- The new body path sets
request["data"] = aws_event.get("body", "")with no size guard. data_collection._map_from_send_default_piidocuments bodies as bounded bymax_request_body_size.aiohttp.get_aiohttp_request_data()and_wsgi_common.RequestExtractorboth callrequest_body_within_bounds()before attaching bodies.- This path is now the default for all clients because
data_collectionis always resolved.
Identified by Warden · code-review, find-bugs · AD4-33Z
| def _get_transaction_name(request: "Any") -> str: | ||
| try: | ||
| if transaction_style == "url": | ||
| name = bottle_request.route.rule or "bottle request" | ||
| else: | ||
| name = ( | ||
| bottle_request.route.name | ||
| or transaction_from_function(bottle_request.route.callback) | ||
| or "bottle request" | ||
| ) | ||
|
|
||
| sentry_sdk.get_current_scope().set_transaction_name( | ||
| name, | ||
| source=SEGMENT_SOURCE_FOR_STYLE[transaction_style], | ||
| ) | ||
| return request.route.rule or "bottle request" | ||
| except RuntimeError: | ||
| pass | ||
|
|
||
|
|
||
| def _set_transaction_name_and_source( | ||
| event: "Event", transaction_style: str, request: "Any" | ||
| ) -> None: | ||
| name = "" | ||
|
|
||
| if transaction_style == "url": | ||
| try: | ||
| name = request.route.rule or "" | ||
| except RuntimeError: | ||
| pass | ||
|
|
||
| elif transaction_style == "endpoint": | ||
| try: | ||
| name = ( | ||
| request.route.name | ||
| or transaction_from_function(request.route.callback) | ||
| or "" | ||
| ) | ||
| except RuntimeError: | ||
| pass | ||
|
|
||
| event["transaction"] = name | ||
| event["transaction_info"] = { | ||
| "source": TRANSACTION_SOURCE_FOR_STYLE[transaction_style] | ||
| } | ||
| return "bottle request" |
There was a problem hiding this comment.
Unmatched Bottle requests can fail while resolving the transaction route
When Bottle handles an unmatched path, request.route can be None. _patched_handle then dereferences .rule without guarding against None—in its HTTP-route block when tracing is enabled, and in _get_transaction_name otherwise. The resulting AttributeError can prevent Bottle's normal 404 response from being returned. Check that the route exists before reading its rule at both access sites.
Evidence
_patched_handlecalls Bottle's original handler, then readsbottle_request.route.rule; that access catchesRuntimeErroronly and runs when a server span exists.- Regardless of tracing,
_patched_handlethen calls_get_transaction_name, which also dereferencesrequest.route.ruleand catches onlyRuntimeError. - Bottle's route property may be
Nonewhen no route matched, so either dereference raisesAttributeErrorinstead of allowing the normal 404 response to proceed.
Identified by Warden · find-bugs · 68K-JMP
Also: consolidated some tests. Closes https://linear.app/getsentry/issue/PY-2861/remove-send-default-pii-from-httpx
Basically the same as httpx Closes https://linear.app/getsentry/issue/PY-2862/remove-send-default-pii-from-httpx2
…#7831) ### Description - drop legacy test cases from `tests/integrations/utils.py` (`DATA_COLLECTION_USER_INFO_CASES_LEGACY`; `DATA_COLLECTION_REMOTE_ADDR_CASES_LEGACY`; `DATA_COLLECTION_QUEUES_CASES_LEGACY`). - removes `**init_kwargs` from `sentry_init(...)`.
| def set_conversation_id(conversation_id: str) -> None: | ||
| """ | ||
| Set the conversation_id in the scope. |
There was a problem hiding this comment.
AI input spans include raw, unbounded blob content
When data_collection.gen_ai.inputs is enabled, Anthropic base64 content and corresponding OpenAI/LangChain formats are copied into span message attributes verbatim. The message serialization path has no local size cap, so large image or document payloads can also inflate spans or cause them to exceed payload limits. Replace blob contents with a substitute and bound or truncate serialized messages before attaching them.
Evidence
transform_anthropic_content_partcopiessource["data"]directly into blobcontent; the OpenAI and generic transformers also return inline content verbatim.- Anthropic, LiteLLM, and LangChain pass transformed messages to
set_data_normalizedwhen GenAI input collection is enabled. set_data_normalizedJSON-serializes messages and callsspan.set_attributewithout a message-size limit.- Google GenAI and PydanticAI explicitly replace blob contents with
BLOB_DATA_SUBSTITUTE, but these shared transforms do not.
Identified by Warden · code-review · 4JF-TJE
| def continue_trace(incoming: "Dict[str, Any]") -> None: | ||
| """ | ||
| Sets the propagation context from environment or headers and returns a transaction. | ||
| Continue a trace from headers or environment variables. | ||
|
|
||
| This function sets the propagation context on the scope. Any span started | ||
| in the updated scope will belong under the trace extracted from the | ||
| provided propagation headers or environment variables. | ||
|
|
||
| continue_trace() doesn't start any spans on its own. Use the start_span() | ||
| API for that. | ||
| """ | ||
| return get_isolation_scope().continue_trace( | ||
| environ_or_headers, op, name, source, origin | ||
| ) | ||
| return traces.continue_trace(incoming) |
There was a problem hiding this comment.
Document the continue_trace migration
The 3.x migration guide does not explain that continue_trace no longer accepts op, name, source, or origin, or returns a Transaction. Add the migration pattern: call continue_trace(headers) to set propagation context, then use start_span(...) to create a span.
Evidence
sentry_sdk.api.continue_trace(incoming)now accepts onlyincomingand returnsNone; its docstring says it does not start spans.sentry_sdk.traces.continue_tracesets propagation context, whilestart_spanis a separate API for creating spans.- The 3.x
MIGRATION_GUIDE.mdlists other removed APIs but does not describe thiscontinue_tracesignature and behavior change.
Identified by Warden · code-review · UAA-4PR
| identifier = "anthropic" | ||
| origin = f"auto.ai.{identifier}" | ||
|
|
||
| def __init__(self: "AnthropicIntegration", include_prompts: bool = True) -> None: | ||
| self.include_prompts = include_prompts | ||
|
|
||
| @staticmethod | ||
| def setup_once() -> None: | ||
| version = package_version("anthropic") | ||
| version = parse_version(ANTHROPIC_VERSION) |
There was a problem hiding this comment.
include_prompts removed without migration path
Removing AnthropicIntegration(include_prompts=...) is a breaking API/behavior change—document the move to data_collection.gen_ai (and that the old default True is now False unless send_default_pii/data_collection enables it) in MIGRATION_GUIDE.md.
Evidence
AnthropicIntegrationno longer defines__init__;AnthropicIntegration(include_prompts=...)will raiseTypeError.- Prompt/response capture now uses
data_collection["gen_ai"]["inputs"|"outputs"](e.g. around_set_common_input_data/_set_output_data). - When
data_collectionis unset, those flags map fromsend_default_pii(default False), so the old defaultinclude_prompts=Trueis no longer preserved. MIGRATION_GUIDE.mdhas no entry forinclude_promptsor this AI integration option change.
Identified by Warden · code-review · Y5C-REM
| collect_response = ( | ||
| "outgoing_response" in client_options["data_collection"]["http_bodies"] | ||
| ) |
There was a problem hiding this comment.
Ariadne responses bypass the legacy PII gate
When data_collection was not explicitly provided, Ariadne previously gated response capture on send_default_pii. The new processor instead relies only on http_bodies; the default mapping includes outgoing_response even when send_default_pii is false, so error events can now include the full GraphQL response payload, including partial data. Preserve the legacy gate for this configuration, while continuing to honor explicit data_collection settings.
Evidence
_map_from_send_default_piisetshttp_bodiesto all body types regardless ofsend_default_pii;has_data_collection_enableddistinguishes this default mapping from user-provided configuration.- Ariadne’s
_make_response_event_processorchecks only foroutgoing_responseandresponse.get("errors"), then stores the entire response undercontexts.response.data. - The Ariadne error handlers attach this processor to events for GraphQL errors, so a response containing both errors and partial
datais included. - Strawberry and gql retain a
should_send_default_pii()fallback when data collection was not user-provided, unlike Ariadne.
Identified by Warden · code-review · L64-KV5
We're preparing our next major on this branch.
The project is tracked in Linear. If you don't have access, we'll try to tag issues belonging to the project with the
SDK3.0 label on GitHub so that you can follow along.Notable changes
Context
You might have read this announcement about us discontinuing work on a 3.0. This is referring to the work done on the
potel-basebranch, which included two types of changes: a huge refactor of our tracing code on the one hand, and various unrelated changes, improvements and fixes on the other. We're dropping the huge refactor part, and only porting the rest, to a new branch and eventually a new 3.0 release.