PYTHON-6135 Fix perf regression by explicitly choosing the python binary - #3086
Merged
Merged
Conversation
PYTHON-6084 (dda56d3) made every uv-invoking step source setup-uv-python.sh, which prefers the toolchain interpreter (UV_PYTHON_PREFERENCE=system). On the performance-benchmarks variant the venv created by the "run server" step was bound to the unoptimized toolchain build (3.10.11 [GCC 11.5.0 Red Hat]) instead of uv's optimized managed build (3.10.11 [Clang 16.0.3]), and the test step reused it, regressing CPU-bound benchmarks by 15-30%. Pin the perf tasks to the managed interpreter in every step that invokes uv: generate_config.py passes task-level UV_PYTHON and UV_PYTHON_PREFERENCE=only-managed to the "run server" and "run tests" functions, setup_tests.py persists that selection in test-env.sh, and setup-uv-python.sh lets a task-level preference win over the toolchain defaults. run-tests.sh re-sources setup-uv-python.sh for direct/local runs.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Member
Author
|
All of the failures are flakes, upstream service outages, or tracked by DRIVERS-3666 (the amazon latest failure). |
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The managed interpreter selection is consistently propagated and the generated configuration matches its source.
Review effort: Balanced
Findings: None
What changed in this PR
Pins performance tasks to uv’s optimized managed Python 3.10.11 build, restoring benchmark performance.
Changes:
- Configures perf tasks with
UV_PYTHON_PREFERENCE=only-managed. - Preserves task-level Python preferences through setup scripts.
- Regenerates Evergreen task and function configurations.
| File | Description |
|---|---|
.evergreen/scripts/setup-uv-python.sh |
Preserves task-defined Python preferences. |
.evergreen/scripts/setup_tests.py |
Selects managed Python for perf tests. |
.evergreen/scripts/generate_config.py |
Adds managed-Python variables to perf tasks. |
.evergreen/run-tests.sh |
Initializes Python selection for direct runs. |
.evergreen/generated_configs/tasks.yml |
Regenerates perf task configuration. |
.evergreen/generated_configs/functions.yml |
Passes the preference through Evergreen functions. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
aclark4life
requested changes
Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PYTHON-6135
Changes in this PR
PYTHON-6084 bound the
performance-benchmarksvenv to the unoptimized toolchain python instead of uv's optimized managed build, causing BSON/JSON benchmarks to regress 12–30%.Perf tasks now pin
UV_PYTHON=3.10.11+UV_PYTHON_PREFERENCE=only-managedto ensure that the uv binaries are used.Test Plan
Perf build. I did a spot check of several of the benchmarks and confirmed that the throughput matches previous levels.
Checklist
Checklist for Author
Checklist for Reviewer