Skip to content

gh-158097: Make test_staggered_race_with_eager_tasks deterministic - #158469

Open
abhinav-phi wants to merge 1 commit into
python:mainfrom
abhinav-phi:fix-158097-staggered-race-flake
Open

abhinav-phi wants to merge 1 commit into
python:mainfrom
abhinav-phi:fix-158097-staggered-race-flake

Conversation

@abhinav-phi

@abhinav-phi abhinav-phi commented Sep 30, 2026 •

Copy link
Copy Markdown

Fixes #158097.

Root cause

asyncio.staggered.staggered_race() awaits the coroutines it is given inline — result = await coro_fn() in run_one_coro() (Lib/asyncio/staggered.py:121) — so each one runs inside the frame of the run_one_coro() task that awaits it, not in a task of its own.

The failing helper in the test yielded once before raising:

async def fail():
    await asyncio.sleep(0)
    raise ValueError("no good")

That left it still unfinished at the moment the race was won, and staggered_race() cancels every coroutine that has not finished as soon as one of them wins. Concretely, await asyncio.sleep(0) suspends the run_one_coro() task that is running fail() and leaves a scheduled resumption behind. If the winner completes before that resumption runs, its cancel() sets _must_cancel, the pending resumption throws CancelledError into fail(), and run_one_coro() stores that in excs[2] instead of the ValueError.

The schedule the test silently depends on, traced on a --with-pydebug build:

time event
0.00 s blocked() starts (index 0, never finishes)
0.25 s stagger delay elapses, asyncio.sleep(1) starts (index 1), due at 1.25 s
0.50 s stagger delay elapses, fail() starts (index 2) and suspends on its sleep(0)
1.25 s asyncio.sleep(1) finishes → winner → every unfinished coroutine is cancelled

Pending timers are only moved onto the ready queue at the top of BaseEventLoop._run_once(), and they are appended after the handles that are already queued. So the t=0.50 s stagger timer and the t=1.25 s asyncio.sleep(1) timer run in due-time order, and when both come due in the same _run_once() the stagger timer starts fail(), which suspends again immediately, and the very next handle is the winner. Anything that deschedules the process for more than ~0.75 s does it — routine on a busy CI machine, which is where the OpenEmbedded report came from.

This is not an asyncio bug: a loser that has not finished is documented to be cancelled. The test was asserting on an exception that fail() had not yet been given a chance to raise.

The fix

Raise immediately, without suspending first. Because the coroutine is awaited inline, the ValueError is then raised and stored in excs[2] in the very same step of run_one_coro() that started it, so there is no scheduled resumption left for a cancellation to overtake, and the outcome no longer depends on the event loop getting another iteration.

         async def fail():
-            await asyncio.sleep(0)
+            # Fail without suspending first.  staggered_race() awaits the
+            # coroutines inline and cancels every one of them that has not
+            # finished once another one wins, so a coroutine that yields before
+            # raising may be cancelled instead, and excs[2] would report a
+            # CancelledError.  Raising right away keeps the outcome
+            # independent of how busy the event loop is.
             raise ValueError("no good")

blocked() still suspends and is still cancelled by the winner, and asyncio.sleep(1) still suspends and still wins, so the assertions on excs[0] and excs[index] are unchanged. The sibling test_staggered_race_with_eager_tasks_no_delay already uses an immediately-raising fail(), and test_asyncio/test_staggered.py covers staggered_race()'s own cancellation semantics with fully deterministic timing. No asyncio library code is touched.

Verification

Built main (00307b0, 3.16.0a0) with ./configure --with-pydebug.

The failure is real and reproducible. Blocking the event loop while it waits for the stagger timer reproduces the issue's traceback verbatim — line 240, in run / self.assertIsInstance(excs[2], ValueError) / AssertionError: CancelledError() is not an instance of <class 'ValueError'> — and the threshold is exactly the ~0.75 s the test depends on:

event-loop deschedule unmodified main with this fix
0.2 s, 0.5 s, 0.7 s, 0.75 s, 0.8 s pass pass
1.0 s fail pass
1.5 s, 2 s, 5 s, 10 s fail pass

The test now passes consistently, with and without that injection:

check runs failures
test method, full setUp/tearDown per run, 10 workers × 500 iterations 5000 0
regrtest, fresh interpreter per run, 10 workers × 150 runs 1500 0
./python -m test -j10 test_asyncio 2927 0 (2 Windows-only skips)
$ ./python -m test -j10 test_asyncio
== Tests result: SUCCESS ==
33 tests OK.
Total duration: 39.8 sec
Total tests: run=2,927 skipped=21
Total test files: run=35/35 skipped=2
Result: SUCCESS

AI tools were used assistively; I reviewed and tested every change and take responsibility for it.

staggered_race() awaits the coroutines it is given inline, so the
"await asyncio.sleep(0)" in the test's fail() helper suspended the
run_one_coro() task running it and left a scheduled resumption behind.
If the asyncio.sleep(1) winner (started 0.25s late by the stagger delay)
completes before that resumption runs, the winner's cancel() sets
_must_cancel and the resumption throws CancelledError into fail(), so
excs[2] reports a CancelledError instead of a ValueError.

Pending timers are only moved onto the ready queue at the top of
BaseEventLoop._run_once() and are appended after the handles already
queued, so the t=0.50s stagger timer and the t=1.25s sleep(1) timer run
in due-time order.  When both are due in the same _run_once() - which is
what a busy CI machine causes - the stagger timer starts fail(), which
immediately suspends again, and the very next handle is the winner.

Raise straight away instead, so the ValueError is stored in excs[2] in
the same step that starts the coroutine and there is no longer a
scheduled resumption for a cancellation to overtake.  The outcome no
longer depends on the event loop getting another iteration.
@python-cla-bot

python-cla-bot Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

All commit authors signed the Contributor License Agreement.

CLA signed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review tests Tests in the Lib/test dir

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test_staggered_race_with_eager_tasks fails with CancelledError

1 participant