Repository navigation
boundary: synchronous-return arm delivers a host result into an instance poisoned during the import call #359
Description
Activity
- addedbugSomething isn't workingSomething isn't workingp3Lowest priority: watchlists, doc-only adjudications, deletion candidatesLowest priority: watchlists, doc-only adjudications, deletion candidates
on Sep 12, 2026 Cross-check against the pinned Component Model reference (
definitions.py@7c67611) and wasmtime4675ee1. Verdict vocabulary: SPEC-BACKED = the reference mandates the expected behavior; CONTRACT-ONLY = spec silent, polyengine's own contract decides; POLICY-QUESTION = neither decides; WEAKENED = part of the claim is overstated (corrections below).Spec: the setup is spec-legal: a sync-typed export re-entered from a host
function while an activation of the same instance is live is allowed —
Task.enter_implicit_threadapplies the backpressure/exclusive gate only
if self.ft.async_:; Concurrency.md:200-202 "functions withoutasyncin
their type are not allowed to block, non-asyncfunctions do not attempt to
acquire the 'exclusive' lock; they just barge in";Store.lift/Store.invoke
permit nested invocation atnesting_depth > 0. The continuation is not
spec-representable: the re-entered activation'strap()raises
Trap(BaseException), which ends the whole computation including the outer
activation; the reference has no "host catches the trap and returns a value"
path (intrinsics.md:22 restates: "Component Model traps must be uncatchable").
So the spec says "the outer call cannot complete" only in the trivial sense
that nothing completes. It does not say what a runtime with narrower fault
scope should do with an activation already inside the newly-poisoned instance.Contract:
embedder-api.md:145-147"A trap escaping a guest activation
poisons its instance ... Catching its host-side rejection does not restore
that instance" — quoted correctly; it says the instance stays poisoned, not
what happens to activations already running inside it.architecture.md:302-303
"permanently poisons that instance; later entry names the original cause"
— the only mechanism named is an entry gate. The ladder item 1
(embedder-api.md:275-276) "Reentrance into a live instance is allowed" is
also quoted correctly. So the contract is silent on the exact point at issue;
the issue's authority is really the runtime's own deferred-arm policy
(boundary.ts:1865produceguard,:1887/:1897async-arm discard, and the
pinned "poisoned own never registers" test), i.e. internal consistency.Wasmtime: same behavior as polyengine's plain-mode actual, and the
preamble's "exclusive borrow prevents the window" does not apply here.
Synchronous re-entry from inside a host fn body is exactly what the borrowed
StoreContextMutpermits, and it is tested:
tests/all/component_model/func.rs:3051-3120(recurse_a_then_a: the
import's host fn callsrun.call(&mut store, ())on the same instance and
the outer call succeeds). If that nested call traps,call_rawsets the
store-wide flag (component/func.rs:389-391) and returnsErrto the host
fn; a host fn that swallows it and returnsOk(value)has its result lowered
normally (func/host.rs:398-423: notrapped()check anywhere in
call_sync_lower), the outer wasm frame continues, and the outercall_raw
returnsOk—trapped()is only ever consulted throughmay_enter()at
entry (component/func.rs:468,concurrent.rs:6278,
resources/any.rs:204). So wasmtime, like polyengine's sync arm, treats
poisoning as an entry gate and lets the live activation finish and return its
value; the next entry fails withTrap::CannotEnterComponent.Verdict: POLICY-QUESTION. The decision: is poisoning (a) an entry gate only
(architecture §6 wording, wasmtime, polyengine's sync arm), or (b) also an
abort of activations already live in the instance (polyengine's deferred
arms, the issue's "Expected")? Neither spec nor contract chooses; the spec's
only statement is that the whole computation dies. What survives regardless
of the choice: (1) polyengine is inconsistent across its own arms — the same
poisoning event is fatal to a deferred settlement and invisible to a
synchronous one; (2) the JSPI-mode outcome (outer rejects with a false
"deadlock detected" verdict rather than either the value or the poison cause)
is wrong under both policies.
wasmtime: same behavior as plain-mode actual (entry-gate semantics; the outer
call returns its value into a trapped store).Severity: keep P3 — the inconsistency and the bogus JSPI diagnostic are real
but the trigger requires a host to swallow aTrapfrom a reentrant call,
and one of the two candidate policies says the current sync-arm behavior is
correct.Notes: Retitle toward "poisoning is an entry gate in the sync arm but an
abort in the deferred arms; JSPI mode gives a false deadlock verdict", and
move the "Expected" into a policy proposal. Drop the implication that the
Error-model sentence mandates aborting live activations; it does not. Cite
wasmtimerecurse_a_then_aas evidence that entry-gate semantics is a
defensible choice, and the runtime's:1865/:1887guards as evidence it has
already half-chosen the other one. If (b) is chosen, the fix sketched in the
issue (mirrorproduce's guard after a synchronoushostFnreturn) is the
right shape; if (a), the deferred arms' guards should be revisited instead.- addedpolicyDecision owed on deltic-owned semanticsDecision owed on deltic-owned semantics
on Sep 12, 2026 Re-evaluated on current main
5616bce, Deno 2.9.5. Keep P3 / policy. Plain mode continues and disposes the returned own; JSPI mode rejects with the misleading deadlock diagnosis; the deferred control discards the result.The setup requires the host to catch a reentrant guest trap and continue. Current contract wording does not clearly choose entry-only poisoning versus aborting already-live activations. Decide that boundary together with #347 and async-import trap routing (#349), rather than treating a copied deferred-arm guard as proof of the intended semantics. Prefer checking poison at host-return/resume boundaries without implying arbitrary interruption of running Wasm. The false deadlock diagnosis needs correction under either policy.
Rechecked current main
ae86a4aafter #369, Deno 2.9.5. Partially resolved; keep P3/open for the JSPI outcome.The synchronous
sync(trap)-then-return setup now discards the returned own in both modes (valueDisposed:0, no registration). Plain mode rejects with the originalTrap: guest trapped: unreachableinstead of completing successfully. This closes the missing synchronous-result eligibility check identified in the original report.JSPI still rejects that same setup with the false
deadlock detected ... no thread is ready and no host call is outstandingdiagnosis. The deferred control also still gives a deadlock diagnosis rather than the original poison cause. Thus do not close this issue solely because guarded delivery is now shared.The additional
trap()Promise-surface variant in JSPI logspoisoned-at-return:falseand can still dispose the returned value before poisoning is observed; it is not evidence of delivery into an instance already poisoned at host return. Distinguish actual poison timing from the later rejection when narrowing the remaining case.
synchronous-return arm delivers a host result into an instance poisoned during the import call
Severity: P3 Confidence: medium
Location:
runtime/src/exec/boundary.ts:1920-1922(else { onResolve(toResults(raw)); }),:1930-1942(sync tail),:1944-1953(eager async tail). Compare the poison guards the two deferred arms carry::1865(produce:if (isInstancePoisoned(inst)) throw instancePoisonCause(inst)) and:1887/:1897(async arm discard). (baseline 1a4f5e1)Authority:
contracts/embedder-api.md§"Error model" — "A trap escaping a guest activation poisons its instance under the runtime's policy (docs/architecture.md §6). Catching its host-side rejection does not restore that instance."docs/architecture.md§6 #173 — "A trap escaping a guest activation permanently poisons that instance". Runtime's own stated intent at boundary.ts:1885-1886 ("Discard cancelled or poisoned recipients before lowering can write guest memory") and the pinned testhost_settlement_test.ts"poisoned own never registers".Expected: When the host import re-enters the same instance synchronously (allowed: §"
sync()" failure ladder item 1, "Reentrance into a live instance is allowed"), the re-entered activation traps, and the host catches that Trap and returns a value, the runtime treats the outer activation like every other recipient in a poisoned instance: the result is not lowered / theown<R>is not registered, and the outer export fails naming the poison cause.Actual: The synchronous return path has no poison check. In plain mode the outer guest activation keeps executing inside the poisoned instance: the
own<R>result is registered and handed to the guest (valueDisposed:1— the guest randrop-ron it), and the outer export resolves successfully (runSync=0) althoughisInstancePoisonedis already true at return time; the very next entry is refused "instance poisoned by: Trap: guest trapped: unreachable". In JSPI mode the result is also delivered (valueDisposed:1) and the outer export rejects, but with a falsedeadlock detected ... no thread is ready and no host call is outstandingverdict rather than the poison cause (diagnostic only, noted, not the claim). The deferred-arm control (same reentrant poisoning, settled 5 ms later through the suspendable arm) discards as designed:valueDisposed:0.Repro:
f2_sync_return_into_poisoned.ts(fixturehost-settlement.wasm,runSync; the import callssync(c.exports.trap)()and catches):Distinct from known issues: #347 is a TOCTOU in the async non-suspendable arm's existing guard (accessor reentrancy after the check at :1887). This is the synchronous return arm (:1921 and the tails), which has no guard at all and needs no accessor trick — plain host reentrancy through the documented
sync()surface reaches it.Fix direction: after
hostFnreturns synchronously (and before the eageronResolve), mirrorproduce:if (isInstancePoisoned(inst)) { subtask.unwindLenders(); deferred?.endScope(); throw instancePoisonCause(inst); }so the outer activation unwinds with the recorded cause instead of continuing. Confidence is medium because the contract speaks of "later entry"; the claim rests on the Error-model sentence plus the runtime's own deferred-arm policy.Running the repro
Scripts below were run from the repository root at baseline
1a4f5e1withthe shim and fixtures built (
just shim fixtures):<REPO>/in the scripts is the absolute path of the checkout (the review ranthem from outside the tree). Each repro was run at least twice and once under a
nonzero
POLYENGINE_SCHED_SEED; output was identical across runs.support.tsf2_sync_return_into_poisoned.ts