Skip to content

Component output is not byte-reproducible for identical inputs #247

Description

@starsang1995

Summary

Running componentize-py componentize twice with the same command and byte-identical
inputs can produce valid but byte-different WebAssembly components. This prevents use in
content-addressed build and artifact-verification pipelines.

I would like maintainer guidance on the shape of an opt-in reproducible-build mode before
opening the larger implementation PR. A small independent PR that makes WASI stub-adapter
emission deterministic is ready separately.

Reproduction

Using current main at aa3d6d1, Rust 1.95.0, WASI SDK 34, and the repository sandbox
example:

for output in first.wasm second.wasm; do
  componentize-py \
    -d sandbox.wit \
    componentize \
    --stub-wasi \
    guest \
    -o "$output"
done

shasum -a 256 first.wasm second.wasm
cmp first.wasm second.wasm

Expected: both files are byte-identical.

Actual: the files differ while remaining valid components.

Investigation

An experimental --reproducible implementation fixed or controlled all of the following:

  • build-time secure and insecure randomness;
  • build-time wall and monotonic clocks;
  • PYTHONHASHSEED;
  • CPython allocator initialization with PYTHONMALLOC=malloc_debug;
  • bytecode cache writes;
  • WASI stub-adapter map iteration order;
  • metadata on generated helper, standard-library, bindings, and input trees.

Even with all of those controls, two current-main outputs still differed: one run produced
19,323,845 bytes, the other 19,325,725 bytes, with the first difference at component offset
9.

I then repeated the experiment with ambient user/system site-packages excluded and the
same input, generated bindings, bundled helpers, and standard library exposed through stable
host directories. The outputs still differed (292a9320... versus f35b0e6d...). Both had
the same 208-page memory and 1,710 data segments, but reconstructing linear memory from those
segments found roughly 4.1 million differing bytes across 141 pages.

The difference is localized more narrowly than an encoder or linker problem:

  • both outputs contain 205 top-level component sections;
  • 204 sections are byte-identical;
  • only the first embedded core module differs;
  • within that module, all type, import, function, table, global, export, and code sections are
    byte-identical;
  • only the memory snapshot/data section differs.

Additional two-run probes showed identical values for Python string hashes, random, wall and
monotonic clocks, process/thread identity, visible directory ordering, generated Symbols, and
the component immediately before pre-initialization. Full linear-memory fingerprints taken
inside the guest were identical after app import and after every do_init phase (exports, type
tables, constructors, environment, runtime hooks, argv, and full-generation GC).

However, immediately after call_init returned through the Component Model boundary, before
component-init-transform measured globals or memory, the linear-memory fingerprints differed.
Disabling either or both WASI adapter/libc reset calls did not change that result. Ordering the
internal component-init-transform maps, reusing stable-inode directories, collecting Python
free lists after return through a second no-argument guest call, and retaining the Rust-level
app_name/Symbols values also did not fix it.

This localizes the remaining nondeterminism to the generated canonical-ABI post-return path for
the large nested init(app-name, symbols, stub-wasi) -> result<_, string> call (or allocator
state changed by that path), rather than Python initialization, component encoding, or snapshot
measurement. The current-main update moved from wit-bindgen 0.53.1 to 0.61.0 and Wasmtime 46.0.1
to 48.0.0; the exact responsible layer still needs a minimal reduction.

The experiment deliberately did not use post-generation byte rewriting or a volatile-byte
allowlist.

Proposed direction

  1. Add a regression test that componentizes one fixture twice and compares the complete
    bytes.
  2. Add an opt-in reproducible mode so existing build-time clock/random semantics do not
    change by default.
  3. Give pre-initialization fixed clocks, random sources, Python hash seed, allocator state,
    and bytecode behavior in that mode.
  4. Canonicalize generated directory metadata and provide input files through a staged or
    virtualized tree. The implementation should not modify user source-file metadata.
  5. Keep output ordering deterministic throughout stub generation and component encoding.
  6. Add a reduced test around the init canonical-ABI boundary. Candidate fixes include capturing
    memory before post-return cleanup, making that cleanup allocator-deterministic, or changing
    the private init protocol so the large nested argument graph does not leave volatile allocator
    state in the captured memory. Any fix must preserve matching allocator globals and memory.

Would the maintainers prefer this as one opt-in feature PR, or as smaller PRs after the
stub-ordering fix?

Additional context

This was found while building a digest-pinned Python/WASI runtime. Reproducibility is a
supply-chain requirement there: the generated component digest is part of the runtime
identity, not merely a build-cache optimization.

Activity

  1. dicej commented on Sep 1, 2026

    @dicej
    Collaborator

    Thanks for reporting this. It looks like a duplicate of #243, which was also opened recently. Perhaps you and @composia-dave could collaborate on addressing this?

  2. dicej commented on Sep 1, 2026

    @dicej
    Collaborator

    Your proposed plan sounds fine to me. A single large PR for this would be fine if that's easiest; otherwise items 3 and 4 in your "Proposed direction" list could be PR'd independently. #243 mentioned wasmtime-wasi changes may be needed, also; that issue has since been edited, so you'll need to look at the earlier revision of the PR description for details.

  3. starsang1995 commented on Sep 2, 2026

    @starsang1995
    ContributorAuthor

    I dug into this a little further. It looks like the remaining differences come from host filesystem metadata entering CPython during import and surviving in memory captured by the snapshot. malloc_debug confirmed this, but it is too expensive and changes the allocator, so it does not seem like a good final fix.

    I tried an alternative implementation here: reproducible component generation. It canonicalizes build-time filesystem metadata and clears transient interpreter state before snapshotting. It produces identical bytes across independent source copies while keeping CPython's default allocator, with no measurable performance impact in a small benchmark.

    It is not as elegant as I would like because replacing the filesystem provider requires quite a bit of forwarding code, but it is the cleanest solution I have found so far without relying on allocator internals or rewriting snapshot memory.

  4. starsang1995 commented on Sep 3, 2026

    @starsang1995
    ContributorAuthor

    Yeah @dicej more permanent and elegant fix would need to be done at wasmtime-wasi side.

  5. rwbh commented on Oct 9, 2026

    @rwbh

    I tested the exact candidate revision linked above, starsang1995/componentize-py@25b4e4783fbc17d590206f8c1d9187ea5c841c7a, from source archive SHA-256 6200277a77c98831a86dd00963c55d9d7af286597cab0d2b2c4ae8140e603a with locked dependencies and WASI SDK 34.

    Both a Rust 1.97.1 release build and a clean-target build using the reported Rust 1.95.0 toolchain fail before tests/component generation:

    scrub-stack.o: undefined symbol: __stack_low
    

    The reviewed assembly reads __stack_low@GOT, and the build script links that object into componentize-py-runtime; neither toolchain resolves the symbol. Exact commands, normalized logs and resource bounds are retained here: https://github.com/wasmagents/browser-use-lab/blob/e961b0d0a82516b78e24d279adbbcd06713ecce6/docs/component-reproducibility-candidate.md

    This is only a buildability observation for that exact unmerged revision. It does not test the proposed reproducible mode or establish anything about a corrected version.

  6. rwbh commented on Oct 9, 2026

    @rwbh

    Correction to my earlier build report for starsang1995/componentize-py@25b4e4783fbc17d590206f8c1d9187ea5c841c7a: the candidate is buildable under the contributor-reported Rust 1.95.0 toolchain and WASI SDK 34 with one additional shared-link flag:

     -Clink-args=-shared \
    +-Clink-args=-Wl,--unresolved-symbols=import-dynamic \

    The existing global.get __stack_low@GOT assembly should remain unchanged. WASI SDK 34 uses LLVM/LLD revision 895aa2c896ada719451be2e3673c83da8ddf1141; at that revision the Wasm linker driver synthesizes optional __stack_low only for non-PIC links, so the PIC shared runtime needs to retain the GOT symbol as a dynamic import.

    After that change:

    • cargo +1.95.0 build --release --locked passed.
    • reproducible_mode_produces_identical_component_bytes passed 1/1.
    • Two larger downstream browser components built from an unchanged 13,696-file input inventory were byte identical: 57,869,636 bytes each, SHA-256 c8fbe59a8c1a0c38836af567c19232fa6005ef49e4aff6a6b43415729821282e.

    Downstream record and exact patch: wasmagents/browser-use-lab@bf19fb6. That repository may require access. This verifies the candidate plus the one-line repair; it does not address upstream merge readiness or downstream compatibility across the Python 3.14/Wasmtime 48/bindings changes.

  7. rwbh commented on Oct 9, 2026

    @rwbh

    Follow-up on the repaired candidate: the byte-identical downstream component requires a matching newer runtime boundary. Wasmtime 38.0.3 refuses before entry because it cannot satisfy wasi:cli/exit@0.2.12#exit-with-code; the official Wasmtime 48.0.0 CLI runs both component copies successfully.

    Under Wasmtime 48, two complete downstream public-site workflows each passed 26 guest/host exchanges, seven screenshots, exact CSV/report, policy and cleanup limits. Exact evidence and limitations are merged at wasmagents/browser-use-lab@e70556e (repository access may be required).

    This confirms the one-line linker repair plus matching runtime for this downstream case. It does not establish compatibility with older Wasmtime releases or upstream merge readiness.

  8. rwbh commented on Oct 9, 2026

    @rwbh

    Follow-up from the browser-use lab: the repaired candidate was also exercised against the actual package entry point, separate from the earlier public-site component.

    At exact lab revision 83988865e0a7c047c135704f1cb3ba27db85fb4c, identical before/after inventories over 13,700 input files (033aeeab…) produced two byte-identical validated browser_bridge_bootstrap components at 75b5f52e… (57,835,694 bytes). Packaging each output also produced identical archives at 27a8c5a7…; an installed archive completed both the 22-exchange Python source workflow and unchanged 11-exchange Rust consumer under Wasmtime 48.

    A useful negative control: the earlier byte-identical public-site component c8fbe59a… failed after one exchange when placed in the owned-fixture package because the response schemas differ. That failure reinforces that reproducibility and workflow completion must be checked for each embedded entry point.

    Evidence and commands: checkpoint, component comparison, work record.

    The tested fork remains unmerged and uses the lab's one-line --unresolved-symbols=import-dynamic repair plus a generated-bindings compatibility shim. This is additional experimental evidence, not an adoption or upstream fix claim.

  9. rwbh commented on Oct 9, 2026

    @rwbh

    Downstream browser lifecycle qualification for exact candidate revision 25b4e4783fbc17d590206f8c1d9187ea5c841c7a merged in wasmagents/browser-use-lab#141 at 2acaa1bf659df57f0d63750d96be29ffc981548b.

    The repaired candidate produced the previously reported byte-identical source component 75b5f52e… and reproducible archive 27a8c5a7…. Under exact Wasmtime 48, the installed source now additionally passes download/screenshot policy denial, outbound/inbound DLP, caller SIGTERM with acknowledged browser stop/HTTP EOF, and real Browser.crash fencing. Separate-process fresh recovery completes the full source and independent-consumer workflows. Exact evidence and limitations; merged-main CI 37882242790 passed.

    This remains downstream lab evidence for the unmerged fork plus the lab's one-line linker repair and bindings shim. It does not claim upstream acceptance, Wasmtime 38 compatibility, original governed-runtime qualification, or production adoption.

  10. rwbh commented on Oct 9, 2026

    @rwbh

    Downstream browser-use follow-up at exact candidate revision 25b4e4783fbc17d590206f8c1d9187ea5c841c7a: https://github.com/wasmagents/browser-use-lab/pull/142

    After the previously reported lab linker repair, two 57,835,045-byte owned-fixture components from one input snapshot are byte-identical at b1f26d55…. The component completes the full 22-request workflow under Wasmtime 48 and under the retained original XMP runtime/owner with a downstream stdio compatibility entry point.

    Observed limitation: the XMP runtime forwards and answers a newline-terminated stdout request, then Componentize-Py/Python 3.14 reports BrokenPipeError from both buffered flush() and direct os.write. The downstream shim remains fail-closed by proceeding only after a bounded response validates for the exact request sequence. This is environment-specific downstream evidence, not a proposed upstream fix; the candidate is still unmerged and the lab-maintained linker/bindings/stdio adaptations remain explicit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions