diff --git a/.claude/skills/memory-hygiene-in-rust/SKILL.md b/.claude/skills/memory-hygiene-in-rust/SKILL.md new file mode 100644 index 00000000..8605be0c --- /dev/null +++ b/.claude/skills/memory-hygiene-in-rust/SKILL.md @@ -0,0 +1,246 @@ +--- +name: memory-hygiene-in-rust +description: + How to keep peak stack memory down for free in any Rust crypto code, and how to measure it honestly. Applies to every crate under crypto/*, not only the *_lowmemory ones: those crates trade performance for memory by algorithmic design, whereas everything here costs nothing and there is no reason to be wasteful anywhere. Use this whenever work adds or changes a function signature, return type, constructor, key or signature decoder, or anything holding a polynomial, vector, matrix or byte array of a kilobyte or more; whenever a change touches `mem_usage_benches` or a "Memory Usage" table in a crate's docs; and for any claim about copies, stack frames, out-parameters, `.try_into()`, return-by-value, `#[inline(never)]`, `.clone()`, `Copy` derives, valgrind/massif, or "does this change affect memory". Also use it when reviewing a PR that says a change is "free", "zero-cost" or "just a small struct", and when a bench number moved and nobody knows why. This work is easier to get right with a strong model (Fable or higher); if the session model is weaker, say so and measure twice as much. +--- + +# Memory Hygiene in Rust + +Peak stack usage is a property of the compiled binary, not of the source. Two pieces of code that are equivalent to a +human can differ by tens of kilobytes at run time depending on what LLVM inlines, which copies it elides, and how the +*caller* is shaped. Everything below follows from that. The single rule: **a claim about stack cost is a hypothesis +until massif has measured it.** +Type size is not a result. "It's only 32 bytes" cost 16 bytes of peak in this repo; "it's a 56 kB matrix, of course it +copies" turned out to be already elided. Measure, then say. + +Read `references/case-studies.md` when you want the measured numbers behind any rule here; each rule cites its case. + +**Scope: every crate, not just the low-memory ones.** In this library "lowmemory" names an algorithmic trade: those +crates drop intermediate values immidiately after use and re-derive them as-needed on-demand instead of holding them, +and pay for it in throughput. That trade is a design decision confined to the +`*_lowmemory` crates and is out of scope here. This skill is about the other kind of saving, the copies and frames that +cost nothing to remove and nothing to keep removed: a borrowed field instead of a clone, an out-parameter instead of a +return slot copy, a repeat expression instead of an array `map`. Those apply identically to `mldsa`, `mlkem`, the hashes +and everything else under +`crypto/`. Being wasteful in a full-featured crate is not a feature of that crate. + +## 1. Where copies come from + +Rust moves are memcpys unless the optimizer proves it can build the value in place. It often can, and the exceptions are +what this section is about. In descending order of how often they bit: + +- **Return-by-value of a large type.** The callee builds the value in its own frame and copies to the caller's slot, + unless LLVM forwards the slot. Passing an out-parameter (`out: &mut T`) makes the destination explicit. The lowmem + crates use this everywhere; the full crates should too. *But see §2: converting a tail-call return into an + out-parameter can add a copy.* +- **Tuples and `Result`/`Option` of large types.** `fn f() -> Result<(A, B, C), E>` for kilobyte + `A, B, C` left three copies live at once when not inlined: the local, the return slot, and the destructured bindings. + Write into caller buffers and return `Result<(), E>`. +- **`.try_into()` from a slice.** `let x: [u8; N] = s.try_into().unwrap()` copies N bytes. + `let x: &[u8; N] = s.try_into().unwrap()` borrows. Always write the type; the difference is one + `&`. +- **`.clone()` of a stored field to pass a reference.** `f(&self.matrix.clone())` when + `f(&self.matrix)` was possible. Give the field a `pub(crate)` or a `&`-returning accessor. +- **Array `map` / `array::from_fn` to build a big array.** `[[(); l]; k].map(|_| ...)` goes through `core::array::drain` + and materialises the whole array in a temporary before copying it into place. Use a repeat expression instead: + `[[Polynomial::new(); l]; k]` for a `Copy` element, or `[[ZERO; l]; k]` with a `const ZERO` for any type. +- **Copying a `Copy` type by naming it.** `let r0 = w;` on a `Copy` polynomial vector is a copy unless `w` is dead + afterwards and LLVM reuses the slot. Reuse the buffer explicitly (`w.sub_vector(&x)` in place) when the old value is + not read again, and say so in a comment. +- **Pass-by-value parameters** for anything bigger than a couple of machine words. Take `&T`. +- **Building `Self` in a local, mutating it, returning it.** + `let mut s = Self { a, big: Big::new() }; fill(&mut s.big); s` was three copies of `Big` + (local, return slot, binding). Build parts in locals and return a struct literal, or better, expand directly into a + caller-owned struct through an in-place method. + +### The structural answer: kilobyte types should not be `Copy` + +Most of the list above exists because `Polynomial` (1 kB) and `Vector` (4 to 8 kB) derive +`Copy`. A `Copy` type duplicates silently: `let r0 = w;` is an 8 kB memcpy the compiler never mentions, and whether it +survives depends on the optimizer. Wrapping the array in a newtype that derives `Clone` but not `Copy` changes the +accounting rather than the code: moves still compile to the same memcpys and are elided exactly as before, but every +duplication is now either a move (the source becomes unusable, so it cannot be an accident) or an explicit `.clone()` +that shows up in a grep. Implicit copies become compile errors, and the copy audit in §3 becomes complete by +construction. `Secret` already uses this shape for zeroization, so the pattern is established. + +Two couplings to handle when doing it here: + +- `ZeroizablePrimitive` currently requires `Copy` on the wrapped type as a proxy for "has no + `Drop`", so `Secret` stops compiling until that bound is replaced by a sealed marker trait that says the + same thing directly. +- Array repeat expressions need `Copy` or a `const` operand. `[[Polynomial::new(); l]; k]` becomes + `[[ZERO; l]; k]` with `const ZERO: Polynomial = Polynomial::new();`, which is allowed for any type. + +## 2. Frames, not values: what actually sets the peak + +Peak = the deepest stack pointer ever reached = sum of all frames live at that instant. Frames are allocated in full at +function entry, so what matters is *which function a value's slot lands in* +and *what else is live in that function*. + +- **Inlining merges frames.** When `verify` is inlined into its caller, verify's 36 kB of working vectors are allocated + for the caller's entire lifetime, including while the caller is still doing something else (expanding a key, + decoding). Two frames that would have been siblings become one large frame. This is the mechanism behind most "why did + the number move?" mysteries. +- **A `let` in an unused `match` arm still costs its full size.** The slot exists on every path. + `None => { let mut a = Matrix::new(); ... }` charged every `Some` caller 57 kB. Move the arm's body into a separate + `#[inline(never)]` function so the slot exists only when taken. +- **Every boundary costs a copy of what crosses it.** A non-inlined call returning a 2.4 kB signature by value costs 2.4 + kB in the caller and 2.4 kB in the callee. A closure that returns a value has the same cost. If you add a boundary to + isolate a frame, make nothing cross back. +- **`#[inline(never)]` is a tool, not a fix.** It guarantees a frame is popped before the next call, which is exactly + right for one-shot setup such as key loading or decoding. It costs a few kB of spills on hot paths and should never be + sprinkled on the algorithm itself without a measurement showing why. +- **LLVM's copy elision is fragile.** A tail call `fn a() -> M { expand(rho) }` has its return slot forwarded straight + through; the copy never exists. Rewriting `expand` to an out-parameter turns that into + `let mut m = M::new(); expand(rho, &mut m); m`, a local plus a move that LLVM does *not* + reliably elide, so the "optimisation" added a 57 kB copy. Whenever you change a return convention, measure both the + callers that were fine and the ones you were fixing. + +## 3. Rules of engagement for a change + +1. **Baseline first.** Measure the benches the change could touch on the unmodified tree. Use a git worktree if you need + to keep editing meanwhile; never `git stash` while a bench runner is rewriting files in the same tree. +2. **Change one thing.** Each structural edit (a signature, a constructor shape, an inlining attribute) can move numbers + in unrelated benches through layout. If you bundle three edits and the numbers move, you will not know which one did + it. +3. **Measure after, in bytes, for every affected bench**, not a sample. Report a table of before / after / delta and + state the *cause* of each delta, not just the sign. Deltas you cannot explain are not done. +4. **Verify outputs are identical** (same signature, same shared secret, verify still succeeds). A memory optimisation + that changes output is a bug, and the bench is the cheapest place to catch it. +5. **Audit every `.clone()` on a kilobyte-sized type**, and put each into one of four bins: + - *Field-wise `Clone` impl of a key struct*: the copy is the point. Keep. + - *Storing a key or seed into long-lived state* (a streaming verifier's `Option`): once per session. Keep. + - *Clone-then-transform*: `let mut y_hat = y.clone(); y_hat.ntt();`. A second buffer is real, but + copy-then-transform writes every coefficient twice. Prefer `ntt_into(&y, &mut y_hat)` + when the transform can write its output directly; same peak, less work, explicit destination. + - *Clone to hand out a reference*: `f(&self.matrix.clone())`. Never. Borrow the field or add a + `&`-returning accessor. This bin held the 57 kB case. Twenty-odd clones per crate is normal; the audit is a + ten-minute grep and it is where the large, avoidable copies hide. With non-`Copy` newtypes (§1) these bins are the + *only* + duplications in the crate. +6. **Prefer structural fixes over lucky ones.** If a copy disappears only because the inliner happened to cooperate, it + will reappear when the caller changes shape. The fix is real when the copy is *impossible*: an out-parameter, a + repeat expression, a borrowed field. +7. **Expect noise and know its size.** On this repo's harness, lowmem crates move by about ±0.1 kB between layouts; full + crates by ±1 to ±9 kB. A change inside that band on a full crate is not evidence of anything. Run twice if it + matters; massif itself is deterministic. +8. **Never trade a table-row increase for an off-table benefit silently.** If a change helps + `Verify_with_expanded_key` (no table row) and costs plain `Verify` (a table row) 6 kB, say so and let the maintainer + choose. + +## 4. Building a measurement harness + +The repo's harness is `mem_usage_benches/`: one binary per algorithm family, `main()` calls exactly one bench function, +and the peak is read from valgrind massif with `--heap=no +--stacks=yes`. `scripts/massif_peak.sh` in this skill does the whole loop for a list of bench functions. What it took +several iterations to learn about the *shape* of a bench function: + +```rust +/// Loads the key. #[inline(never)] so its byte array and decode temporaries are popped +/// before the operation runs; otherwise the bench reports load + op instead of max(load, op). +#[inline(never)] +fn load_mldsa44_sk() -> MLDSA44PrivateKey { + MLDSA44PrivateKey::from_bytes(&[ /* hard-coded encoded key */ ]).unwrap() +} + +/// Runs the operation in its own frame. Returns nothing so no result crosses the boundary. +#[inline(never)] +fn measure(f: impl FnOnce()) { f() } + +fn bench_mldsa44_sign() { + eprintln!("MLDSA44/Sign"); + let sk = load_mldsa44_sk(); + let msg = b"..."; + measure(|| { + let mu = MLDSA44::compute_mu_from_sk(&sk, msg, None).unwrap(); + let sig = MLDSA44::sign_mu_deterministic(&sk, None, &mu, [0u8; 32]).unwrap(); + print!("{:x?}", sig); // inside: keeps the optimizer honest, and nothing is returned + }); +} +``` + +Why each part is the way it is: + +- **Hard-coded keys, never `keygen` in the bench.** Keygen's frame is often larger than the operation's, so the old + benches were reporting keygen. Dump the encoded key once from a commented-out setup block and paste the bytes. +- **Key load in a non-inlined helper.** The operation gets inlined into whatever function calls it, so that function's + frame is allocated on entry; a load done in the same function stacks on top of it. Sibling frames give max (load, op); + parent/child gives the sum. This mattered by up to 14.6 kB per row. +- **Operation in a non-inlined, non-returning closure.** For the same reason, in the other direction: it keeps the + operation's frame from being allocated during the load, and it keeps a by-value result from being copied back. +- **Inputs (ciphertext, signature) on the stack in the bench body, never in a heap `Vec`.** + Massif with `--heap=no` does not measure the heap cheaply, it ignores it. A `hex::decode` + into a `Vec` silently deleted 3 to 5 kB from two table rows for months. If you want to exclude inputs from the number, + put them in `static` storage and say so in the methodology text; the heap is never the honest answer for a `no_std` + target. +- **Expanded-key variants** expand inside `measure`, since the expansion is the cost being measured, but from a plain + key loaded by the helper. +- **Keep the do-nothing baseline** (`bench_do_nothing`) and quote it under the table; a row equal to it is at the noise + floor, not a measurement. + +Massif gives a time series too. When a peak moves, look at *when* it occurs (`time=` vs +`mem_stacks_B=` in the massif file) before theorising: a peak at 4 % of the run is key loading, a peak at 40 % is the +operation. + +## 5. Finding the copy: frame layout remarks + +Do not read `objdump` prologues for this: any frame over 4 kB is allocated by a stack-probe loop, so every large +function shows `sub rsp, 0x1000`. Use LLVM's remarks, which work on stable: + +``` +RUSTFLAGS="-C remark=stack-frame-layout -C remark=prologepilog -C debuginfo=1" \ + CARGO_TARGET_DIR=/tmp/remarks cargo build --release -p mem_usage_benches --bin 2> remarks.txt +python3 .claude/skills/memory-hygiene-in-rust/scripts/frame_layout.py remarks.txt [name-filter] +``` + +`prologepilog` prints `N stack bytes in function 'name'` per function; `stack-frame-layout` +lists every stack object with its size. Compare the two builds (before/after, or slice-fed vs array-fed) and the extra +copy is the object that appears twice or the frame that appeared at all. Two traps: names are v0-mangled (the script +demangles crudely), and **filter on nothing at first**: the 57 kB `core::array::drain::drain_array_with` frame hid for +an hour behind a +`grep mldsa`. Use a separate `CARGO_TARGET_DIR` so the remark build does not invalidate the normal one, and note +`debuginfo=1` did not change any frame size in practice. + +## 6. A different class of copy: byte views, and the `zerocopy` crate + +Everything above is about copies of *typed* values: moves at return boundaries, temporaries, frames. There is a second +class this repo has not yet hit: copying bytes out of a buffer just to give them a type, i.e. parsing a wire format into +a struct by memcpy when the bytes were already laid out correctly. Signs that you are looking at it: a `from_bytes` that +does nothing but +`copy_from_slice` field by field into a `#[repr(C)]`-shaped struct, a `[u8; N]` that is only ever reinterpreted as +`[u32; N/4]`, or a parser whose output is byte-identical to its input. + +For that class the right tool is the `zerocopy` crate (Google; v0.8.x, `no_std`): derive +`FromBytes`, `IntoBytes`, `KnownLayout`, `Immutable` or `Unaligned` on the type and view a `&[u8]` +as a `&T` or `&[T]` with no copy and no `unsafe` in this repo's code. The `unsafe` lives inside the crate, which is one +of the most heavily reviewed in the ecosystem and is what the Rust standard library ecosystem itself leans on for this +job. + +Two cautions. First, check that the problem is really reinterpretation: the ML-DSA and ML-KEM decoders bit-unpack 10-, +13-, 18- and 20-bit fields into `i32` arrays, which is a transformation, so there is nothing there for `zerocopy` to +remove, and the only pure reinterpretation in those paths, `&[u8]` to `&[u8; N]`, is already free via `try_into` on a +reference. Second, QUALITY_AND_STYLE.md's rule is zero external runtime dependencies, and every crypto crate is +`#![forbid(unsafe_code)]`. `zerocopy` is a plausible candidate for a deliberate exception to the dependency rule on +reputation grounds, but that is a maintainer decision to raise explicitly in the PR, with the measured copy it removes, +not something to pull in quietly. + +## 7. Related lessons that are not about memory but were learned the same way + +- **A step that only matters in the worst case is invisible to `cargo mutants` and KATs.** The plain `reduce32` before + `inv_ntt` was deleted in March 2026 because mutants and the full bc-test-data set passed without it; six months later + a Wycheproof vector showed it let a forgery through. Before deleting a reduction, range check, or bound-keeping step + because "nothing fails", find the input-bound argument it serves. Prefer a `debug_assert!` on the bound to deletion. +- **Perf claims get the same treatment as memory claims.** Criterion with `--save-baseline` / + `--baseline`, all benches, then rerun any single outlier with a longer window before believing it. A 5 % "regression" + on one bench whose siblings show −0.5 % is noise until it reproduces. +- **In-repo tests are for developer mistakes; adversarial vectors live in bc-test-data and wycheproof.** Do not copy + external vectors into the repo as regression tests. Do add a unit test of a helper's contract (congruence, range, + boundaries) when you add the helper. + +## 8. Model note + +Most of the wrong turns in the history behind this skill were plausible-sounding hypotheses about what the compiler did, +stated confidently and then contradicted by measurement. The work needs a model that will hold a claim loosely, run the +measurement, read a stack-frame dump without filtering it to what it expects, and change its mind in public. That has +worked with Fable-class models; with a weaker model, insist on the before/after table for every change and do not accept +"should be free" from it or from yourself. diff --git a/.claude/skills/memory-hygiene-in-rust/references/case-studies.md b/.claude/skills/memory-hygiene-in-rust/references/case-studies.md new file mode 100644 index 00000000..6143fd0c --- /dev/null +++ b/.claude/skills/memory-hygiene-in-rust/references/case-studies.md @@ -0,0 +1,120 @@ +# Case studies behind the rules (bc-rust, September 2026) + +Every number is a massif peak in bytes on the `mem_usage_benches` harness, release build, x86_64, +measured before and after a single change. "Full" is `bouncycastle-mldsa` / `-mlkem`; "lowmem" is +the `-lowmemory` crate. + +## 1. "Only 32 bytes" was not free + +`make_hint_row` in mldsa-lowmemory took `out: &mut HintRow` (32 bytes) and returned the weight. +Changing it to return `(HintRow, i32)` by value, on the argument that 32 bytes cannot matter, added +exactly **+16 bytes** of peak on all three lowmem sign benches. Returning just `HintRow` and computing +the weight with popcount: also +16. Reverted. The sibling `unpack_h_row` on the verify path went to +`Option` at measured **0** bytes, because it was already the deepest frame there. Same +type, same size, opposite results: the frame context decides, not the type. + +## 2. Decaps benches were measuring keygen + +Every ML-KEM decaps bench called `keygen_from_seed` inside the measured binary. Hard-coding the +encoded private key instead: + +| | old | in-main decode | helper decode | +|---|---|---|---| +| ML-KEM-512 decaps | 24136 | 28408 | 21592 | +| ML-KEM-768 decaps | 39800 | 44072 | 33416 | +| ML-KEM-1024 decaps | 63800 | 62792 | 49224 | + +The "in-main" column is the first attempt, with the key byte array and `from_bytes` in the bench +body: worse than keygen for 512 and 768, because the array and the `Result` temporary stayed live +under decaps. Moving the load into an `#[inline(never)]` helper gave the third column. The old table +had over-reported decaps by up to 14.6 kB. + +## 3. Sign 44 went up 8.7 kB with no code change in sign + +After item 2's helper pattern was applied to ML-DSA, `Sign/ML-DSA-44` rose 93720 → 102456 with +identical instruction counts. Frame remarks: the bench `main` was 73608 bytes in both builds with +sign fully inlined into it. The key-load helper (12.5 kB) and `from_bytes` (14.6 kB) were being +called from that `main`, so they stacked on top of the already-allocated 73.6 kB frame. Fix: run the +operation in a separate non-inlined, non-returning closure so load and op are sibling frames. +Residual after the fix: **+2.1 kB**, the boundary's own spills. A wrapper that *returned* the +signature cost 4.6 kB instead: the 2.4 kB result was copied across the boundary. + +## 4. Plain/expanded pairs reporting identical peaks + +Before the harness rework, every `Encaps`/`Encaps_expanded_pk` pair and every +`Sign`/`Sign_expanded_sk` pair reported byte-identical peaks. That only happens when the key decode +in the bench body, not the operation, sets the peak. Treat identical numbers across variants that +do different work as a harness bug, not a coincidence. + +## 5. A heap `Vec` deleted 3 to 5 kB from two table rows + +`bench_mldsa65_lowmemory_verify` and the 87 variant decoded their signature with `hex::decode` into +a `Vec`. Massif `--heap=no` ignores the heap, so those rows under-reported by one signature +each (3309 and 4627 bytes) relative to the 44 row, which used a stack array. Converting them moved +the rows 15864 → 19096 and 17784 → 22328. + +## 6. `sig_decode` returned an 11 kB tuple + +`fn sig_decode(sig) -> Result<(SigCTilde, VecL, VecK), ()>`. When LLVM inlined it, fine. When fed a +runtime-length slice it did not inline it, and the remarks showed three copies of the 11.3 kB tuple +live at once (local, return slot, destructured bindings) plus `sig_decode`'s own 23.6 kB frame under +verify: **+18 to +24 kB** depending on caller shape. Out-parameters (`c_tilde: &mut, z: &mut, +h: &mut`, returning `Result<(), ()>`) removed it structurally: full-crate verify −6 to −20 kB across +parameter sets, 44 sign −9 to −12 kB. + +## 7. `Matrix::new()` built a 56 kB temporary through `array::map` + +`Self { elems: [[(); l]; k].map(|_| [(); l].map(|_| Polynomial::new())) }` goes through +`core::array::drain::drain_array_with`, which materialises the full matrix before copying it. It was +invisible until the fix in item 6 changed inlining and a **57368-byte** `drain_array_with` frame +appeared under `expandA`, adding 47 kB to one bench. Repeat expression `[[Polynomial::new(); l]; k]` +(elements are `Copy`, `new` is `const`) removed it. The same pattern in the full `mlkem` crate was +*not* being materialised (its matrix is 4 to 16 kB and LLVM folded it); changing it there moved +three table rows up 2 to 3 kB from inlining alone, so it was reverted. Structural correctness did +not win over measurement. + +## 8. An unused `match` arm cost 57 kB + +```rust +match a_hat { + Some(a) => sign_internal(sk, a, ...), + None => { let mut a = Matrix::new(); sk.expand_into(&mut a); sign_internal(sk, &a, ...) } +} +``` +The `None` arm's local is allocated in the frame on both paths, so every expanded-key caller paid +for a matrix it never used (`sign_mu_deterministic` frame 62 kB). Fix: the arm calls an +`#[inline(never)]` helper that owns the local. Cost: 5 to 6 kB on the plain paths for the extra +boundary, which was accepted. + +## 9. Converting a return to an out-parameter added a copy + +`fn A_hat(&self) -> M { expandA(&self.rho) }` was a tail call; LLVM forwarded the return slot into +`expandA`, so only `expandA`'s own local existed. After `expandA` took `&mut M`, `A_hat` became +`let mut m = M::new(); expandA(&self.rho, &mut m); m`, and LLVM did not elide the move: the by-value +`A_hat()` path gained a full extra matrix. Keygen and plain verify improved by 13 to 49 kB and 12 to +50 kB from the same change, expanded-key construction regressed by 10 to 53 kB. Constructors of the +form `let mut s = Self { big: M::new(), .. }; fill(&mut s.big); s` showed three copies; returning a +struct literal built from locals showed two. Getting to one copy for a by-value constructor was not +achievable in safe Rust without an out-parameter API; that became a follow-up. + +## 10. The deleted reduction + +Not a memory case, but found with the same tools and worth keeping next to them. A plain `reduce32` +before `inv_ntt` existed from the first ML-DSA commits, was commented out on 2026-03-20 because +`cargo mutants` and the full bc-test-data set passed without it, and was deleted the next day. In +September 2026 Wycheproof's `MissingReduction` vectors showed the 8-level inverse-NTT butterflies +overflow `i32` when fed an unreduced sum of `l+1` Montgomery products: a valid signature rejected and +a forgery accepted in release. Restored at all nine accumulation sites in both crates; Criterion +showed the cost below the noise floor (median −0.7 % across 30 benches, one +5 % outlier that +re-ran at −1.3 %). Bound-keeping steps are dead code to every test that uses honest inputs. + +## Measured noise floors on this harness + +| Crate family | Layout noise between builds | +|---|---| +| lowmem (ML-KEM, ML-DSA) | ≤ 0.2 kB | +| full ML-KEM | 1 to 4 kB | +| full ML-DSA | 2 to 9 kB | + +Massif itself is deterministic: identical binaries give identical peaks. The noise is in what the +compiler does with a differently shaped caller, not in the measurement. diff --git a/.claude/skills/memory-hygiene-in-rust/scripts/frame_layout.py b/.claude/skills/memory-hygiene-in-rust/scripts/frame_layout.py new file mode 100755 index 00000000..3c4966fe --- /dev/null +++ b/.claude/skills/memory-hygiene-in-rust/scripts/frame_layout.py @@ -0,0 +1,69 @@ +#!/usr/bin/env python3 +"""Summarise rustc/LLVM stack-frame remarks: per-function frame sizes and large stack objects. + +Build with (stable toolchain is fine): + RUSTFLAGS="-C remark=stack-frame-layout -C remark=prologepilog -C debuginfo=1" \ + CARGO_TARGET_DIR=/tmp/remarks cargo build --release -p --bin 2> remarks.txt +then: + frame_layout.py remarks.txt [name-regex] [--min-object BYTES] [--top N] + +Prints the largest frames ('N stack bytes in function' from the prologepilog remark) and, for each +matching function, its stack objects at or above --min-object (from the stack-frame-layout remark). +Names are demangled crudely from the v0 scheme. Start with NO name filter: the copy you are looking +for is often in a std/core frame such as core::array::drain::drain_array_with. +""" +import re, sys + +def demangle(sym): + parts = [] + for m in re.finditer(r"(\d+)(_?)([A-Za-z_][A-Za-z0-9_]*)", sym): + n = int(m.group(1)); ident = m.group(3)[:n] + if len(ident) == n and not ident.startswith("Cs"): + parts.append(ident) + return "::".join(parts) if parts else sym + +def main(): + args = sys.argv[1:] + if not args: + print(__doc__); sys.exit(1) + path = args.pop(0) + min_obj = 4096; top = 20; pattern = None + while args: + a = args.pop(0) + if a == "--min-object": min_obj = int(args.pop(0)) + elif a == "--top": top = int(args.pop(0)) + else: pattern = re.compile(a) + text = open(path, errors="replace").read() + text = re.sub(r"_R[A-Za-z0-9_]+", lambda m: demangle(m.group(0)), text) + text = re.sub(r"_ZN[0-9]+_?", "", text) + text = re.sub(r"17h[0-9a-f]{16}E", "", text) + for a, b in (("$LT$", "<"), ("$GT$", ">"), ("$u20$", " "), ("$C$", ","), ("$RF$", "&"), ("..", "::")): + text = text.replace(a, b) + + frames = {} + for m in re.finditer(r"(\d+) stack bytes in function '([^']+)'", text): + frames[m.group(2)] = max(frames.get(m.group(2), 0), int(m.group(1))) + objects = {} + cur = None + for line in text.splitlines(): + m = re.match(r"\s*Function: (.*)", line) + if m: cur = m.group(1).strip(); objects.setdefault(cur, []); continue + m = re.match(r"\s*Offset: \[SP[-+]\d+\], Type: (\w+), Align: \d+, Size: (\d+)", line) + if m and cur is not None: + objects[cur].append((int(m.group(2)), m.group(1))) + + noise = re.compile(r"backtrace|gimli|driftsort|panicking|rustc_demangle|std::sys|addr2line|miniz") + rows = sorted(((sz, n) for n, sz in frames.items() if not noise.search(n) and (pattern is None or pattern.search(n))), reverse=True) + print(f"largest frames (top {top}):") + for sz, n in rows[:top]: + print(f"{sz:9d} {n[:140]}") + print(f"\nstack objects >= {min_obj} bytes:") + for sz, n in rows[:top]: + big = sorted((o for o in objects.get(n, []) if o[0] >= min_obj), reverse=True) + if big: + print(f" {n[:120]}") + for osz, kind in big: + print(f" {osz:8d} {kind}") + +if __name__ == "__main__": + main() diff --git a/.claude/skills/memory-hygiene-in-rust/scripts/massif_peak.sh b/.claude/skills/memory-hygiene-in-rust/scripts/massif_peak.sh new file mode 100755 index 00000000..11bf2d9b --- /dev/null +++ b/.claude/skills/memory-hygiene-in-rust/scripts/massif_peak.sh @@ -0,0 +1,38 @@ +#!/usr/bin/env bash +# Peak stack of one or more mem_usage_benches functions, via valgrind massif. +# +# massif_peak.sh