Skip to content

senior-dev: read its window from codeaf offline, refuse tiny models, keep progress on compaction - #1791

Open
ZeroPoint95 wants to merge 5 commits into
devfrom
zeropoint95/senior-dev-small-windows
Open

ZeroPoint95 wants to merge 5 commits into
devfrom
zeropoint95/senior-dev-small-windows

Conversation

@ZeroPoint95

Copy link
Copy Markdown
Collaborator

Why

A CyberGym do-stage run in a container that could not reach models.dev sized deepseek/deepseek-v4.1-flash (1,048,576 tokens) at senior-dev's offline guess of 16,384. It compacted 115 times in 25 minutes. 113 of those rebuilds installed a stub that erased the session's progress ("No generated completion claim survived the capacity rebuild."). The agent re-read its brief 83 times and delivered nothing. codeaf's own model catalog in that same container already listed the real window.

What changes

Sizing (the root cause)

  • Senior-dev now sizes a model from models.dev, then the catalog codeaf keeps for the profile (config.CachedContextWindow: disk only, no key, no network), and only then the 16K guess.
  • A guessed window is said: a [senior-dev] line in the log, and a compaction-capacity/guessed record that reads on the run's page as "could not learn how much its model holds; assumed 16,384 tokens". A guess is never refused.
  • A senior-dev shell run waits up to catalog.FetchTimeout (15s) for the profile's model list before it starts, so a fresh profile has the file senior-dev reads.

Tiny models refused up front

  • A model known to hold ≤ 32,768 tokens is refused before its first call, with nothing spent: senior-dev cannot work with <model> (<n> tokens): ….

Compaction hardening (on top of the cherry-picked d38643441)

  • The summary length budget measures the system prompt, tools and pins instead of assuming 2,048. When the summary would have too little room beside the recent messages, those messages are dropped first. The budget never exceeds the summarizer's output limit.
  • The capacity stub is measured bare first, then carries the newest summary cut to the room left. If that still doesn't fit, it falls back to the bare stub instead of ending the session with a capacity error.
  • Kept from the cherry-pick: the stub carries the changed-files list, and overflow pins are per model, not per session.

Proof

  • TestASixteenKSessionKeepsItsProgressThroughRepeatedCapacityRebuilds drives three compactions on a 16K window and checks that progress survives each one. It fails on dev as it is.
  • New tests cover which source sizes a model, the 32,768 cutoff, a run refused before any stage, a connected service's own catalog compartment, and the line on the run's page.
  • make pr-ready passes. One cmd/codeaf test, TestRunSurfaceWiresTheDeferredLaunchCheckAndInstallerThroughRealInit, fails on dev as well (from Keep codeaf current with quiet background updates that preserve running work #1790) and was classified "already failing on the base".

Manual

  • senior-dev.md: a new section, "senior-dev keeps compacting, or refused a model as too small". The passages that said the run uses "conservative limits" offline are rewritten. Three new probes.

Not in this PR

The .senior-dev "directory vanishes mid-run" report was the benchmark's own validate.py (restore_src runs sudo rm -rf /src). That is being fixed in the rig, not here.

🤖 Generated with Claude Code

ZeroPoint95 and others added 5 commits October 7, 2026 13:27
…city rebuilds, pin per model

Three compaction fixes from the cybergym smoke forensics:

- The summary prompt now carries a Length-budget instruction sized from
  the same watermarks the trigger used (high minus continuation headroom,
  the kept tail, and a fixed baseline; clamped to a floor, absent on
  unbounded windows), so a generated summary is one the rebuild can hold.
- The deterministic capacity stub no longer erases progress: it carries
  the previous summary verbatim (bounded) and the changed-files list the
  install loop would otherwise zero, beside the observed watermark numbers.
- Overflow pins are keyed by provider/model, not session: every session
  routed to the same endpoint faces the same window, so the rejection tax
  is paid once per run and a second session inherits the pin.

(cherry picked from commit d3864344172d10df7aed2316c113999749fc6667)
…dows, keep progress through capacity rebuilds

With models.dev out of reach, senior-dev sized every model at 16,384
tokens, though codeaf's own catalog on the same machine listed the real
window. A 1M-token model then compacted every few thousand tokens, and
113 of 115 rebuilds installed a stub that erased the session's progress.

- Sizing: models.dev first, then the catalog codeaf keeps for the
  profile (config.CachedContextWindow: disk only, no key, no network),
  then the 16K guess. A guessed window is now said, in the run's log and
  on its page ("could not learn how much its model holds; assumed 16,384
  tokens").
- A model known to hold 32,768 tokens or fewer is refused before its
  first call, by name and size; a guess is never refused.
- A senior-dev shell run waits up to one catalog fetch
  (catalog.FetchTimeout) for the profile's model list, so a fresh profile
  has the file senior-dev reads.
- Hardening of the cherry-picked compaction fix: the summary budget
  measures the system prompt, tools and pins instead of assuming 2,048,
  gives the tail up before the summary, and never exceeds the
  summarizer's output limit. The capacity stub is measured bare first,
  then carries the NEWEST summary cut to the room left, and falls back to
  the bare stub rather than ending the session when a carry does not fit.
- A 16K reproduction drives three compactions and checks that progress
  survives each one; it fails on dev as it is.

Manual and change entry updated in the same change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and 1.27

The Go 1.26 gofmt CI runs breaks a map's alignment around a far shorter
key, and the empty-string case did exactly that. The blank-model case is
now asserted on its own line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rees with

Go 1.26.5's gofmt, which CI runs, splits a map literal's alignment between
long and short keys where 1.27's does not, so the two disagreed on this
file. A slice of cases has no key column to align. Every changed Go file
is now clean under both versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y what the capacity stub keeps

Review follow-ups on #1791:

- A crew seat known to hold 32,768 tokens or fewer is now left out with
  a note, as a seat nothing can size already was, and a coder's pool it
  empties routes on senior-dev's own list. Only models a person asked for
  are refused by name. The manual's crew passage and its new section say
  the same.
- The manual no longer says senior-dev never throws its progress away:
  when none of the newest summary fits, only the changed-files list
  carries over.
- The shell host's catalog wait is covered: the end-to-end senior-dev
  shell test checks that it runs once, before any model call, and is
  bounded by catalog.FetchTimeout.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ZeroPoint95
ZeroPoint95 marked this pull request as ready for review October 7, 2026 18:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant