Repository navigation
senior-dev: read its window from codeaf offline, refuse tiny models, keep progress on compaction - #1791
Open
ZeroPoint95 wants to merge 5 commits into
Open
senior-dev: read its window from codeaf offline, refuse tiny models, keep progress on compaction#1791ZeroPoint95 wants to merge 5 commits into
ZeroPoint95 wants to merge 5 commits into
Conversation
…city rebuilds, pin per model Three compaction fixes from the cybergym smoke forensics: - The summary prompt now carries a Length-budget instruction sized from the same watermarks the trigger used (high minus continuation headroom, the kept tail, and a fixed baseline; clamped to a floor, absent on unbounded windows), so a generated summary is one the rebuild can hold. - The deterministic capacity stub no longer erases progress: it carries the previous summary verbatim (bounded) and the changed-files list the install loop would otherwise zero, beside the observed watermark numbers. - Overflow pins are keyed by provider/model, not session: every session routed to the same endpoint faces the same window, so the rejection tax is paid once per run and a second session inherits the pin. (cherry picked from commit d3864344172d10df7aed2316c113999749fc6667)
…dows, keep progress through capacity rebuilds
With models.dev out of reach, senior-dev sized every model at 16,384
tokens, though codeaf's own catalog on the same machine listed the real
window. A 1M-token model then compacted every few thousand tokens, and
113 of 115 rebuilds installed a stub that erased the session's progress.
- Sizing: models.dev first, then the catalog codeaf keeps for the
profile (config.CachedContextWindow: disk only, no key, no network),
then the 16K guess. A guessed window is now said, in the run's log and
on its page ("could not learn how much its model holds; assumed 16,384
tokens").
- A model known to hold 32,768 tokens or fewer is refused before its
first call, by name and size; a guess is never refused.
- A senior-dev shell run waits up to one catalog fetch
(catalog.FetchTimeout) for the profile's model list, so a fresh profile
has the file senior-dev reads.
- Hardening of the cherry-picked compaction fix: the summary budget
measures the system prompt, tools and pins instead of assuming 2,048,
gives the tail up before the summary, and never exceeds the
summarizer's output limit. The capacity stub is measured bare first,
then carries the NEWEST summary cut to the room left, and falls back to
the bare stub rather than ending the session when a carry does not fit.
- A 16K reproduction drives three compactions and checks that progress
survives each one; it fails on dev as it is.
Manual and change entry updated in the same change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and 1.27 The Go 1.26 gofmt CI runs breaks a map's alignment around a far shorter key, and the empty-string case did exactly that. The blank-model case is now asserted on its own line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rees with Go 1.26.5's gofmt, which CI runs, splits a map literal's alignment between long and short keys where 1.27's does not, so the two disagreed on this file. A slice of cases has no key column to align. Every changed Go file is now clean under both versions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y what the capacity stub keeps Review follow-ups on #1791: - A crew seat known to hold 32,768 tokens or fewer is now left out with a note, as a seat nothing can size already was, and a coder's pool it empties routes on senior-dev's own list. Only models a person asked for are refused by name. The manual's crew passage and its new section say the same. - The manual no longer says senior-dev never throws its progress away: when none of the newest summary fits, only the changed-files list carries over. - The shell host's catalog wait is covered: the end-to-end senior-dev shell test checks that it runs once, before any model call, and is bounded by catalog.FetchTimeout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ZeroPoint95
marked this pull request as ready for review
October 7, 2026 18:54
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A CyberGym do-stage run in a container that could not reach models.dev sized
deepseek/deepseek-v4.1-flash(1,048,576 tokens) at senior-dev's offline guess of 16,384. It compacted 115 times in 25 minutes. 113 of those rebuilds installed a stub that erased the session's progress ("No generated completion claim survived the capacity rebuild."). The agent re-read its brief 83 times and delivered nothing. codeaf's own model catalog in that same container already listed the real window.What changes
Sizing (the root cause)
config.CachedContextWindow: disk only, no key, no network), and only then the 16K guess.[senior-dev]line in the log, and acompaction-capacity/guessedrecord that reads on the run's page as "could not learn how much its model holds; assumed 16,384 tokens". A guess is never refused.catalog.FetchTimeout(15s) for the profile's model list before it starts, so a fresh profile has the file senior-dev reads.Tiny models refused up front
senior-dev cannot work with <model> (<n> tokens): ….Compaction hardening (on top of the cherry-picked
d38643441)Proof
TestASixteenKSessionKeepsItsProgressThroughRepeatedCapacityRebuildsdrives three compactions on a 16K window and checks that progress survives each one. It fails ondevas it is.make pr-readypasses. Onecmd/codeaftest,TestRunSurfaceWiresTheDeferredLaunchCheckAndInstallerThroughRealInit, fails ondevas well (from Keep codeaf current with quiet background updates that preserve running work #1790) and was classified "already failing on the base".Manual
senior-dev.md: a new section, "senior-dev keeps compacting, or refused a model as too small". The passages that said the run uses "conservative limits" offline are rewritten. Three new probes.Not in this PR
The
.senior-dev"directory vanishes mid-run" report was the benchmark's ownvalidate.py(restore_srcrunssudo rm -rf /src). That is being fixed in the rig, not here.🤖 Generated with Claude Code