Skip to content

feat(evals): add doc-sync quality evals and hillclimb loop - #2036

Draft
soheimam wants to merge 41 commits into
masterfrom
docs/improve-agent-doc-writing-skills
Draft

soheimam wants to merge 41 commits into
masterfrom
docs/improve-agent-doc-writing-skills

Conversation

@soheimam

Copy link
Copy Markdown
Contributor

What changed? Why?

Adds a lightweight way to measure and improve the quality of the base-std docs sync bot (.github/workflows/base-std-docs-sync.yml → scripts/sync-from-base-std/). Today nothing measures whether a reviewer would merge the bot's output: of 10 bot PRs, 1 merged (#1939, which humans rewrote ~85%), 1 closed, 8 still open. Reviewer comments show the same failures repeatedly: scope creep, paraphrasing the upstream changelog, wrong or ungrounded facts, housekeeping callouts, style slips.

Method follows "Automating eval design and hillclimbing" (Claude blog): real cases first, cheapest grader that works, claim-based LLM judge calibrated against a human, train/held-out split, keep a change only when both improve beyond noise.

Everything is new, under scripts/doc-evals/ (plan: PLAN.md, usage: README.md):

  • Replay harness (build-cases.mjs, cases/, replay/): 9 frozen cases built from past bot PRs (source sha, pre-sync docs commit, payload, human-merged reference where one exists, confirmed scope, verbatim review findings). replay/run.mjs reruns the sync in a throwaway git worktree against the historical docs tree. Nothing is pushed or opened.
  • Graders (graders/, grade.mjs, calibrate.mjs): code checks (scope precision/recall and forbidden paths, identifier grounding, selector recomputation via keccak-256, lint-mdx, housekeeping callouts, changelog shape and upstream fidelity, no-op), a yes/no claim judge, a blinded two-order pairwise vs the human reference, judge variance, and a human calibration tool. Overall = 0.4 scope + 0.25 code + 0.2 judge + 0.15 pairwise.
  • Metrics (metrics/, .github/workflows/doc-quality-report.yml): bot PR merge rate, human rewrite ratio, superseded PRs, review-comment taxonomy, candidate cases. Weekly, non-blocking; updates one issue titled "Docs sync quality report". No PR trigger.
  • Hillclimb (hillclimb/): proposes one root-cause change to llm/prompts.mjs or route-table.json, validates it against the sync's own tests in a scratch copy, replays train and test, and keeps it only if train improves beyond noise and held-out test improves. Test cases never reach the proposer. Budget cap, reflection after 2 non-keeps. Makes no commits and never touches the real sync code.

Owner decisions baked in (PLAN.md "Decisions"): changelog entries follow the upstream entry closely; the bot never edits docs/build-on-base/; evals run with CLAUDE_MAX_TOKENS=16000; failed judge calls are excluded from scores and block hillclimb decisions.

Notes to reviewers

  • Draft: baseline numbers and the first hillclimb result will be added to this description when the runs finish.
  • No changes to docs/, the sync code, or the existing workflows. The only non-new file is the root package.json test glob, which now includes scripts/doc-evals/__tests__/*.test.mjs. Those tests are offline and never load @anthropic-ai/sdk, since CI runs npm test without scripts/node_modules.
  • Findings about the sync itself, not fixed here:
    • At the default 4096 output tokens, long pages are truncated and rejected (one case always produces nothing). Proposed fix: set the CLAUDE_MAX_TOKENS repo variable to 16000.
    • Pages over 10,000 chars are skipped silently ("edit mode pending").
    • Symbol-manifest routing fans one source change out to 9–30 pages when 2–4 are right. This is the main driver of scope creep, and it lives in index.mjs, outside the hillclimb's edit surface.
    • A gateway 403 kills a whole sync run because it isn't retried.
    • base-std-routing.test.mjs ("upstream docs tree routes to the pages the IA guidelines assign") already fails on master.

How has it been tested?

  • npm test at the repo root: 346 tests pass, including with scripts/node_modules removed (CI mode; 1 skip for an SDK-only equality check).
  • Live smoke runs:
    • A replay of 64bd955 produced every output file and cleaned up its worktree.
    • grade.mjs scored that run end to end with the Opus judge; judge variance was 9/9 agreement per claim.
    • The metrics dry run printed 10 bot PRs, 1 merged, 1 closed, 8 open.
    • A one-round hillclimb ($4.15) proposed, validated, replayed, scored, and correctly reverted a patch.
  • Baseline replay (6 cases × 2 reps): the bot touched 9–30 pages per case where 2–4 were right, including 8–11 Build on Base pages on two cases. Rep-to-rep variation was low.

Screenshots

N/A (no user-facing changes; no rendered docs pages change).

Generated with Toshi

listBotPullRequests uses the list-PRs endpoint, which never returns
additions/deletions (confirmed live against base/docs and via the
GitHub REST docs) — only the single-PR endpoint does. merge-rate.mjs
was reading pr.additions/pr.deletions straight off list-endpoint
objects, so humanRewriteRatios silently came back empty on every real
run even though the offline fixture (which sets those fields by hand)
made the unit test pass. Add fetchPRTotals and use it for merged bot
PRs only.
@cb-heimdall

Copy link
Copy Markdown
Collaborator

🟡 Heimdall Review Status

Requirement Status More Info
Reviews 🟡 0/1
Denominator calculation
Show calculation
1 if user is bot 0
1 if user is external 0
2 if repo is sensitive 0
From .codeflow.yml 1
Additional review requirements
Show calculation
Max 0
0
From CODEOWNERS 0
Global minimum 0
Max 1
1
1 if commit is unverified 0
Sum 1


const f3 = (n) => (n === null || n === undefined ? "n/a" : Number(n).toFixed(3));
const usd = (n) => `$${Number(n || 0).toFixed(2)}`;
const cell = (s) => String(s ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").trim();

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants