Conversation
…he report workflow
listBotPullRequests uses the list-PRs endpoint, which never returns additions/deletions (confirmed live against base/docs and via the GitHub REST docs) — only the single-PR endpoint does. merge-rate.mjs was reading pr.additions/pr.deletions straight off list-endpoint objects, so humanRewriteRatios silently came back empty on every real run even though the offline fixture (which sets those fields by hand) made the unit test pass. Add fetchPRTotals and use it for merged bot PRs only.
…s, mirror sync env in replay
…ent at base, lazy-load sync code for CI
…tests run without the sdk
…per model, cap source diff in prompts
…ssert graders receive the payload diff
…r rubric; win/loss only when orders agree
…only source diffs; report variance token cost
…rules, grading errors excluded, no pairwise on changelog entries
…et guard, reflection
… the overall score
…ns for evals, no judge in first hillclimb runs
Collaborator
🟡 Heimdall Review Status
|
|
|
||
| const f3 = (n) => (n === null || n === undefined ? "n/a" : Number(n).toFixed(3)); | ||
| const usd = (n) => `$${Number(n || 0).toFixed(2)}`; | ||
| const cell = (s) => String(s ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").trim(); |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed? Why?
Adds a lightweight way to measure and improve the quality of the base-std docs sync bot (
.github/workflows/base-std-docs-sync.yml→scripts/sync-from-base-std/). Today nothing measures whether a reviewer would merge the bot's output: of 10 bot PRs, 1 merged (#1939, which humans rewrote ~85%), 1 closed, 8 still open. Reviewer comments show the same failures repeatedly: scope creep, paraphrasing the upstream changelog, wrong or ungrounded facts, housekeeping callouts, style slips.Method follows "Automating eval design and hillclimbing" (Claude blog): real cases first, cheapest grader that works, claim-based LLM judge calibrated against a human, train/held-out split, keep a change only when both improve beyond noise.
Everything is new, under
scripts/doc-evals/(plan:PLAN.md, usage:README.md):build-cases.mjs,cases/,replay/): 9 frozen cases built from past bot PRs (source sha, pre-sync docs commit, payload, human-merged reference where one exists, confirmed scope, verbatim review findings).replay/run.mjsreruns the sync in a throwaway git worktree against the historical docs tree. Nothing is pushed or opened.graders/,grade.mjs,calibrate.mjs): code checks (scope precision/recall and forbidden paths, identifier grounding, selector recomputation via keccak-256, lint-mdx, housekeeping callouts, changelog shape and upstream fidelity, no-op), a yes/no claim judge, a blinded two-order pairwise vs the human reference, judge variance, and a human calibration tool. Overall = 0.4 scope + 0.25 code + 0.2 judge + 0.15 pairwise.metrics/,.github/workflows/doc-quality-report.yml): bot PR merge rate, human rewrite ratio, superseded PRs, review-comment taxonomy, candidate cases. Weekly, non-blocking; updates one issue titled "Docs sync quality report". No PR trigger.hillclimb/): proposes one root-cause change tollm/prompts.mjsorroute-table.json, validates it against the sync's own tests in a scratch copy, replays train and test, and keeps it only if train improves beyond noise and held-out test improves. Test cases never reach the proposer. Budget cap, reflection after 2 non-keeps. Makes no commits and never touches the real sync code.Owner decisions baked in (PLAN.md "Decisions"): changelog entries follow the upstream entry closely; the bot never edits
docs/build-on-base/; evals run withCLAUDE_MAX_TOKENS=16000; failed judge calls are excluded from scores and block hillclimb decisions.Notes to reviewers
docs/, the sync code, or the existing workflows. The only non-new file is the rootpackage.jsontestglob, which now includesscripts/doc-evals/__tests__/*.test.mjs. Those tests are offline and never load@anthropic-ai/sdk, since CI runsnpm testwithoutscripts/node_modules.CLAUDE_MAX_TOKENSrepo variable to 16000.index.mjs, outside the hillclimb's edit surface.base-std-routing.test.mjs("upstream docs tree routes to the pages the IA guidelines assign") already fails on master.How has it been tested?
npm testat the repo root: 346 tests pass, including withscripts/node_modulesremoved (CI mode; 1 skip for an SDK-only equality check).grade.mjsscored that run end to end with the Opus judge; judge variance was 9/9 agreement per claim.Screenshots
N/A (no user-facing changes; no rendered docs pages change).
Generated with Toshi