Skip to content

feat: grader file format with test-style assertions - #45

Merged
queso merged 6 commits into
mainfrom
feat/grader-file-format
Sep 30, 2026
Merged

queso merged 6 commits into
mainfrom
feat/grader-file-format

Conversation

@queso

@queso queso commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Closes #32.

What this adds

Grader files. A scenario can reference a named grader in a TS/JS file:

"grader": { "file": "./grade.eval.ts", "name": "product-named" }

The file's default export is grade(artifactPath, { name: ({ result }) => ... }). result has assert(condition, message, context?), json(), text, path, shouldHave, shouldNotHave, and shouldMatch. Every check runs, so one run reports every failure, with the context data or an artifact excerpt next to each. An exception in the grader counts as one more failure. A grader that makes no checks fails. The docs and the new example (examples/07-grader-file) lead with semantic result.assert checks, per design constraint 1.

The command contract stays the base layer. A file grader compiles to a command grader: bun <pkg>/promptdiff grade --file <abs> --name <name> runs in the sandbox with the same cwd, env, and timeout handling. promptdiff grade is also a public subcommand, so you can run a grader by hand inside a kept sandbox. The artifact path resolves against the grader's cwd, the same place a command grader looks. Grader files import grade from @theaiteam/promptdiff (or bare promptdiff). A Bun plugin resolves both names to the running copy, so no install is needed next to the file. package.json now has an exports entry for consumers.

Missing artifacts are a separate outcome.

  • File graders exit 77 when the artifact does not exist. Command graders can exit 77 too; the code is exported to them as $PROMPTDIFF_NO_ARTIFACT_EXIT_CODE. A grader with an explicit expectExitCode: 77 keeps its old meaning.
  • Pass rates and Fisher's p use graded runs only. Summaries read 4/4 pass (100%), 1 no-artifact and list each such run as run 3 no artifact: ....
  • compare prints a NOTE with both arms' no-artifact counts. An arm with zero graded runs fails its target or regression assertion.
  • Reports and receipts include noArtifact only when it is non-zero, so existing output does not change.

Load-time validation. A missing grader file, a bad default export, or an unknown grader name fails when the scenario loads, before any paid run. The check runs promptdiff grade --list in a child process.

compare points at measure. On a baseline-only scenario the error now reads compare requires proposed skill paths; to run one instruction set without a proposed arm, use \promptdiff measure``.

Cache. The baseline cache key includes a hash of the grader file's content.

Open question for review

In compare, a proposed arm that produces the artifact less often than baseline only gets a NOTE. For example, proposed at 1/1 pass plus 4 no-artifact, against baseline at 3/5, passes a target assertion. That follows the rule that a missing artifact is never a failure. For a regression assertion, though, a skill change that stops the agent writing the artifact is arguably the regression you want to catch. One option is to fail regression when proposed's no-artifact count clearly exceeds baseline's. I left that out rather than guess at the threshold.

Known limits

  • The cache hashes the grader file but not modules it imports (documented).
  • Loading a scenario now executes grader file code in a child process (documented in the security note).
  • The promptdiff grade output in example 07 was captured for real. Its measure cost is marked "not measured" because no paid run was made.

Testing

  • bun run check: 155 tests pass, typecheck clean. The new tests in test/grade-file.test.ts cover collecting every failure, no-artifact classification for file and command graders, summary lines, config validation errors, and the compare hint.
  • Ran promptdiff grade by hand from a directory outside the repo with a bare import { grade } from "promptdiff". Checked --list, a missing artifact (exit 77), two failing checks reported together (exit 1), all checks passing (exit 0), and an unknown name (exit 2).

🤖 Generated with Claude Code

queso and others added 4 commits September 30, 2026 18:37
A grader file is a TS/JS file whose default export is
grade(artifactPath, { name: ({ result }) => { ... } }). A scenario picks
one grader with "grader": { "file": "./grade.eval.ts", "name": "..." }.
result offers assert(condition, message, context?), json(), text, path,
shouldHave, shouldNotHave, and shouldMatch. Every check is collected
instead of stopping at the first failure, and each failure prints its
message with its context or with what the artifact held instead.

A file grader compiles to a command grader: the engine runs
`promptdiff grade --file <file> --name <name>` in the sandbox, so the
artifact path resolves against the sandbox like any command grader's
paths. The new `grade` subcommand exits 0 (pass), 1 (failed checks),
77 (no artifact), or 2 (unusable file or name). Scenario loading checks
the file and name in a child process before any paid run, and the
baseline cache key includes the grader file's content.

No artifact is now its own outcome. Command graders signal it by exiting
77 (also exported as $PROMPTDIFF_NO_ARTIFACT_EXIT_CODE) unless their
expectExitCode is 77. Such runs are neither passes nor failures: pass
rates cover graded runs only, summaries print "4/4 pass (100%), 1
no-artifact", compare notes each arm's count and fails an assertion for
an arm with no graded run, and reports and receipts carry the count.

The package now exports grade and its types from src/index.ts. Under
`promptdiff grade`, "@theaiteam/promptdiff" and "promptdiff" resolve to
the running copy, so grader files work without a local install.

Refs #32

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
compare on a baseline-only scenario failed with "compare requires
proposed skill paths". The error now says to use `promptdiff measure`,
which runs one instruction set without a proposed arm.

Refs #32

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Covers check collection (every failure reported, thrown errors and
invalid JSON recorded), exit-code mapping, artifact paths resolved
against the grader cwd, command graders exiting 77, scenario validation
of the grader-file shape (missing file, unknown name, bad export),
measure and compare summary lines with no-artifact runs, the grader file
in the cache key, report records, and the compare error that points at
measure.

Refs #32

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
README and SPEC describe the grader file format, leading with
result.assert for semantic checks, how a file grader runs as a command
grader, the exit-77 no-artifact convention for command graders, and how
measure and compare count no-artifact runs. examples/07-grader-file adds
two semantic graders with sample artifacts that run without a model
call.

Refs #32

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nitpick review — comment

The grader-file feature is well-structured and largely well-tested, but the exit-77 reclassification applies to all command graders, not just file graders, so a broken grader script exiting 77 would now be excluded from pass rates instead of counted as a failure — a silent semantic change for existing suites. A few lower-severity items remain: prototype-inherited lookups in grader-file, a potentially degenerate samplingP when an arm has zero graded runs, a test depending on a prebuilt binary, and new timeout/no-artifact semantics that ship without pinning tests. No blocking issues; recommend addressing the exit-77 scoping before merge.

6 inline comment(s).

Comment thread src/engine/grader.ts
Comment thread src/engine/grader-file.ts Outdated
Comment thread src/engine/compare.ts Outdated
Comment thread test/grade-file.test.ts
Comment thread test/grade-file.test.ts
Comment thread src/engine/compare.ts
queso and others added 2 commits September 30, 2026 19:47
- Fix: look up grader names by own property, so "constructor" or
  "toString" exit 2 instead of running an inherited function
- Fix: leave samplingP undefined when an arm has no graded runs
- Fix: kill the grader's whole process group on timeout; a hung file
  grader (sh -> bun) kept the pipes open past timeoutMs
- Test: inherited names, samplingP with graded denominators, file
  grader timeout

Addresses review comments from github-actions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@queso
queso merged commit 52236f9 into main Sep 30, 2026
2 checks passed

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nitpick review — comment

The grader-file feature is well-structured and mostly well-tested, but the exit-77 reclassification in gradeCommand changes semantics for any pre-existing command grader that happens to exit 77 without declaring it: such graders silently move from 'visible failure' to 'noArtifact' (excluded from pass-rate denominators), altering compare statistics with no config change. Two smaller items remain: an unguarded JSON.parse that degrades a load-time diagnostic into an opaque crash, and a CLI help test that depends on a prebuilt binary at the repo root. Worth addressing before merge, none blocking.

1 inline comment(s).

2 previously-acknowledged finding(s) not re-posted (resolved review threads).

Comment thread src/engine/grader-file.ts
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: grader file format with test-style assertions (result.shouldHave / result.assert)

1 participant