From 4530feffbf21862345f7791ba84d22265acbbc0e Mon Sep 17 00:00:00 2001 From: Josh Owens Date: Wed, 30 Sep 2026 20:21:58 +0000 Subject: [PATCH] docs: capture example 07 measure output and cost Co-Authored-By: Claude Opus 5.5 --- examples/07-grader-file/README.md | 23 ++++++++++++++++++----- examples/README.md | 6 +++--- 2 files changed, 21 insertions(+), 8 deletions(-) diff --git a/examples/07-grader-file/README.md b/examples/07-grader-file/README.md index e09582a..b4af66e 100644 --- a/examples/07-grader-file/README.md +++ b/examples/07-grader-file/README.md @@ -28,9 +28,9 @@ The scenario references each grader by name: "grader": { "file": "./grade.eval.ts", "name": "lunch-stays-free" } ``` -**Cost:** free for the grader runs below. The `measure` run (8 haiku runs -in artifact mode) was not run for this README, so its cost is not measured. -**Requires:** bun; the claude CLI for `measure` +**Cost:** free for the grader runs below; ~$0.29 for `measure` (8 haiku +runs in artifact mode) · **Time:** ~97s for `measure` · **Requires:** bun; +the claude CLI for `measure` ## Try the graders without a model run @@ -90,8 +90,21 @@ no artifact: schedule.json does not exist in the grader's working directory ./promptdiff measure --scenario ./examples/07-grader-file/scenario.json ``` -This output was not captured for this README. Each scenario prints a line -like `3/3 pass (100%), 1 no-artifact`. +``` +promptdiff measure: day-planner (haiku via claude-p) + +no-overlaps + 4/4 pass (100%) | $0.1570 + +lunch-stays-free + 4/4 pass (100%) | $0.1351 + +total cost: $0.2921 +``` + +Haiku wrote `schedule.json` in all 8 runs, so no `no-artifact` count +appears. When a run writes nothing, the line reads +`3/3 pass (100%), 1 no-artifact` and that run is listed below it. ## What to notice diff --git a/examples/README.md b/examples/README.md index cf06397..9d67a2b 100644 --- a/examples/README.md +++ b/examples/README.md @@ -8,8 +8,8 @@ you can see what promptdiff gives you before spending anything yourself. The output in 01-04 and 06 is captured from a real run. 05 is the exception: it needs a local ollama server, which was not available where these examples were built, so its README shows an illustrative block and says so. 07's -grader output is real and free (it grades sample files with no model run); -its `measure` output was not captured. +grader output is free (it grades sample files with no model run), and its +`measure` output is from a real run. Run them from the repo root. All except 05 need the `claude` CLI installed; none need an API key beyond what `claude` already uses. @@ -22,7 +22,7 @@ none need an API key beyond what `claude` already uses. | [04-judge-grader](./04-judge-grader/) | A calibrated LLM judge end to end: rubric, labeled fixtures, `calibrate`, then a compare — and the refuse-to-grade gate when calibration is missing. | ~$0.10 | | [05-any-model](./05-any-model/) | The `openai` runner against local ollama: not Claude-only, zero API keys. Same shape works for vLLM, llama.cpp, OpenRouter. Output illustrative, not a real run. | free (local) | | [06-token-economics](./06-token-economics/) | Turn caps (`--max-turns`) + raw per-run results (`--raw-out`): token-category analysis, and how a capped run is failed without grading. | ~$0.19 | -| [07-grader-file](./07-grader-file/) | A grader file: named graders with `result.assert` for semantic checks, every failure reported with its context, and `no-artifact` (exit 77) kept apart from failures. Try the graders on sample files before any paid run. | free (graders); `measure` not measured | +| [07-grader-file](./07-grader-file/) | A grader file: named graders with `result.assert` for semantic checks, every failure reported with its context, and `no-artifact` (exit 77) kept apart from failures. Try the graders on sample files before any paid run. | free (graders); ~$0.29 (`measure`) | Every `scenario.json` here is loaded by `test/examples.test.ts` in CI, and every `*.eval.ts` grader file is typechecked, so the examples can't silently