Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 18 additions & 5 deletions examples/07-grader-file/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,9 +28,9 @@ The scenario references each grader by name:
"grader": { "file": "./grade.eval.ts", "name": "lunch-stays-free" }
```

**Cost:** free for the grader runs below. The `measure` run (8 haiku runs
in artifact mode) was not run for this README, so its cost is not measured.
**Requires:** bun; the claude CLI for `measure`
**Cost:** free for the grader runs below; ~$0.29 for `measure` (8 haiku
runs in artifact mode) · **Time:** ~97s for `measure` · **Requires:** bun;
the claude CLI for `measure`

## Try the graders without a model run

Expand Down Expand Up @@ -90,8 +90,21 @@ no artifact: schedule.json does not exist in the grader's working directory
./promptdiff measure --scenario ./examples/07-grader-file/scenario.json
```

This output was not captured for this README. Each scenario prints a line
like `3/3 pass (100%), 1 no-artifact`.
```
promptdiff measure: day-planner (haiku via claude-p)

no-overlaps
4/4 pass (100%) | $0.1570

lunch-stays-free
4/4 pass (100%) | $0.1351

total cost: $0.2921
```

Haiku wrote `schedule.json` in all 8 runs, so no `no-artifact` count
appears. When a run writes nothing, the line reads
`3/3 pass (100%), 1 no-artifact` and that run is listed below it.

## What to notice

Expand Down
6 changes: 3 additions & 3 deletions examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@ you can see what promptdiff gives you before spending anything yourself.
The output in 01-04 and 06 is captured from a real run. 05 is the exception:
it needs a local ollama server, which was not available where these examples
were built, so its README shows an illustrative block and says so. 07's
grader output is real and free (it grades sample files with no model run);
its `measure` output was not captured.
grader output is free (it grades sample files with no model run), and its
`measure` output is from a real run.

Run them from the repo root. All except 05 need the `claude` CLI installed;
none need an API key beyond what `claude` already uses.
Expand All @@ -22,7 +22,7 @@ none need an API key beyond what `claude` already uses.
| [04-judge-grader](./04-judge-grader/) | A calibrated LLM judge end to end: rubric, labeled fixtures, `calibrate`, then a compare — and the refuse-to-grade gate when calibration is missing. | ~$0.10 |
| [05-any-model](./05-any-model/) | The `openai` runner against local ollama: not Claude-only, zero API keys. Same shape works for vLLM, llama.cpp, OpenRouter. Output illustrative, not a real run. | free (local) |
| [06-token-economics](./06-token-economics/) | Turn caps (`--max-turns`) + raw per-run results (`--raw-out`): token-category analysis, and how a capped run is failed without grading. | ~$0.19 |
| [07-grader-file](./07-grader-file/) | A grader file: named graders with `result.assert` for semantic checks, every failure reported with its context, and `no-artifact` (exit 77) kept apart from failures. Try the graders on sample files before any paid run. | free (graders); `measure` not measured |
| [07-grader-file](./07-grader-file/) | A grader file: named graders with `result.assert` for semantic checks, every failure reported with its context, and `no-artifact` (exit 77) kept apart from failures. Try the graders on sample files before any paid run. | free (graders); ~$0.29 (`measure`) |

Every `scenario.json` here is loaded by `test/examples.test.ts` in CI, and
every `*.eval.ts` grader file is typechecked, so the examples can't silently
Expand Down
Loading