Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .claude/skills/add-lang/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,8 +163,8 @@ Tiers (match `corpus.json`): **Small** <~150 files · **Medium** ~150–1500 ·
**Large** >~1500. Skip repos that are tagged `<lang>` but mostly another
language. Write one cross-file architecture **question** per repo (the kind that
needs tracing across files). Add a `"<Language>"` block to
`.claude/skills/agent-eval/corpus.json` (fields: `name`, `repo`, `size`,
`files`, `question`) so `/agent-eval` can reuse them.
`.claude/skills/codegraph-lift/corpus.json` (fields: `name`, `repo`, `size`,
`files`, `question`) so `/codegraph-lift` can reuse them.

### Step 8 — Benchmark all 3 (extraction + A/B)

Expand Down Expand Up @@ -210,7 +210,7 @@ releases go through the GitHub Actions Release workflow.
## Notes
- The A/B spawns real **paid** `claude -p` runs (opus, `--max-budget-usd`),
2 arms × 3 repos. The corpus dir `/tmp/codegraph-corpus` is shared with
`/agent-eval`, so clones are reused across runs.
`/codegraph-lift`, so clones are reused across runs.
- Any new `*.wasm` must live in `src/extraction/wasm/` — `copy-assets` (run by
`npm run build`) ships it; otherwise it won't be in `dist/`.
- An index must be served by the **same** binary that built it. Step 8 builds +
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: agent-eval
description: Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo.
name: codegraph-lift
description: Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /codegraph-lift or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo.
---

# CodeGraph Quality Audit
Expand All @@ -10,9 +10,20 @@ codegraph version on a chosen real-world repo. Drives the harness in
`scripts/agent-eval/`.

## Prerequisites
- `tmux` 3+, a logged-in `claude` CLI, `node`, `git` (macOS/Linux).
- `node`, `git`, a logged-in agent CLI (Claude Code today — see Runners).
- `tmux` 3+ for the interactive harness only.
- Run from the codegraph repo root.

## Runners

Headless is the portable arm; the tmux arm drives the Claude TUI specifically.

- Claude Code: `claude -p` with stream-json — native; `parse-run.mjs` / `parse-session.mjs` are built for its formats.
- opencode: `opencode run --format json` — proven in other panels; needs its own stream parser (follow-up).
- Cursor: `cursor-agent -p` — proven; avoid `--mode plan` (swallows print output), pass `--trust` headless.
- Devin: `devin -p` is help-asserted but unverified here.
- `AskUserQuestion` below means the host's question tool (name varies by host).

## Workflow

Copy this checklist:
Expand All @@ -27,12 +38,12 @@ Copy this checklist:

**Step 1 — version.** Ask with `AskUserQuestion`: which codegraph version to test.
Offer "Local dev build" and "Latest published"; the free-text "Other" lets the
user type a specific version (e.g. `0.7.10`). Map the answer to a VERSION token:
user type a specific version. Map the answer to a VERSION token:
- "Local dev build" → `local`
- "Latest published" → `latest`
- a typed version → that string (e.g. `0.7.10`)
- a typed version → that string

**Step 2 — language.** Read `.claude/skills/agent-eval/corpus.json`. Ask with
**Step 2 — language.** Read `.claude/skills/codegraph-lift/corpus.json`. Ask with
`AskUserQuestion` which language to test, listing the languages that have entries.

**Step 3 — repo.** From the chosen language's entries, ask which repo. Label each
Expand All @@ -48,7 +59,7 @@ the answer to a MODE token:
- "Both" → `all` — headless + interactive (4 runs).

**Step 5 — run.** Launch in the background (sets the version, clones if missing,
wipes + re-indexes, runs the chosen arms — several minutes):
wipes + re-indexes, runs the chosen arms — several minutes, paid runs):
```bash
scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>
```
Expand Down
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"_comment": "Test corpus for /agent-eval. Add entries freely. size: Small (<~150 files), Medium (~150-1500), Large (>~1500). 'question' is a representative architectural question that exercises cross-file understanding.",
"_comment": "Test corpus for /codegraph-lift. Add entries freely. size: Small (<~150 files), Medium (~150-1500), Large (>~1500). 'question' is a representative architectural question that exercises cross-file understanding.",
"TypeScript": [
{
"name": "ky",
Expand Down
2 changes: 1 addition & 1 deletion docs/design/dynamic-dispatch-coverage-playbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -141,7 +141,7 @@ from adoption.

### Step 1 — Pick the framework's canonical *flow* question
Every framework has a signature data/control flow. Pick the "how does X reach/become Y"
question and a real repo (add to `.claude/skills/agent-eval/corpus.json`). Examples:
question and a real repo (add to `.claude/skills/codegraph-lift/corpus.json`). Examples:
- React state→DOM, Vue reactive→render, Svelte store→update
- Rails request→controller→view, Spring request→`@Controller`→service
- Express/Koa request→middleware→handler, FastAPI request→route→dependency
Expand Down