From 2e4a03ed1624b1f47621c23518b6ab8abdc6a8bd Mon Sep 17 00:00:00 2001 From: Aaron Queen Date: Thu, 17 Sep 2026 15:53:27 -0600 Subject: [PATCH 1/3] skills: rename agent-eval to codegraph-lift --- .claude/skills/add-lang/SKILL.md | 6 +++--- .claude/skills/{agent-eval => codegraph-lift}/SKILL.md | 6 +++--- .claude/skills/{agent-eval => codegraph-lift}/corpus.json | 2 +- docs/design/dynamic-dispatch-coverage-playbook.md | 2 +- 4 files changed, 8 insertions(+), 8 deletions(-) rename .claude/skills/{agent-eval => codegraph-lift}/SKILL.md (92%) rename .claude/skills/{agent-eval => codegraph-lift}/corpus.json (98%) diff --git a/.claude/skills/add-lang/SKILL.md b/.claude/skills/add-lang/SKILL.md index 37cbdce55..c8328e68c 100644 --- a/.claude/skills/add-lang/SKILL.md +++ b/.claude/skills/add-lang/SKILL.md @@ -163,8 +163,8 @@ Tiers (match `corpus.json`): **Small** <~150 files · **Medium** ~150–1500 · **Large** >~1500. Skip repos that are tagged `` but mostly another language. Write one cross-file architecture **question** per repo (the kind that needs tracing across files). Add a `""` block to -`.claude/skills/agent-eval/corpus.json` (fields: `name`, `repo`, `size`, -`files`, `question`) so `/agent-eval` can reuse them. +`.claude/skills/codegraph-lift/corpus.json` (fields: `name`, `repo`, `size`, +`files`, `question`) so `/codegraph-lift` can reuse them. ### Step 8 — Benchmark all 3 (extraction + A/B) @@ -210,7 +210,7 @@ releases go through the GitHub Actions Release workflow. ## Notes - The A/B spawns real **paid** `claude -p` runs (opus, `--max-budget-usd`), 2 arms × 3 repos. The corpus dir `/tmp/codegraph-corpus` is shared with - `/agent-eval`, so clones are reused across runs. + `/codegraph-lift`, so clones are reused across runs. - Any new `*.wasm` must live in `src/extraction/wasm/` — `copy-assets` (run by `npm run build`) ships it; otherwise it won't be in `dist/`. - An index must be served by the **same** binary that built it. Step 8 builds + diff --git a/.claude/skills/agent-eval/SKILL.md b/.claude/skills/codegraph-lift/SKILL.md similarity index 92% rename from .claude/skills/agent-eval/SKILL.md rename to .claude/skills/codegraph-lift/SKILL.md index 8d06ac733..4d320105a 100644 --- a/.claude/skills/agent-eval/SKILL.md +++ b/.claude/skills/codegraph-lift/SKILL.md @@ -1,6 +1,6 @@ --- -name: agent-eval -description: Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo. +name: codegraph-lift +description: Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /codegraph-lift or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo. --- # CodeGraph Quality Audit @@ -32,7 +32,7 @@ user type a specific version (e.g. `0.7.10`). Map the answer to a VERSION token: - "Latest published" → `latest` - a typed version → that string (e.g. `0.7.10`) -**Step 2 — language.** Read `.claude/skills/agent-eval/corpus.json`. Ask with +**Step 2 — language.** Read `.claude/skills/codegraph-lift/corpus.json`. Ask with `AskUserQuestion` which language to test, listing the languages that have entries. **Step 3 — repo.** From the chosen language's entries, ask which repo. Label each diff --git a/.claude/skills/agent-eval/corpus.json b/.claude/skills/codegraph-lift/corpus.json similarity index 98% rename from .claude/skills/agent-eval/corpus.json rename to .claude/skills/codegraph-lift/corpus.json index 150b4a601..6cece83fe 100644 --- a/.claude/skills/agent-eval/corpus.json +++ b/.claude/skills/codegraph-lift/corpus.json @@ -1,5 +1,5 @@ { - "_comment": "Test corpus for /agent-eval. Add entries freely. size: Small (<~150 files), Medium (~150-1500), Large (>~1500). 'question' is a representative architectural question that exercises cross-file understanding.", + "_comment": "Test corpus for /codegraph-lift. Add entries freely. size: Small (<~150 files), Medium (~150-1500), Large (>~1500). 'question' is a representative architectural question that exercises cross-file understanding.", "TypeScript": [ { "name": "ky", diff --git a/docs/design/dynamic-dispatch-coverage-playbook.md b/docs/design/dynamic-dispatch-coverage-playbook.md index 8ec0a8e38..a5ec16e1d 100644 --- a/docs/design/dynamic-dispatch-coverage-playbook.md +++ b/docs/design/dynamic-dispatch-coverage-playbook.md @@ -141,7 +141,7 @@ from adoption. ### Step 1 — Pick the framework's canonical *flow* question Every framework has a signature data/control flow. Pick the "how does X reach/become Y" -question and a real repo (add to `.claude/skills/agent-eval/corpus.json`). Examples: +question and a real repo (add to `.claude/skills/codegraph-lift/corpus.json`). Examples: - React state→DOM, Vue reactive→render, Svelte store→update - Rails request→controller→view, Spring request→`@Controller`→service - Express/Koa request→middleware→handler, FastAPI request→route→dependency From c62c413ef63b32cd5fa3137e9d5e4053b3795a9a Mon Sep 17 00:00:00 2001 From: Aaron Queen Date: Thu, 17 Sep 2026 15:57:40 -0600 Subject: [PATCH 2/3] skills: drop stale version example; tmux only for interactive harness --- .claude/skills/codegraph-lift/SKILL.md | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/.claude/skills/codegraph-lift/SKILL.md b/.claude/skills/codegraph-lift/SKILL.md index 4d320105a..c4c0631d4 100644 --- a/.claude/skills/codegraph-lift/SKILL.md +++ b/.claude/skills/codegraph-lift/SKILL.md @@ -10,7 +10,8 @@ codegraph version on a chosen real-world repo. Drives the harness in `scripts/agent-eval/`. ## Prerequisites -- `tmux` 3+, a logged-in `claude` CLI, `node`, `git` (macOS/Linux). +- `node`, `git`, a logged-in `claude` CLI (macOS/Linux). +- `tmux` 3+ for the interactive harness only. - Run from the codegraph repo root. ## Workflow @@ -27,10 +28,10 @@ Copy this checklist: **Step 1 — version.** Ask with `AskUserQuestion`: which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the -user type a specific version (e.g. `0.7.10`). Map the answer to a VERSION token: +user type a specific version. Map the answer to a VERSION token: - "Local dev build" → `local` - "Latest published" → `latest` -- a typed version → that string (e.g. `0.7.10`) +- a typed version → that string **Step 2 — language.** Read `.claude/skills/codegraph-lift/corpus.json`. Ask with `AskUserQuestion` which language to test, listing the languages that have entries. From c0e90524122b5cc9278ab9b8cb59e95dc155f017 Mon Sep 17 00:00:00 2001 From: Aaron Queen Date: Thu, 17 Sep 2026 15:59:55 -0600 Subject: [PATCH 3/3] skills: codegraph-lift runners table; paid-run disclosure --- .claude/skills/codegraph-lift/SKILL.md | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/.claude/skills/codegraph-lift/SKILL.md b/.claude/skills/codegraph-lift/SKILL.md index c4c0631d4..f55d33f76 100644 --- a/.claude/skills/codegraph-lift/SKILL.md +++ b/.claude/skills/codegraph-lift/SKILL.md @@ -10,10 +10,20 @@ codegraph version on a chosen real-world repo. Drives the harness in `scripts/agent-eval/`. ## Prerequisites -- `node`, `git`, a logged-in `claude` CLI (macOS/Linux). +- `node`, `git`, a logged-in agent CLI (Claude Code today — see Runners). - `tmux` 3+ for the interactive harness only. - Run from the codegraph repo root. +## Runners + +Headless is the portable arm; the tmux arm drives the Claude TUI specifically. + +- Claude Code: `claude -p` with stream-json — native; `parse-run.mjs` / `parse-session.mjs` are built for its formats. +- opencode: `opencode run --format json` — proven in other panels; needs its own stream parser (follow-up). +- Cursor: `cursor-agent -p` — proven; avoid `--mode plan` (swallows print output), pass `--trust` headless. +- Devin: `devin -p` is help-asserted but unverified here. +- `AskUserQuestion` below means the host's question tool (name varies by host). + ## Workflow Copy this checklist: @@ -49,7 +59,7 @@ the answer to a MODE token: - "Both" → `all` — headless + interactive (4 runs). **Step 5 — run.** Launch in the background (sets the version, clones if missing, -wipes + re-indexes, runs the chosen arms — several minutes): +wipes + re-indexes, runs the chosen arms — several minutes, paid runs): ```bash scripts/agent-eval/audit.sh "" ```