Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion engine/hooks/llm-judge/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,10 @@ the dictionary's `on_hit` text.

`ask(prompt)` tries these in order and stops at the first one that answers:

1. **codex**: `codex exec --skip-git-repo-check -m gpt-5.3-codex-spark --sandbox read-only -c notify=[] PROMPT`
1. **codex**: `codex exec --skip-git-repo-check --sandbox read-only -c notify=[] PROMPT`
(no `-m`: codex runs the model set in `~/.codex/config.toml`, so the judge
uses a model the account can already call; a ChatGPT-account login refuses
API-only models)
2. **claude**: `claude -p --model haiku --settings '{"disableAllHooks": true}' PROMPT`
3. **cursor**: `cursor-agent -p --output-format text PROMPT`

Expand Down
2 changes: 1 addition & 1 deletion engine/hooks/llm-judge/judge.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@
RUNNERS_ENV = "CATSTACK_LLM_JUDGE_RUNNERS"
STATE_ENV = "CATSTACK_LLM_JUDGE_STATE_DIR"
DEFAULT_RUNNERS = (
("codex", ["codex", "exec", "--skip-git-repo-check", "-m", "gpt-5.3-codex-spark", "--sandbox", "read-only", "-c", "notify=[]", PROMPT_SLOT]),
("codex", ["codex", "exec", "--skip-git-repo-check", "--sandbox", "read-only", "-c", "notify=[]", PROMPT_SLOT]),
("claude", ["claude", "-p", "--model", "haiku", "--settings", '{"disableAllHooks": true}', PROMPT_SLOT]),
("cursor", ["cursor-agent", "-p", "--output-format", "text", PROMPT_SLOT]),
)
Expand Down
7 changes: 7 additions & 0 deletions engine/hooks/llm-judge/tests/test_judge.py
Original file line number Diff line number Diff line change
Expand Up @@ -179,6 +179,13 @@ def test_default_runner_order_is_codex_then_claude_then_cursor(self):
os.environ.pop(judge.RUNNERS_ENV)
self.assertEqual([name for name, _ in judge.runners()], ["codex", "claude", "cursor"])

def test_default_codex_runner_uses_the_account_model_not_a_pinned_one(self):
with patch.dict(os.environ):
os.environ.pop(judge.RUNNERS_ENV)
codex_argv = dict(judge.runners())["codex"]
self.assertNotIn("-m", codex_argv)
self.assertNotIn("--model", codex_argv)

def test_investigate_runner_argv_is_read_only_and_excludes_cursor(self):
os.environ.pop(judge.RUNNERS_ENV)
self.assertEqual(
Expand Down
2 changes: 1 addition & 1 deletion engine/hooks/wrong-check-reflect/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ window left the `/reflect` typed beside it outside. Inside a judge run (`CATSTAC
refuses the job.

The model call runs in a detached background process, so the reply is never
held up. Runners are tried in `llm-judge` order: `codex` (gpt-5.3-codex-spark),
held up. Runners are tried in `llm-judge` order: `codex` (the account's configured model),
then `claude` (haiku, hooks off), then `cursor-agent`, first answer wins.

The verdict reports one turn later. On the next prompt the `llm-judge` inbox
Expand Down
Loading