Skip to content

skill-usage-log: record every skill use in Claude, Cursor and Codex, on by default - #826

Merged
mergify[bot] merged 2 commits into
mainfrom
hooks/skill-usage-metrics-all-harnesses
Sep 24, 2026
Merged

mergify[bot] merged 2 commits into
mainfrom
hooks/skill-usage-metrics-all-harnesses

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

A skill is a saved set of steps an agent can load and follow. Until now nobody could tell which skills the agents really use.

Each agent now writes one line every time it uses a skill, into the same shared file the other usage counts already read.

All three agents do this, not just one, and it is on unless you switch it off.

A new option on the usage summary tool groups those lines by skill, and names the skills no agent has ever loaded.

Review Claim

Each of the four ways an agent can load a skill writes exactly one line, naming the skill, which agent, and how it was loaded. Talking about a skill without loading it writes nothing.

Review Lane

behavior

Review Unit

engine-runtime

Safety Invariant

The hook never blocks or changes a tool call or prompt. Input it cannot read writes skill_usage_unchecked and a catstack-hook-error line, so it is never counted as "no skill used".

Slice Rationale

The detector, the three harness entry points, their install wiring and the report that reads the rows are one feature; a harness wired without the others would report zero for skills it simply cannot see.

Non-goals

Does not update the root README's env-var row (a separate docs slice stacked on this one). Does not count a skill whose SKILL.md a subagent is only told to read; that subagent's own read is counted.

Test Plan

Test Plan
python3 engine/hooks/skill-usage-log/tests/test_hooks.py      Ran 13 tests  OK
python3 engine/hooks/_runner/tests/test_report.py             Ran 12 tests  OK
python3 -m unittest tests.test_install.TestSkillSymlinks.test_skill_usage_log_wired_for_claude_cursor_and_codex   OK

Every CI step, run locally at this commit:

bash scripts/test/run_all_tests.sh        Ran 644 tests  FAILED (failures=1: the hook-coverage gate, fixed in this commit)
python3 scripts/ci/check_hook_test_coverage.py   check_hook_test_coverage: OK (41 hook(s) checked)
python3 tests/test_check_hook_test_coverage.py   OK
uvx ruff check . --select E9,F                   All checks passed!
shellcheck install.sh                            exit 0
every other scripts/ci/check_*.py step in ci.yml exit 0

test_git_path_churn.py (2 failures) is in a file this diff does not touch.

Backtest of the detector on recent real transcripts (tool name and input as recorded):

claude counted {'shell_read': 6, 'skill_tool': 4}
cursor counted {'Read': 670, 'Shell': 16, 'ReadFile': 22}
codex  counted {'shell_read': 1}
cursor rejected {'Task': 25, 'Glob': 80, 'Write': 42, 'StrReplace': 44, 'Shell': 140, 'CreatePlan': 8, 'Read': 1, 'Grep': 13, 'TodoWrite': 1, 'WebFetch': 7}

Real path, Claude: claude -p --settings <this branch's hook through the runner>, asked to load diu with the Skill tool and cat the cat-mode SKILL.md:

{'hook': 'skill-usage-log', 'harness': 'claude', 'action': 'skill_used', 'reason': 'skill_tool', 'skill': 'diu'}
{'hook': 'skill-usage-log', 'harness': 'claude', 'action': 'skill_used', 'reason': 'shell_read', 'skill': 'cat-mode'}
{'hook': 'skill-usage-log', 'harness': 'claude', 'action': 'skill_used', 'reason': 'read', 'skill': 'cat-mode'}

Not exercised on a live harness: Codex and Cursor. Codex runs only hooks whose hash it has trusted, and cursor-agent is not logged in on this machine. Their payload shapes are covered by the backtest above and by entry-script tests that run each script as a real process.

Revert Plan

Revert Plan

Revert this commit, then delete every hook entry whose command contains skill-usage-log/ from ~/.cursor/hooks.json, ~/.codex/hooks.json and the UserPromptSubmit list in ~/.claude/settings.json, and rerun install.sh. The old installer only replaces the Claude PreToolUse entry, so the other entries would keep calling scripts the revert deletes. Rows already written stay in the events file; nothing else reads skill_used.

🤖 Generated with Claude Code


Note

Low Risk
Metrics-only hook that never blocks agent actions; main risk is extra PreToolUse/prompt hook runs and local event file growth, not security or data handling.

Overview
Replaces the old Claude-only, opt-in skill-usage-log (private JSONL under ~/.cache/catstack-skill-usage-log) with silent metrics on Claude, Cursor, and Codex that append skill_used rows to the shared hook metrics event log.

Detection covers Skill tool calls, direct/shell reads of SKILL.md, leading /skill prompts, and Codex $skill mentions, while ignoring edits, searches, and incidental path mentions. Shared log.py + detect.py back thin harness entry scripts; write_stage_event now accepts extra fields (e.g. skill name). Unreadable input or skill dirs emit skill_usage_unchecked and surface via report.py --skills (exit 2).

Install merges hook fragments for all three agents via install_common.py; hooks.toml drops the old CATSTACK_SKILL_USAGE_LOG enable gate (opt out with =0). Docs and tests cover reporting, detection, and install wiring.

Reviewed by Cursor Bugbot for commit ed7e548. Bugbot is set up for automated code reviews on this repo. Configure here.

EdbertChan pushed a commit that referenced this pull request Sep 24, 2026
@EdbertChan
EdbertChan force-pushed the hooks/skill-usage-metrics-all-harnesses branch from 5cc2331 to fe27181 Compare September 24, 2026 03:58
@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_529425e3-ef8b-4b83-8f3f-97f6b703fb54)

@EdbertChan
EdbertChan force-pushed the hooks/codex-notify-through-metrics-runner branch 2 times, most recently from 1345e62 to 489189c Compare September 24, 2026 04:14
EdbertChan pushed a commit that referenced this pull request Sep 24, 2026
@EdbertChan
EdbertChan force-pushed the hooks/skill-usage-metrics-all-harnesses branch from fe27181 to f90ccbd Compare September 24, 2026 04:14
@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_0b887549-cd92-49a8-b45f-cb9093f1a430)

@EdbertChan
EdbertChan changed the base branch from hooks/codex-notify-through-metrics-runner to main September 24, 2026 05:25
@mergify

mergify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Queued — the merge queue status continues in this comment ↓.

@EdbertChan

Copy link
Copy Markdown
Owner Author

Mergify repair stopped: GitHub reports merge conflict. The retry cap was reached for current head f90ccbd.

EdbertChan and others added 2 commits September 24, 2026 15:52
…on by default

skill-usage-log was switched off unless CATSTACK_SKILL_USAGE_LOG=1, covered
only Claude's Skill tool, swallowed write errors, and wrote a private file no
report read. ~/.cache/catstack-skill-usage-log did not exist on this machine,
so no skill use was on record.

It now records one catstack.hook_event.v1 row per use (action skill_used,
reason = source, skill = name) in the shared events file, from all three
harnesses:

- Claude PreToolUse (Skill|Read|Bash) and UserPromptSubmit
- Cursor preToolUse and beforeSubmitPrompt
- Codex PreToolUse and UserPromptSubmit

Sources: skill_tool, read (a Read of a SKILL.md, including Cursor's
skills-cursor/), shell_read (cat/sed/head/tail/nl/less/more/bat on one, also
inside a Codex exec code string), slash (a prompt starting /<installed
skill>), mention (Codex $<installed skill>).

Backtest on recent real transcripts: Claude 10 uses counted, Cursor 708
(670 of 671 SKILL.md Read calls, 22 of 22 ReadFile), Codex 1; rejected
mentions were Glob/Grep patterns, Write/StrReplace content, Task/Agent
prompts, WebFetch URLs and rg searches.

Unreadable input writes skill_usage_unchecked plus catstack-hook-error.
report.py --skills prints uses per skill per harness, `no record` for an
installed skill never used, and exits 2 on unchecked runs.
CATSTACK_SKILL_USAGE_LOG=0 opts out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Change-Id: I4c41be3269cd80190cbabad6d973de15008065f3
@EdbertChan
EdbertChan force-pushed the hooks/skill-usage-metrics-all-harnesses branch from f90ccbd to ed7e548 Compare September 24, 2026 07:55
@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_413fd765-f368-483f-a7e5-2f8a1db82e3a)

@EdbertChan

Copy link
Copy Markdown
Owner Author

@Mergifyio requeue

@mergify

mergify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merge Queue Status

This pull request spent 30 minutes 22 seconds in the queue, including 29 minutes 58 seconds running CI.

Required conditions to merge
  • check-success = lint
  • check-success = test
  • check-success = validate

@mergify mergify Bot added the queued label Sep 24, 2026
@mergify mergify Bot mentioned this pull request Sep 24, 2026
6 tasks done
@mergify
mergify Bot merged commit dab263c into main Sep 24, 2026
6 checks passed
@mergify mergify Bot removed the queued label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant