skill-usage-log: record every skill use in Claude, Cursor and Codex, on by default - #826
Conversation
5cc2331 to
fe27181
Compare
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_529425e3-ef8b-4b83-8f3f-97f6b703fb54) |
1345e62 to
489189c
Compare
fe27181 to
f90ccbd
Compare
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_0b887549-cd92-49a8-b45f-cb9093f1a430) |
|
Queued — the merge queue status continues in this comment ↓. |
|
Mergify repair stopped: GitHub reports merge conflict. The retry cap was reached for current head f90ccbd. |
…on by default skill-usage-log was switched off unless CATSTACK_SKILL_USAGE_LOG=1, covered only Claude's Skill tool, swallowed write errors, and wrote a private file no report read. ~/.cache/catstack-skill-usage-log did not exist on this machine, so no skill use was on record. It now records one catstack.hook_event.v1 row per use (action skill_used, reason = source, skill = name) in the shared events file, from all three harnesses: - Claude PreToolUse (Skill|Read|Bash) and UserPromptSubmit - Cursor preToolUse and beforeSubmitPrompt - Codex PreToolUse and UserPromptSubmit Sources: skill_tool, read (a Read of a SKILL.md, including Cursor's skills-cursor/), shell_read (cat/sed/head/tail/nl/less/more/bat on one, also inside a Codex exec code string), slash (a prompt starting /<installed skill>), mention (Codex $<installed skill>). Backtest on recent real transcripts: Claude 10 uses counted, Cursor 708 (670 of 671 SKILL.md Read calls, 22 of 22 ReadFile), Codex 1; rejected mentions were Glob/Grep patterns, Write/StrReplace content, Task/Agent prompts, WebFetch URLs and rg searches. Unreadable input writes skill_usage_unchecked plus catstack-hook-error. report.py --skills prints uses per skill per harness, `no record` for an installed skill never used, and exits 2 on unchecked runs. CATSTACK_SKILL_USAGE_LOG=0 opts out. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Change-Id: I4c41be3269cd80190cbabad6d973de15008065f3
… validate; Exit code: 0
f90ccbd to
ed7e548
Compare
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_413fd765-f368-483f-a7e5-2f8a1db82e3a) |
|
@Mergifyio requeue |
Merge Queue Status
This pull request spent 30 minutes 22 seconds in the queue, including 29 minutes 58 seconds running CI. Required conditions to merge
|
Summary
A skill is a saved set of steps an agent can load and follow. Until now nobody could tell which skills the agents really use.
Each agent now writes one line every time it uses a skill, into the same shared file the other usage counts already read.
All three agents do this, not just one, and it is on unless you switch it off.
A new option on the usage summary tool groups those lines by skill, and names the skills no agent has ever loaded.
Review Claim
Each of the four ways an agent can load a skill writes exactly one line, naming the skill, which agent, and how it was loaded. Talking about a skill without loading it writes nothing.
Review Lane
behavior
Review Unit
engine-runtime
Safety Invariant
The hook never blocks or changes a tool call or prompt. Input it cannot read writes
skill_usage_uncheckedand acatstack-hook-errorline, so it is never counted as "no skill used".Slice Rationale
The detector, the three harness entry points, their install wiring and the report that reads the rows are one feature; a harness wired without the others would report zero for skills it simply cannot see.
Non-goals
Does not update the root README's env-var row (a separate docs slice stacked on this one). Does not count a skill whose
SKILL.mda subagent is only told to read; that subagent's own read is counted.Test Plan
Test Plan
Every CI step, run locally at this commit:
test_git_path_churn.py(2 failures) is in a file this diff does not touch.Backtest of the detector on recent real transcripts (tool name and input as recorded):
Real path, Claude:
claude -p --settings <this branch's hook through the runner>, asked to loaddiuwith the Skill tool andcatthe cat-mode SKILL.md:Not exercised on a live harness: Codex and Cursor. Codex runs only hooks whose hash it has trusted, and
cursor-agentis not logged in on this machine. Their payload shapes are covered by the backtest above and by entry-script tests that run each script as a real process.Revert Plan
Revert Plan
Revert this commit, then delete every hook entry whose command contains
skill-usage-log/from~/.cursor/hooks.json,~/.codex/hooks.jsonand theUserPromptSubmitlist in~/.claude/settings.json, and reruninstall.sh. The old installer only replaces the ClaudePreToolUseentry, so the other entries would keep calling scripts the revert deletes. Rows already written stay in the events file; nothing else readsskill_used.🤖 Generated with Claude Code
Note
Low Risk
Metrics-only hook that never blocks agent actions; main risk is extra PreToolUse/prompt hook runs and local event file growth, not security or data handling.
Overview
Replaces the old Claude-only, opt-in
skill-usage-log(private JSONL under~/.cache/catstack-skill-usage-log) with silent metrics on Claude, Cursor, and Codex that appendskill_usedrows to the shared hook metrics event log.Detection covers Skill tool calls, direct/shell reads of
SKILL.md, leading/skillprompts, and Codex$skillmentions, while ignoring edits, searches, and incidental path mentions. Sharedlog.py+detect.pyback thin harness entry scripts;write_stage_eventnow accepts extrafields(e.g.skillname). Unreadable input or skill dirs emitskill_usage_uncheckedand surface viareport.py --skills(exit 2).Install merges hook fragments for all three agents via
install_common.py;hooks.tomldrops the oldCATSTACK_SKILL_USAGE_LOGenable gate (opt out with=0). Docs and tests cover reporting, detection, and install wiring.Reviewed by Cursor Bugbot for commit ed7e548. Bugbot is set up for automated code reviews on this repo. Configure here.