Skip to content
#

terminal-bench

Here are 53 public repositories matching this topic...

coder_eval

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

  • Updated Sep 6, 2026
  • Python

Add this topic to your repo

To associate your repository with the terminal-bench topic, visit your repo's landing page and select "manage topics."

Learn more