CodeGraph Quality Audit

SkillAI & models

agent-eval is a skill that benchmarks CodeGraph retrieval quality on a real codebase. It runs a comparison of agent behavior with and without CodeGraph indexing on a selected open-source repository, then reports costs, tool-call counts, and side-by-side comparison tables so the person using the agent can see whether a given codegraph version actually helps.

Available today. Use it from your connected AI after setup.

Have tmux 3+, a logged-in claude CLI, node, and git installed on macOS or Linux.

Then ask your AI: use the CodeGraph Quality Audit skill

What your AI can do with it

  • Compare agent performance with and without CodeGraph indexing on a real repo
  • Test a local dev build, the latest published npm version, or a specific version like 0.7.1
  • Run headless audits with exact token and cost metrics via parse-run.mjs
  • Run interactive tmux audits that drive the real TUI and read session logs
  • Report residual context occupancy, explore sufficiency, and allocation efficiency
  • Produce a side-by-side ARM COMPARISON table with a contamination check

Getting started

  1. Have tmux 3+, a logged-in claude CLI, node, and git installed on macOS or Linux.
  2. Run from the codegraph repo root, where the harness lives in scripts/agent-eval/.
  3. Trigger the skill by running /agent-eval or asking the agent to test, benchmark, audit, or validate a codegraph version.
  4. Answer the agent's questions to pick the codegraph version, language, target repo, and harness mode.
  5. Let the agent launch audit.sh in the background and report the results when the job finishes.

What this skill tells your AI

The instructions your AI receives, as published by colbymchenry/codegraph in .claude/skills/agent-eval/SKILL.md and read by ahel’s review.

Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in scripts/agent-eval/.

Prerequisites

  • tmux 3+, a logged-in claude CLI, node, git (macOS/Linux).
  • Run from the codegraph repo root.

Workflow

Copy this checklist:

- [ ] 1. Pick version (local or npm)
- [ ] 2. Pick language
- [ ] 3. Pick repo by size
- [ ] 4. Pick harness (headless / tmux / both)
- [ ] 5. Run audit.sh in the background
- [ ] 6. Report results

Step 1 — version. Ask with AskUserQuestion: which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g. 0.7.10). Map the answer to a VERSION token:

  • "Local dev build" → local
  • "Latest published" → latest
  • a typed version → that string (e.g. 0.7.10)

Step 2 — language. Read .claude/skills/agent-eval/corpus.json. Ask with AskUserQuestion which language to test, listing the languages that have entries.

Step 3 — repo. From the chosen language's entries, ask which repo. Label each option with its size and file count, e.g. excalidraw — Medium (~600 files). Each entry carries the repo URL and a representative question.

Step 4 — harness. Ask with AskUserQuestion which harness to run, and map the answer to a MODE token:

  • "Headless" → headless — claude -p with stream-json: exact tokens/cost and a clean tool sequence (2 runs, fast, no TTY).
  • "Interactive (tmux)" → tmux — drives the real Claude TUI in tmux: faithful Explore-subagent behavior, metrics from session logs (2 runs, slower).
  • "Both" → all — headless + interactive (4 runs).

Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes):

scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>

Step 6 — report. When the job finishes, read the log and report per arm:

  • Headless (parse-run.mjs): total tool calls, file Reads, Grep/Bash, codegraph-tool calls, duration, total cost.
  • Interactive (parse-session.mjs): the VERDICT: codegraph_explore used Nx | Read N | Grep/Bash N and TOKENS: lines.
  • Both paths also print the three feedback metrics — residual context occupancy, explore sufficiency, allocation efficiency — and a headless A/B ends with a side-by-side ARM COMPARISON table. Report that table, and check its contamination row first: CLI calls that RETURNED output > 0 means the arm reached codegraph through Bash and its numbers are void. How to read the rest: docs/benchmarks/agent-eval-feedback-metrics.md.

Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer.

Notes

  • The index is rebuilt every run (audit.sh wipes .codegraph) — different versions extract differently, so an index must be served by the same binary that built it.
  • audit.sh temporarily mutates the global codegraph install for the test, then restores your dev link via local-install.sh.
  • Corpus repos are cloned to /tmp/codegraph-corpus (reused if already present).
  • Add or edit repos in corpus.json (fields: name, repo, size, files, question).

Signals

GitHub stars
72k
Forks
5k
Last commit
Sep 2026
Hacker News mentions
2

Questions

How to evaluate Claude skill?
agent-eval evaluates CodeGraph by running a benchmark script that compares agent behavior with and without CodeGraph indexing on a chosen open-source repo, then reporting costs, tool-call counts, and comparison tables.
What is Claude Eval?
In this skill it refers to the CodeGraph quality audit: a harness in scripts/agent-eval/ that measures how much CodeGraph helps an agent versus plain grep/read on a real-world repository.
Advanced
Item type
skill
Key
codegraph-agent-eval
Source
github.com/colbymchenry/codegraph