← Back to the Lab

A code graph did not make my coding agent cheaper

benchmarkmcpcoding-agentsevals

Two numbers sent me back to this benchmark. The README of codebase-memory-mcp, a tree-sitter code graph served over MCP, promises 99% fewer tokens. My own earlier run said that cbm-lean, a thin wrapper I wrote around it, made a coding agent 26% cheaper than plain grep on haiku and 18% cheaper on sonnet. I could never reproduce the first number, and the second came from a flaw in my setup. This is the re-run on public repositories, with that flaw fixed.

The setup

Every question ran under three arms:

Each repository got twelve questions. Six are structural: who calls a function, what it calls, the call chain from an HTTP endpoint down to the database write, the module layout. Six are exact: a config value, where an environment variable is read, an error message, the line where a class is defined. Each question lists the substrings a correct answer has to contain, and I checked every one against the code by hand.

The repositories are three small ones (the FastAPI full-stack template, ltx-2-mlx and avoid-ai-writing, 1.6k to 5k graph nodes) and dub, a Next.js monorepo with 25k nodes. Haiku 4.5 ran on all of them and sonnet 5 on dub: 252 runs, $14.69 at API prices.

Cost is measured as tokenEquiv: input tokens, plus 1.25 × cache writes, plus 0.1 × cache reads. Raw billed input is misleading once prompt caching is on, because an arm that happens to run on a warm cache looks cheap for reasons that have nothing to do with the arm.

Results

Median tokenEquiv per run, with the number of correct answers in brackets:

settingrunscbm-leanraw graphgrep
Small repos, haiku 4.510830,321 (28/36)43,846 (32/36)31,159 (33/36)
dub, haiku 4.57218,791 (22/24)20,687 (22/24)14,955 (20/24)
dub, sonnet 57217,227 (22/24)23,085 (22/24)13,348 (22/24)

Three results hold up:

The p values come from a Wilcoxon signed-rank test over per-question medians.

Where each arm went wrong

Mounting a graph is not using it

The most useful result came from the pilot, not the main run. Claude Code 2.1.283 defers MCP tool schemas behind a ToolSearch step by default: the model sees the tool names and has to fetch the schemas before it can call anything. Headless haiku did not do that. In six pilot runs the graph was never called, with or without the routing rule. On the twelve small-repository questions:

setupcalled the graphcorrectmedian tokenEquiv
Default loading, with the rule2 of 1211 of 1241,008
Schemas loaded, no rule4 of 129 of 1262,008
Schemas loaded, with the rule22 of 3628 of 3630,321

With the schemas loaded (ENABLE_TOOL_SEARCH=false), the same callers question went through two graph calls in three turns. The mounted graph with no rule was the most expensive setup of all: every session pays for the schemas, and the model mostly leaves them unused.

If you ship an MCP server and test it with default settings, first check that the model calls it at all.

Why the earlier number was wrong

The first benchmark ran inside my projects folder. Its CLAUDE.md told the agent to use the graph for structural questions, and Claude Code loads CLAUDE.md files from parent directories, so that instruction reached every arm, including the one with no graph. The control group was given the treatment's instructions, which means the comparison measured the rule as much as the tools.

The re-run keeps the repositories outside any tree with a CLAUDE.md above it, skips user settings and hooks (--setting-sources project), and gives the rule only to the arms that have a graph. The other controls each fixed a wrong conclusion in an earlier round. Bash and Agent are denied in every arm, because scoped Bash grants do not match piped commands and subagents hide turns while still costing tokens. Arm order rotates between repetitions, so no arm always runs on a warm cache. A run that calls a tool its arm does not own is discarded. And answers are graded, because tokens saved on a wrong answer are not savings.

What I would tell a team wiring a graph into an agent

Limits

Two models, four repositories up to 25k nodes, two or three repetitions, twelve questions each. No Opus, no long interactive sessions, nothing beyond 25k nodes. Medians are reported because repeated runs of the same question varied widely.

Reproduce

Everything is in code-graph-vs-grep: the questions with their pinned commits, the benchmark driver, the raw runs with every answer, and cbm-lean itself.

bash bench/setup.sh
node bench/run.mjs --reps 3
python3 bench/stats.py results/*.json
A code graph did not make my coding agent cheaper · AllKeep