My last benchmark said a code graph did not make my coding agent cheaper. It was too small to say more than that: 24 questions, nothing bigger than 25k graph nodes, no Opus. So I ran it again at a size where the answer could change, and wrote the plan down first. The hypotheses, repositories, budget and analysis are in PLAN-v2.md, committed before any run.
The short version: a graph pays when the model would otherwise keep searching. Haiku on kubernetes and vscode is that case. Sonnet saves a little there and opus nothing, and on a small repository the graph costs sonnet more than plain grep.
The setup
The agent is GitHub Copilot CLI 1.0.90, run headless, one question per session, with two arms:
- grep: Copilot's view, grep and glob tools and nothing else;
- graph: the same three tools, plus cbm-lean (my thin wrapper around codebase-memory-mcp 0.9.0) and a routing rule at the top of the prompt: the graph for callers, callees and call chains, grep for exact strings.
Four public repositories, each pinned to a commit:
| repository | language | graph nodes |
|---|---|---|
| full-stack-fastapi-template | Python, TypeScript | 1,601 |
| dub | TypeScript | 25,023 |
| kubernetes | Go | 146,421 |
| vscode | TypeScript | 218,883 |
Each repository has 16 questions. Eight are structural: who calls a function, what it calls, the call path from a
CLI command or an HTTP handler down to a named function, and which functions break if a signature changes. Eight
are exact: a config value, where an environment variable is read, the line an error is raised on, the line a type
is defined on. I checked every answer key by hand at the pinned commit, and an entry counts only as a whole word,
so init_db does not pass for init.
Haiku 4.5 answered all 64 questions three times in each arm, and sonnet 5 twice. Opus 5.5 costs 15 premium requests a prompt in Copilot, so it answered twelve structural questions from kubernetes and vscode once. That is 664 runs and about 744 premium requests.
Cost is tokenEquiv, as in the first note: uncached input, plus 1.25 × cache writes, plus 0.1 × cache reads, summed over the model calls of a run.
Results
The graph's cost as a share of grep's: for each question the graph arm's median over repetitions divided by the grep arm's, then the median over questions, with a 95% bootstrap interval. Below 1 the graph is cheaper.
| model | repository size | structural | exact |
|---|---|---|---|
| haiku 4.5 | small (1.6k nodes) | 0.88 (0.65–1.29) | 0.92 (0.72–1.15) |
| haiku 4.5 | medium (25k) | 0.61 (0.11–1.50) | 0.96 (0.75–1.37) |
| haiku 4.5 | large (146k, 219k) | 0.29 (0.11–0.70) | 1.01 (0.94–1.28) |
| sonnet 5 | small | 1.51 (1.45–1.70) | 1.42 (1.20–1.71) |
| sonnet 5 | medium | 0.94 (0.57–2.00) | 1.14 (1.05–1.30) |
| sonnet 5 | large | 0.74 (0.62–0.92) | 1.13 (1.01–1.29) |
| opus 5.5 | large | 0.97 (0.72–1.10) | – |
Correct answers on the structural questions, graph against grep:
| model | small | medium | large |
|---|---|---|---|
| haiku 4.5 | 13/24 against 19/24 | 20/24 against 22/24 | 38/47 against 33/48 |
| sonnet 5 | 13/16 against 15/16 | 16/16 against 16/16 | 32/32 against 28/32 |
| opus 5.5 | – | – | 12/12 against 12/12 |
Cost and accuracy together, as tokenEquiv per correct answer on the structural questions:

What holds up:
- On kubernetes and vscode, haiku with the graph spent 0.29 of what it spent with grep on structural questions and was cheaper on 15 of 16 of them. It was also right more often, so per correct answer the gap is 47k against 209k tokenEquiv. This is the only result that survives the Holm correction across the 13 tests (p = 0.006).
- Sonnet on the same questions: 0.74, cheaper on 12 of 16, and 32 of 32 correct against 28 of 32. Opus: 0.97, cheaper on 7 of 12, and every answer right in both arms.
- On the small repository the graph never paid. Haiku spent about the same and got structural questions wrong more often. Sonnet spent 42 to 51% more.
- On exact questions the graph made no measurable difference for haiku, and sonnet paid 13 to 42% more for it.
- The graph's tool schemas and the routing rule add a fixed 1.56k prompt tokens to every model call for haiku and 2.07k for sonnet and opus, at every repository size.
Why the small model gains most
Count the model calls. With grep, haiku needed a median of 22 calls per structural answer on kubernetes and 20 on vscode, and one impact question took 114. With the graph it needed 5 and 4. Every call sends the conversation again, so a long search costs far more than its call count suggests: haiku's most expensive grep answer came to 471k tokenEquiv.
Sonnet's searches with grep were shorter, 8 calls against 4 with the graph. Opus took 4 or 5 calls either way, so the graph had nothing to cut, and its fixed overhead ate what little it saved.
My guess, which I did not test directly: a graph replaces search turns, so it pays when the model would otherwise take many of them, which means a weaker model on a larger repository. A strong model already searches efficiently with grep, and on a small repository nobody needs many turns.
Copilot bills premium requests per prompt, whatever the number of calls, so inside Copilot the graph saves no money. It saves time: haiku's structural answers on the large repositories took a median of 71 s with the graph and 157 s with grep. On the small repository the graph arm was slower.
Where each arm went wrong
- A field called
lines. Asked on which lineHorizontalControlleris defined, the graph arm answered horizontal.go:43 in all five runs, haiku and sonnet alike. The struct starts on line 93 and is 43 lines long. The snippet that codebase-memory-mcp returns carries"start_line": 93,"end_line": 135and"lines": 43, and the models read the last one as the line number. It happened again withLemonSqueezyClient(client.ts:129 for a class on line 59 that is 129 lines long), and haiku did the same withCursorsController. My first note blamed "line numbers from the graph" for that client.ts:129; this is where it came from. Sonnet with grep got all three right. cbm-lean now passes the field on asline_count. - Holes in the graph, trusted. codebase-memory-mcp 0.9.0 records no call edge for a bare call to a function imported
with
from x import y. Asked whatrecover_passwordcalls, the graph arm answered "only get_user_by_email" in all five runs, haiku and sonnet alike, straight from the graph and without a single grep. The function also callsgenerate_password_reset_token,generate_reset_password_emailandsend_email, and the grep arm found all four every time. - A missing dash. Copilot's grep tool shows line numbers only when it gets
"-n": true, and its view tool returns a file without them. Sonnet passed-nin 508 of its 571 grep calls and opus in all 53. Haiku passedn, without the dash, 920 times and-n22 times out of 1,949. The tool ignores the unknown key, so haiku got no numbers, counted lines in the viewed file and missed by a few: for an error raised on line 112 it answered 105, 107, 108 and 113. All but one of haiku's 79 wrong answers to exact questions are a wrong line number. - Long searches that still fail. On the call chain from
kubeadm token createdown toencodeTokenSecretData, haiku with grep took 18 to 22 calls and got the path wrong all three times. It was wrong with the graph too. Sonnet was right in both arms. - Haiku tried to call bash three times in the grep arm. Copilot refused, since the tool does not exist there, so no run used a tool outside its arm.
What changed since the first note
The first note said not to expect token savings from a graph with Claude models on repositories up to 25k nodes. That still holds at those sizes. It missed what happens above them: on kubernetes and vscode the smaller models gain, and the smaller the model, the more.
What I would tell a team wiring a graph into an agent
- Measure on your largest repository with your cheapest model first. That is where a graph earns its place.
- With a strong model on a small or medium repository, leave it out. You pay for its schemas on every call and get nothing back.
- Keep routing exact questions to grep. The graph never won them.
- Do not let the model take line numbers or "this is the complete list" from graph output on trust. Drop or rename
fields like
linesbefore they reach the model.
Limits
- One agent harness (Copilot CLI) and one graph (codebase-memory-mcp 0.9.0 behind my wrapper). Claude Code, another graph or a newer version could give other numbers.
- Eight questions per cell on the small and medium repositories. With 13 tests under the Holm correction, a cell of eight cannot reach significance even if the graph wins every question, so only the large-repository haiku result is significant.
- Opus answered twelve questions, once each.
- tokenEquiv leaves output tokens out. Counting them would widen haiku's gap on the large repositories: its grep answers there produced a median of 7.0k output tokens against 0.9k with the graph.
- One haiku run failed before Copilot opened a session and is left out. Three sonnet runs failed to fetch Copilot's model list before any model call and were run again. Both are recorded in the plan.
Reproduce
Everything is in code-graph-vs-grep: the plan written before the runs, the questions with the evidence for every answer key, the driver, all 664 runs with their answers, and the analysis.
node bench/run-copilot.mjs --model claude-haiku-4.5 --reps 3 --out results/v2/claude-haiku-4.5.json
python3 bench/stats-v2.py results/v2/*.json