← Back to the Lab

Where a code graph pays off: big repositories, small models

benchmarkmcpcoding-agentsevals

My last benchmark said a code graph did not make my coding agent cheaper. It was too small to say more than that: 24 questions, nothing bigger than 25k graph nodes, no Opus. So I ran it again at a size where the answer could change, and wrote the plan down first. The hypotheses, repositories, budget and analysis are in PLAN-v2.md, committed before any run.

The short version: a graph pays when the model would otherwise keep searching. Haiku on kubernetes and vscode is that case. Sonnet saves a little there and opus nothing, and on a small repository the graph costs sonnet more than plain grep.

The setup

The agent is GitHub Copilot CLI 1.0.90, run headless, one question per session, with two arms:

Four public repositories, each pinned to a commit:

repositorylanguagegraph nodes
full-stack-fastapi-templatePython, TypeScript1,601
dubTypeScript25,023
kubernetesGo146,421
vscodeTypeScript218,883

Each repository has 16 questions. Eight are structural: who calls a function, what it calls, the call path from a CLI command or an HTTP handler down to a named function, and which functions break if a signature changes. Eight are exact: a config value, where an environment variable is read, the line an error is raised on, the line a type is defined on. I checked every answer key by hand at the pinned commit, and an entry counts only as a whole word, so init_db does not pass for init.

Haiku 4.5 answered all 64 questions three times in each arm, and sonnet 5 twice. Opus 5.5 costs 15 premium requests a prompt in Copilot, so it answered twelve structural questions from kubernetes and vscode once. That is 664 runs and about 744 premium requests.

Cost is tokenEquiv, as in the first note: uncached input, plus 1.25 × cache writes, plus 0.1 × cache reads, summed over the model calls of a run.

Results

The graph's cost as a share of grep's: for each question the graph arm's median over repetitions divided by the grep arm's, then the median over questions, with a 95% bootstrap interval. Below 1 the graph is cheaper.

modelrepository sizestructuralexact
haiku 4.5small (1.6k nodes)0.88 (0.65–1.29)0.92 (0.72–1.15)
haiku 4.5medium (25k)0.61 (0.11–1.50)0.96 (0.75–1.37)
haiku 4.5large (146k, 219k)0.29 (0.11–0.70)1.01 (0.94–1.28)
sonnet 5small1.51 (1.45–1.70)1.42 (1.20–1.71)
sonnet 5medium0.94 (0.57–2.00)1.14 (1.05–1.30)
sonnet 5large0.74 (0.62–0.92)1.13 (1.01–1.29)
opus 5.5large0.97 (0.72–1.10)–

Correct answers on the structural questions, graph against grep:

modelsmallmediumlarge
haiku 4.513/24 against 19/2420/24 against 22/2438/47 against 33/48
sonnet 513/16 against 15/1616/16 against 16/1632/32 against 28/32
opus 5.5––12/12 against 12/12

Cost and accuracy together, as tokenEquiv per correct answer on the structural questions:

Tokens per correct answer on structural questions, with the code graph and with grep only, on the FastAPI template and on kubernetes and vscode

What holds up:

Why the small model gains most

Count the model calls. With grep, haiku needed a median of 22 calls per structural answer on kubernetes and 20 on vscode, and one impact question took 114. With the graph it needed 5 and 4. Every call sends the conversation again, so a long search costs far more than its call count suggests: haiku's most expensive grep answer came to 471k tokenEquiv.

Sonnet's searches with grep were shorter, 8 calls against 4 with the graph. Opus took 4 or 5 calls either way, so the graph had nothing to cut, and its fixed overhead ate what little it saved.

My guess, which I did not test directly: a graph replaces search turns, so it pays when the model would otherwise take many of them, which means a weaker model on a larger repository. A strong model already searches efficiently with grep, and on a small repository nobody needs many turns.

Copilot bills premium requests per prompt, whatever the number of calls, so inside Copilot the graph saves no money. It saves time: haiku's structural answers on the large repositories took a median of 71 s with the graph and 157 s with grep. On the small repository the graph arm was slower.

Where each arm went wrong

What changed since the first note

The first note said not to expect token savings from a graph with Claude models on repositories up to 25k nodes. That still holds at those sizes. It missed what happens above them: on kubernetes and vscode the smaller models gain, and the smaller the model, the more.

What I would tell a team wiring a graph into an agent

Limits

Reproduce

Everything is in code-graph-vs-grep: the plan written before the runs, the questions with the evidence for every answer key, the driver, all 664 runs with their answers, and the analysis.

node bench/run-copilot.mjs --model claude-haiku-4.5 --reps 3 --out results/v2/claude-haiku-4.5.json
python3 bench/stats-v2.py results/v2/*.json
Where a code graph pays off: big repositories, small models · Rodion Kazennov