CodeGraph, a pre-indexed code knowledge graph for AI coding agents, published a re-measurement on 2026-08-05 that puts its savings at 62% fewer tokens and 44% lower cost on average across seven open-source repositories. The project’s README reports the agent answering architecture questions in one to four codegraph_explore calls with zero file reads on every benchmark repo, against up to 43 tool calls and 19 file reads for an agent working from grep and Read alone.
The surprising part is not the savings. It is the row directly underneath them. In multi-turn sessions on the same seven repos, CodeGraph’s responses leave roughly 80% more retrieval context resident at the end of a session than a file-reading agent’s do. On VS Code, that is 67k tokens against 18k. The tool that processes fewer tokens also occupies more of your window.
The mechanism is not a bug
CodeGraph returns one dense, verbatim payload that answers the question. A grep-and-read agent churns through many small results, most of which get evicted as the session continues. Fewer tokens processed and a larger persistent footprint are both real at once, and the README says so plainly rather than burying it: “If you run long sessions in a small window, budget for it.”
That is an unusual disclosure. Most developer-tool benchmarks report the number that flatters the tool and stop. CodeGraph’s authors measured the axis where they lose, published the per-repo file (docs/benchmarks/residual-context-occupancy.md), and told long-session users to plan around it. The honesty is worth more than the headline number.
What the benchmark actually controls for
The methodology is stricter than most agent-tool claims. Both arms run claude -p headless against the repo with --strict-mcp-config, at the median of four runs per arm, on Claude Opus 4.8 (claude-opus-4-8). Built-in Read, Grep, and Bash stay available to both arms. Same question per repo.
The important control: the codegraph CLI is blocked in both arms via a sanitized PATH plus a PreToolUse hook. Without that block, the authors found the WITHOUT agent locating the CLI on PATH and reaching CodeGraph through Bash in 26 of 28 runs, which would contaminate the comparison in both directions, since a CLI call is not counted as a tool call and its output still enters the window. In the reported run, all 28 WITHOUT runs attempted the CLI and all 28 were blocked. Contamination row: 0 of 28.
This is the detail that separates a benchmark from marketing. An earlier published figure was produced without the block. The team re-measured and said so.
Where the savings hold and where they vanish
The per-repo table is more interesting than the average. VS Code (~11k TypeScript files) went from 28 tool calls to 2, 2m10s to 58s, and $1.80 to $0.53. Excalidraw (~640 files) went from 43 tool calls to 2 and $2.43 to $0.54. Tokio (~790 Rust files) went from 29 calls to 3 and $1.83 to $0.66.
Then there is Gin. Go, ~110 files, 7 tool calls to 1, 46s to 28s, 180k tokens to 87k, and cost roughly even at $0.31 against $0.31. Django shows the same pattern: 14 tool calls to 3, 309k tokens to 183k, but only 13% cheaper.
The README explains why, and the explanation is the most useful sentence in the document: cost tracks how much discovery a question demands, not raw repo size. On repos where the file-reading arm needed 28 to 43 tool calls, CodeGraph saves 57% to 78% on cost. Where the agent got there in 7 or 14 calls, the savings compress toward zero. The with-arm still answered in 1 and 3 calls with zero file reads. It just did not have much waste to eliminate.
Fewer tokens processed and a larger persistent footprint are both real at once.
The Rust kernel and the 20-language claim
The parsing engine is a native Rust kernel covering 20 languages, including TypeScript, Python, Go, Rust, C++, C#, Swift, Kotlin, Scala, Dart, R, Lua, and Luau, with Metal and CUDA riding the C++ path. The README claims each language shipped only after its graphs proved byte-for-byte identical to a reference engine on real repositories, up to the Linux kernel, with per-file fallback for platforms lacking a prebuilt binary and for files with syntax errors. That is a strong claim and an unverified one; we have not independently reproduced it. Treat it as the vendor’s stated bar, not a measured fact.
The distribution story matters for adoption. No Node.js required: a curl or PowerShell one-liner grabs the right build, and npm works on any version. codegraph install auto-detects and wires the MCP server into Claude Code, Cursor, Codex CLI, opencode, Hermes Agent, Gemini CLI, Antigravity IDE, Kiro, and GitHub Copilot across VS Code, Copilot CLI, and JetBrains IDEs. codegraph init builds the graph per project, and auto-sync updates it on every file change by default. codegraph uninstall reverses the agent configuration and leaves project indexes in place.
Why this matters for AI builders
The agent-tooling market has spent two years optimizing for retrieval quality and mostly ignoring retrieval cost. CodeGraph’s numbers make the case that discovery, not reasoning, is where agent budgets go. Up to 43 tool calls and 19 file reads to reconstruct a call graph the model could have been handed in one call is pure overhead, and the 57% to 78% cost savings on those repos is the price of that overhead made visible.
The residual-context finding cuts the other way and deserves more attention than it will get. If dense graph payloads persist in the window while small grep results get evicted, then the tool’s advantage in throughput and its disadvantage in occupancy scale with session length. An agent that answers in one call early in a session may be carrying that answer for the next forty turns.
The platform is also expanding. The README teases a hosted CodeGraph product for pull requests, promising to identify what to test, what could break, which flows are affected, and whether business logic is compromised, with early beta access at getcodegraph.com. That is a different business than a local CLI, and it is where the local-first framing will get tested.
For now, the practical read is narrow. CodeGraph pays off most on large repos and questions that demand real discovery, and it pays off least where a strong model was already going to find the answer in a handful of calls. The benchmark says as much, which is more than most tools in this category are willing to admit.