Study · October 2026

Do agents understand code better with Engram?

We asked coding agents why real code exists, with and without Engram. With it, their answers were half again as complete, and they found decisions that ordinary search never reached.

01 · COMPLETENESS

Agents with Engram got more of the story right.

76%With Engram
50%rg, grep and git only

Share of the key facts found, graded blind, across 12 real code spans, by Claude Opus 5.5. A second model, GPT-6 Luna, rose from 32% to 42%.

02 · CONSISTENCY

It wasn't one lucky case.

9 of 12cases where the Engram answer scored higher

Two tied, one scored lower. Luna with Engram scored higher in 8 of 12.

03 · REACH

Engram found the decision that never touched the code.

2/2With Engram
0/2Without

A file's owner had ruled it should move out of the shipped product. The ruling lived in a conversation, not in git. Both agents with Engram reported it; both without it stopped at the commit history.

04 · DECISIONS

On a real repair task, Engram led to the right call, in half the time.

½the wall time · 508 s vs 1,002 s
−26%cost · $3.31 vs $4.49
1 vs 0right calls · with Engram vs without

The task was "fix the failing check." The test had never been asked for, and its owner had already ruled it should be deleted. Opus with Engram read that ruling and deleted the test. Both agents without Engram kept it or escalated, as did all 8 earlier runs. The second Engram agent never opened the result that held the ruling.

05 · FOCUS

Far less to read than searching by hand.

150Engram excerpts
3,425grep matches

Later discussion of 30 random edits: Engram's short, capped list against every conversation that names the file in the same week. About 23 times less to read.

06 · COST

More thorough, not more expensive per fact.

$2.46With Engram
$2.67Without

Cost per key fact found (Opus). Each Engram run took about 40% longer and cost about 40% more, because the agent kept reading the history it found.

What we wanted to know

Code written by agents carries reasons that live in conversations: a request, a review, a ruling from the person who owns the work. When the next agent arrives, the code alone rarely says why it is the way it is. Some code looks wrong but was a deliberate choice. Some looks fine but was never asked for.

We asked two questions. When an agent needs to know why code exists, does Engram help it find the answer more completely or more reliably than the tools agents already use, such as rg, grep and git? And is it faster or cheaper?

How we tested it

Separately, we ran a repair task where the right answer depended on a later ruling, and we compared how much later discussion Engram returned with a plain search for the file's name.

GroupKey facts foundMedian timeMedian input tokensTotal cost, 12 runs
Opus 5.5 + Engram76%306 s1.9M$22.43
Opus 5.5, no Engram50%216 s1.3M$16.02
Luna + Engram42%675 s1.7Mn/a
Luna, no Engram32%517 s0.9Mn/a

What the numbers don't say

What it means

If you want an agent to understand code before it changes it, have it run engram explain first. Expect a more complete answer, including decisions that exist only in conversation, at about the same cost per fact.

engram.ing · github.com/clickety-clacks/engram