What we wanted to know
Code written by agents carries reasons that live in conversations: a request, a review, a ruling from the person who owns the work. When the next agent arrives, the code alone rarely says why it is the way it is. Some code looks wrong but was a deliberate choice. Some looks fine but was never asked for.
We asked two questions. When an agent needs to know why code exists, does Engram help it find the answer more completely or more reliably than the tools agents already use, such as rg, grep and git? And is it faster or cheaper?
How we tested it
- Real code, chosen at random. We drew 30 code edits at random from a month of transcripts from our own agent-built projects, two codebases written almost entirely by coding agents. We kept the 12 whose history we could establish with confidence.
- Answer keys built without Engram. For each case, we reconstructed the true history from raw transcripts and git: which conversation wrote the code, the task it was doing, who asked for it and any later decisions. Engram was not used, so it was not graded against itself.
- Same question, same agents. Each case went to fresh agents on two models, Claude Opus 5.5 and GPT-6 Luna, once with Engram and once with any tools except Engram. The question was neutral: "Tell me why this code is the way it is: where it came from, what it is for, and what has been decided about it." Each run had 15 minutes and a cutoff date, so no agent could see the future.
- Blind grading. Graders scored each answer against the key's required facts without knowing which group produced it.
- Measured cost. We recorded wall time, tokens and dollar cost for every run.
Separately, we ran a repair task where the right answer depended on a later ruling, and we compared how much later discussion Engram returned with a plain search for the file's name.
| Group | Key facts found | Median time | Median input tokens | Total cost, 12 runs |
|---|---|---|---|---|
| Opus 5.5 + Engram | 76% | 306 s | 1.9M | $22.43 |
| Opus 5.5, no Engram | 50% | 216 s | 1.3M | $16.02 |
| Luna + Engram | 42% | 675 s | 1.7M | n/a |
| Luna, no Engram | 32% | 517 s | 0.9M | n/a |
What the numbers don't say
- This is a small study: 12 graded cases, one run each per group, and one repair task. It shows a consistent effect, not a precise one.
- We tuned the later-discussion results while looking at the repair task, so treat that task as a demonstration. The 12 study cases were not used for tuning.
- Engram does not make a single lookup faster. Agents with Engram spent longer because they found more history and read it.
- The grading was strict. Only 4 of 48 answers were fully correct, mostly because few traced a request all the way back to the person who made it.
- The cases come from our own projects, where transcripts are complete. Results may differ where history is patchier.
- The results used a development build of
engram explainthat also lists later discussion of the code. That part is not yet in a release.
What it means
If you want an agent to understand code before it changes it, have it run engram explain first. Expect a more complete answer, including decisions that exist only in conversation, at about the same cost per fact.