Skip to main content
Research•8 min read

Does memory beat a bigger model? I tested it on my own routing work.

October 2, 2026 by Asif Waliuddin

evalsretrievalcontextjevtypesafemodel-selection
Does memory beat a bigger model? I tested it on my own routing work.

In Part 1, no model could route my company's own operations work from the bare text. Asked who wrote a team post, which program owns a work item and which team owns a commitment, the best model was right 21%, 58% and 44% of the time. Those questions had the names, mentions and IDs stripped out, so no model had any history to go on.

My guess was missing context, not a missing model. So I tested it: before answering, every model got up to five of the most similar earlier records, each with its known answer. I published the plan before a single context call, with the bars and the verdict words written down.

The short answer: it depends on the question. On "who wrote this post", context took Claude Opus 5.5 from 23% to 73%. On "which program owns this work item", neither model showed a clear effect.

What I ran

The same 144 held-out questions as Part 1, 48 in each of the three classes. Each answer is one of a fixed list of options: 20 for authorship, 11 for programs, 9 for owners.

  • Jev (jev-1.13.0) and Claude Opus 5.5, each asked twice in the same run: once with the five records and once without. Asking both ways side by side matters. In Part 1, Opus changed its answer across three asks on 21 of 100 questions, so comparing against last week's run could mistake noise for an effect.
  • Memory alone: no model at all, just a majority vote of the five records' labels.
  • Qwen3-14B with the five records, on one local 4090. It has no side-by-side no-context run, so I report it as a secondary read only.

The records come from an exact nearest-neighbour search over a frozen snapshot of my company's records, embedded with my memory store's own embedding model (nomic-embed-text-v1.5, cut to 384 dimensions). The store itself was not queried. Three rules keep the test honest:

  • Only records created before the question's own record are eligible, because a real router only has the past. On one question that left a single eligible record instead of five.
  • Every option's name is blanked out of the record text. The answer label next to each record stays visible.
  • Each record is cut to 400 characters.

One certifying run, 864 answers. It passed all 17 automated certification gates from Part 1, and a second reviewer reproduced every hypothesis row from the raw records before I wrote a word of this.

The results

In Part 1, the best model without context scored 21%, 58% and 44% on these three questions. Asked again without context in this run, side by side with the context asks, Opus scored 23%, 60% and 35%. That spread is why every comparison below pairs the two asks in the same run.

Accuracy without context, then with the five records:

Question (options)ChanceJevOpusMemory alone
Who wrote this post (20)5%13% to 38%23% to 73%25%
Which program owns this (11)9%42% to 52%60% to 65%25%
Which team owns this (9)11%31% to 35%35% to 58%27%

Against the bar I set in advance (the 95% interval entirely above zero and a gain of at least 10 points), context helps Jev on authorship (+25 points) and helps Opus on authorship (+50) and ownership (+23). Everything else is no clear effect.

Accuracy with and without five retrieved records, per question, for Jev, Opus and memory alone

Where memory helped

Authorship is the clean case. Five similar past posts, with their authors attached, moved Opus from 11 right out of 48 to 35. Jev tripled, from 6 to 18.

On authorship, the right answer was often in front of the model: on 33 of the 48 questions, the right author appeared among the five records' labels. That alone does not explain the gains. On ownership the right answer appeared just as often, 32 of 48, and Jev showed no clear effect there.

Where it showed no clear effect

On "which program owns this work item", context moved Jev by 10 points and Opus by 4. Neither interval clears zero, so by the plan's rules both are no clear effect.

The likely reason: the right program appeared among the five records on only 13 of 48 questions, so there was little for a model to use. My guess at why is that work items read alike across programs, so the closest wording often belongs to a different program. I have not tested that guess.

My hypothesis is that this is a structure question, not a text question: which item sits under which program. My memory store holds that as a graph, and this test never asked the graph. It used the store's embedding model over a frozen snapshot, one part of a three-part store. So this null says nothing about the store as a whole. Testing the graph is next, and I'm not claiming the answer before then.

Opus and Jev, both with context

With the same records, Opus scored 35, 12.5 and 23 points higher than Jev on the three questions. The plan treats this comparison as descriptive, so here are the intervals instead of a verdict: 95% intervals of 21 to 50 points, 0 to 25, and 10 to 37.5. At list price, the Opus answers with context cost about 470 times what Jev's did on these questions, and took a median 6.9 seconds against 0.15.

Memory alone matched Opus-without-context on authorship, 25% against 23%. That sounds like more than it is. Both are about five times chance and both are low. Give Opus the same records and it reaches 73%. On programs, memory alone trailed Opus-without-context by 35 points.

Jev's confidence, again

Part 1 found that Jev's stated confidence runs ahead of its accuracy. Context made that worse on one question. On ownership, the five records raised Jev's average confidence by 25 points and its accuracy by 4. Its calibration error roughly doubled, from 0.171 to 0.335.

Opus moved both together on the same question: confidence up 21 points, accuracy up 23.

The local model is a separate story. Qwen3-14B put 99% or more on its answer for most questions in every class and was right about a third of the time. It is far too confident to use as a probability.

How I kept myself honest

  • The plan went public first. The bars, the arms and the retrieval rules are in the public plan, committed before any context call. The first amendment was public before any context call too. The second, which fixed the memory-alone comparison and its verdict words, went public before the certifying run but after two wiring-check calls that produced no reported number.
  • Side-by-side comparisons. Every with-versus-without number pairs the same question in the same run. Against Part 1's records instead, one verdict changes: Opus on ownership goes from "helps" to "no clear effect", because Opus without context got 17 right this time and 21 last time on the same questions. That is the re-ask noise the plan was built to control.
  • Strict reading. Three intervals have a bound at exactly zero. I read "entirely above zero" and "entirely below zero" strictly, so a bound of zero never counts.
  • A second reviewer reproduced every hypothesis row from the raw records using their own code. A few numbers here are mine alone: how often the right answer appeared among the five records, the average-confidence shifts, the local model's share of 99% answers, and the cost and speed figures. An independent reviewer then recomputed every number in this article from the raw records.
  • You can check it. The questions are private company records, so I can't release them. The public results include a per-question file with no text: right or wrong, and the stated confidence, for every model and question. One script recomputes every accuracy, interval and verdict above from it.

What I'm doing with it

For "who wrote this" style routing, retrieval plus a model works, and the model choice is a cost decision: Opus for accuracy, Jev for speed and price.

For "which program", five similar past records were not enough, and the next test asks the graph directly. For "who owns", context helped Opus and not Jev.

And Jev's confidence still doesn't set a threshold, with or without context. Use its label. Measure its number per class.

The plan, the public results and the reproduction script are at github.com/nxtg-ai/jev-calibration-eval. Part 1 is here.

Ready to build?

Ship AI you can trust

Forge gives you agents, governance, and verification — so your AI ships with confidence, not hope.

Newsletter

Enjoyed this article?

Get more insights like this delivered straight to your inbox.

Email subscription coming soon. Follow along on LinkedIn in the meantime.

Follow on LinkedIn