ArticleResearch8 min
Does memory beat a bigger model? I tested it on my own routing work.
Part 2 of my Jev test, pre-registered before the first call: up to five similar past records added to every prompt. Opus went from 23% to 73% on one routing question, and neither model showed a clear effect on another. Per class results for Jev, Claude Opus 5.5 and memory alone.
evalsretrievalcontext+3
Asif Waliuddin
Oct 2, 2026
