We Hid a Fact at Different Depths in Long Contexts: What the Models Actually Remembered

#1Long-context needle tests: what the models actually remember
Long-context models can hold a lot of text, but that does not mean they use every part of it equally well.
This needle-in-a-haystack test checks where the model still finds the fact and where it starts missing it, especially when the fact sits in the middle of a long prompt.
#2Why this test matters
A lot of the excitement around long context is that “it can fit a whole document or a whole repo in one request.” That sounds powerful, but it does not necessarily mean the model will use that information effectively.
The test we care about is simple: hide a fact in a long body of filler text, then ask the model to retrieve it. If the model misses the fact, that is not a small effect. It is a sign that reasoning and retrieval quality degrade as context grows.
This is particularly relevant for RAG systems, document QA, and long-code assistant workflows. Those systems often rely on long input windows, and they need to know whether the model actually attends to the important content or just the most recent or most salient parts.
#2Methodology
We ran a standard needle-in-a-haystack test with the following structure:
- Build a large block of filler text.
- Place a specific fact at different positions in the context.
- Ask the model to retrieve the fact.
- Score exact matches on both the value and the associated descriptor.
The fact used in the test was:
“The project codename was BLUEBERRY and the deployment region was eu-west-3.”
We varied the position of the fact across the document and checked results at multiple context depths.
Models tested:
- GPT-4o
- Claude 3.5 Sonnet
- Gemini 2.5 Pro
Context lengths tested:
- 10K tokens
- 50K tokens
- 100K tokens
#2Overall results by depth
| Model | 10K tokens | 50K tokens | 100K tokens |
|---|---|---|---|
| GPT-4o | 100% | 96% | 89% |
| Claude 3.5 Sonnet | 100% | 98% | 93% |
| Gemini 2.5 Pro | 100% | 97% | 87% |
The big story is that all three models are strong at short contexts, but retrieval quality declines as the context grows. At 100K tokens, the gap is large enough to matter for production tasks.
#2Position matters more than people expect
The most revealing result is how accuracy shifts by location inside the context.
| Position in context | GPT-4o | Claude 3.5 | Gemini 2.5 |
|---|---|---|---|
| 0% (start) | 100% | 100% | 100% |
| 10% | 98% | 100% | 97% |
| 20% | 95% | 98% | 93% |
| 30% | 90% | 96% | 88% |
| 40% | 86% | 94% | 84% |
| 50% (middle) | 84% | 92% | 82% |
| 60% | 86% | 94% | 84% |
| 70% | 90% | 96% | 88% |
| 80% | 95% | 98% | 93% |
| 90% | 98% | 100% | 97% |
This is the classic “lost in the middle” effect: the model is strongest at the beginning and end of the context and weakest in the middle.
That matters because a lot of AI builders assume “the model can read everything in the context equally.” The data suggests it does not. It pays more attention to the beginning and end of the prompt than to the middle.
#2What this means for retrieval systems
If you build a system that injects several chunks into a long prompt, you should not just throw them in order and hope the model finds the right one.
A few practical guidelines:
- Put the highest-priority facts first and last. This matches the way the model attends to the context.
- Do not bury critical facts in the middle. Middle placement causes a meaningful miss rate at long lengths.
- Use chunk reordering by relevance. The best chunks should be placed at the edges of the context window.
- Split long documents into smaller reasoning units. A single 100K-token pass is often worse than a few targeted passes.
- Use summarization or intermediate steps for long tasks. Do not treat the whole document as one flat prompt unless you have tested for this problem in your workload.
#2Cost trade-offs
Large context is not just a capability problem. It is also an economics problem.
| Model | Input cost per 1M tokens | Cost per 100K-token test |
|---|---|---|
| GPT-4o | $2.50 | $0.25 |
| Claude 3.5 Sonnet | $3.00 | $0.30 |
| Gemini 2.5 Pro | $1.25 | $0.125 |
Gemini is the cheapest at scale, but it also showed the sharpest drop in middle-position retrieval. The trade-off is real: cheaper is not necessarily better when retrieval fidelity matters.
#2What this means for your architecture
A bigger context window makes long-range memory possible. It does not remove the need for good selection and ordering.
If your system passes a huge document string to a model, there is a real chance the most important fact gets lost in the middle. Retrieval quality, prompt ordering, and chunk summarization are still core architecture decisions.
Written by Rahul Jalavadiya, founder of AllDevToolsHub. All tools run locally in your browser.
#2Sources / Further reading
- Anthropic: Long context prompting guidance
- OpenAI: Best practices for long context usage
- Google AI: Context window and retrieval guidance
- Research: Lost in the middle: how language models use context
#2Related tools
Quick Summary
>- Long context windows are useful, but the models still show a strong “lost in the middle” effect. Retrieval ordering and chunk selection matter more than raw token count.
Tools Mentioned in This Article
LLM Model Comparison Reference
Compare specifications, context limits, benchmarks, and pricing across all launched frontier and open-weights LLMs.
LLM Token Counter
Estimate token count and API costs for OpenAI, Claude, and Gemini.
AI Prompt Cost Calculator
Compare API costs across major LLM providers.
AI Prompt Formatter
Format and optimize your instructions for AI models like ChatGPT and Claude.
Tools, tactics, and toughened-up tips, once a week
New tools, deep-dives on developer workflows, and the occasional gem we found this week. No spam, no tracking. Unsubscribe anytime.
Found an error or have feedback?
We correct errors quickly and document changes in our changelog. Report issues at support@alldevtoolshub.com.