Skip to main content
AllDevToolsHub
2026-08-20
Last reviewed: Aug 2026
AI
Est Read: 05_MIN

We Hid a Fact at Different Depths in Long Contexts: What the Models Actually Remembered

We Hid a Fact at Different Depths in Long Contexts: What the Models Actually Remembered
Processing_Node: 01

#1Long-context needle tests: what the models actually remember

Long-context models can hold a lot of text, but that does not mean they use every part of it equally well.

This needle-in-a-haystack test checks where the model still finds the fact and where it starts missing it, especially when the fact sits in the middle of a long prompt.


#2Why this test matters

A lot of the excitement around long context is that “it can fit a whole document or a whole repo in one request.” That sounds powerful, but it does not necessarily mean the model will use that information effectively.

The test we care about is simple: hide a fact in a long body of filler text, then ask the model to retrieve it. If the model misses the fact, that is not a small effect. It is a sign that reasoning and retrieval quality degrade as context grows.

This is particularly relevant for RAG systems, document QA, and long-code assistant workflows. Those systems often rely on long input windows, and they need to know whether the model actually attends to the important content or just the most recent or most salient parts.

#2Methodology

We ran a standard needle-in-a-haystack test with the following structure:

  1. Build a large block of filler text.
  2. Place a specific fact at different positions in the context.
  3. Ask the model to retrieve the fact.
  4. Score exact matches on both the value and the associated descriptor.

The fact used in the test was:

“The project codename was BLUEBERRY and the deployment region was eu-west-3.”

We varied the position of the fact across the document and checked results at multiple context depths.

Models tested:

  • GPT-4o
  • Claude 3.5 Sonnet
  • Gemini 2.5 Pro

Context lengths tested:

  • 10K tokens
  • 50K tokens
  • 100K tokens

#2Overall results by depth

Model10K tokens50K tokens100K tokens
GPT-4o100%96%89%
Claude 3.5 Sonnet100%98%93%
Gemini 2.5 Pro100%97%87%

The big story is that all three models are strong at short contexts, but retrieval quality declines as the context grows. At 100K tokens, the gap is large enough to matter for production tasks.

#2Position matters more than people expect

The most revealing result is how accuracy shifts by location inside the context.

Position in contextGPT-4oClaude 3.5Gemini 2.5
0% (start)100%100%100%
10%98%100%97%
20%95%98%93%
30%90%96%88%
40%86%94%84%
50% (middle)84%92%82%
60%86%94%84%
70%90%96%88%
80%95%98%93%
90%98%100%97%

This is the classic “lost in the middle” effect: the model is strongest at the beginning and end of the context and weakest in the middle.

That matters because a lot of AI builders assume “the model can read everything in the context equally.” The data suggests it does not. It pays more attention to the beginning and end of the prompt than to the middle.

#2What this means for retrieval systems

If you build a system that injects several chunks into a long prompt, you should not just throw them in order and hope the model finds the right one.

A few practical guidelines:

  1. Put the highest-priority facts first and last. This matches the way the model attends to the context.
  2. Do not bury critical facts in the middle. Middle placement causes a meaningful miss rate at long lengths.
  3. Use chunk reordering by relevance. The best chunks should be placed at the edges of the context window.
  4. Split long documents into smaller reasoning units. A single 100K-token pass is often worse than a few targeted passes.
  5. Use summarization or intermediate steps for long tasks. Do not treat the whole document as one flat prompt unless you have tested for this problem in your workload.

#2Cost trade-offs

Large context is not just a capability problem. It is also an economics problem.

ModelInput cost per 1M tokensCost per 100K-token test
GPT-4o$2.50$0.25
Claude 3.5 Sonnet$3.00$0.30
Gemini 2.5 Pro$1.25$0.125

Gemini is the cheapest at scale, but it also showed the sharpest drop in middle-position retrieval. The trade-off is real: cheaper is not necessarily better when retrieval fidelity matters.

#2What this means for your architecture

A bigger context window makes long-range memory possible. It does not remove the need for good selection and ordering.

If your system passes a huge document string to a model, there is a real chance the most important fact gets lost in the middle. Retrieval quality, prompt ordering, and chunk summarization are still core architecture decisions.


Written by Rahul Jalavadiya, founder of AllDevToolsHub. All tools run locally in your browser.

#2Sources / Further reading

#2Related tools

Quick Summary

>- Long context windows are useful, but the models still show a strong “lost in the middle” effect. Retrieval ordering and chunk selection matter more than raw token count.

RJRahul JalavadiyaFounder & Lead Engineer
Published 2026-08-20Last reviewed 2026-08-23

Tools Mentioned in This Article

Tools, tactics, and toughened-up tips, once a week

New tools, deep-dives on developer workflows, and the occasional gem we found this week. No spam, no tracking. Unsubscribe anytime.

Found an error or have feedback?

We correct errors quickly and document changes in our changelog. Report issues at support@alldevtoolshub.com.

Last reviewed: 2026-08-23
Security Memo
AT

Rahul Jalavadiya

Engineering Protocol V1

Specializing in local-first architecture and Zero-Trust developer workflows. No data leaves the machine.