- The short answer Exactly one model advertises a 10 million token context window: Meta's Llama 4 Scout. Every managed commercial API tops out around one million.Meta's own announcement says Scout was pre-trained and post-trained at 256K. The 10M figure is length generalization beyond the trained range, and the evidence Meta cites for it is needle-in-a-haystack retrieval plus perplexity over code, not reasoning benchmarks.Meta also says Scout fits on a single H100 with Int4 weights. Our arithmetic below shows those two claims cannot hold simultaneously: the leftover memory on an 80GB card holds a few hundred thousand tokens of KV cache, not ten million.If you need to reason over millions of tokens, retrieval is the answer, and it is roughly a hundred times cheaper.
Vendor limits read from Google, OpenAI, SpaceXAI and Meta documentation on 17 September 2026. The memory arithmetic is our own and shown in full so you can check it.
Which models actually have a 10M context window?
Start with what you can send today, rather than what a comparison table on a content farm says.

Every figure read from the vendor’s own model documentation this week.
Google’s documentation for Gemini 3.8 Flash gives an input token limit of 1,048,576 — exactly one mebi-token, and exactly one million in round numbers. OpenAI’s model card for GPT-6 Astra gives a 1,050,000 context window with a 922,000 maximum input. SpaceXAI’s largest is Grok 4.3 at 1M, and its current flagship Grok 4.6 is only 500,000.
Several articles currently ranking for this question list Gemini at 10M. As of this week Google’s own model pages do not. Gemini 3 Pro Preview, which some of those articles cite, is listed in Google’s documentation as shut down (see our full gpt-oss-120b vs Gemini 3 Pro comparison for the live numbers and benchmarks).
So there is one 10M model, it has open weights, and you have to serve it yourself. That single fact reframes the whole question, because the moment you are serving it yourself the memory cost stops being someone else’s problem.
What Meta actually claimed
It is worth quoting the announcement rather than the coverage of it, because the announcement is more careful than the headlines were.

Three things Meta stated plainly that most write-ups skipped.
The sentence that matters: Scout is “both pre-trained and post-trained with a 256K context length, which empowers the base model with advanced length generalization capability.” Meta is not hiding this. The model was trained at 256,000 tokens. The 10M number describes how far the architecture is expected to stretch past what it saw in training.
The architecture doing the stretching is what Meta calls iRoPE — interleaved attention layers without positional embeddings, plus inference-time temperature scaling of attention. Meta notes the “i” also gestures at the long-term goal of supporting infinite context length. That is an honest statement of direction, and it is not a claim that 10M works today.
What the cited evidence does and does not show
Meta supports the 10M figure with two things: retrieval needle-in-a-haystack results, and cumulative negative log-likelihoods over ten million tokens of code.
Needle-in-a-haystack plants one distinctive fact in a long document and asks the model to find it. Models are good at it, and have been for a while, because it is a search problem with a bright, unusual target. It tells you almost nothing about whether the model can hold two facts from opposite ends of the window and reason about their relationship.
Negative log-likelihood is a perplexity measure. Low NLL over 10M tokens of code means the model is not producing nonsense at that length. It does not mean the model is using what it read. A model can be unsurprised by text it cannot reason about.
Neither is a bad measurement. They are just not measurements of the thing the marketing number implies, and no published benchmark shows quality holding across a full 10M window.
The memory arithmetic nobody publishes
Here is where the claim comes apart, and it does so on numbers Meta supplied itself.

Our calculation. The inputs are Meta’s own two claims about the same model.
Meta says Scout has 109 billion total parameters and fits in a single NVIDIA H100 with Int4 quantization. At four bits per parameter that is 109e9 x 4 / 8 = 54.5 GB. An H100 has 80 GB. That leaves 25.5 GB for everything else, and everything else includes the KV cache.
The KV cache is the memory that holds the attention keys and values for every token in the context. Its size is set by a simple formula, and it grows linearly with context length:
bytes = 2 x layers x kv_heads x head_dim x context x bytes_per_element
↑
K and V
We do not have Scout’s exact layer and head configuration to hand, so the table above runs a range of plausible values rather than asserting one. Across every configuration in that range the conclusion is identical. At fp16 you fit roughly 130,000 tokens in the headroom. At aggressively quantized int4 KV you fit around 520,000. Neither is within an order of magnitude of ten million.
Run it the other way and the number gets vivid. Holding ten million tokens of KV cache needs somewhere between roughly 490 GB and 2 TB depending on precision. That is six to twenty-five H100s, dedicated to the cache, for a single request.
This is the same method we used in what actually fits in 32GB of RAM, where measuring real file sizes rather than trusting nominal numbers produced answers that surprised most readers. The arithmetic here is deliberately simple so you can substitute your own figures.
None of this means Meta lied. Both claims are true in isolation: the model does fit on one H100, and the architecture does accept a 10M positional range. They are simply not true at the same time, and a spec sheet that lists them side by side invites a reading that the hardware does not support.
How much text is ten million tokens anyway?

Rough conversions, to make the number concrete.
Seventy-five novels. A million lines of code. Roughly fifteen thousand pages. It is a genuinely startling amount of text and that is exactly why the number became a marketing asset.
But ask what you would actually do with it. Nobody needs a model to hold seventy-five novels simultaneously. They need it to answer a question that depends on four paragraphs scattered across seventy-five novels. Those are different problems, and only one of them requires an enormous context window.
What it would cost, if you could

Input tokens only, at published long-context rates. Output is on top.
Two hundred dollars of input tokens for one GPT-6 Astra request, except you cannot make that request, because the cap is 1.05M. Twenty-five dollars on the cheapest long-context option in the table, which also caps below 10M.
There is a second trap in those numbers. Both OpenAI and SpaceXAI reprice the entire request once the prompt crosses a long-context threshold — 272,000 tokens for GPT-6 Astra, 200,000 for Grok. It is a step, not a taper. The rates in that table are already the elevated ones, because any prompt in this territory has long since crossed the line.
What to do instead

Four approaches that work today and cost a fraction of a giant prompt.
Retrieval is the honest answer to almost every question that sounds like it needs 10M tokens. Index the corpus, fetch the twenty thousand tokens that actually bear on the question, and send those. The cost falls from tens of dollars to fractions of a cent, latency falls from minutes to seconds, and quality usually improves, because a model reasoning over twenty thousand relevant tokens beats one skimming ten million mostly-irrelevant ones.
Prompt caching is the second lever, and it is underused. Both major vendors discount cached input by 80 to 90 percent. If you are sending the same large prefix repeatedly — a codebase, a policy document, a schema — caching turns the expensive part into the cheap part.
If the motivation is privacy rather than capability, the answer is different again: running models locally removes the vendor from the loop entirely, and the Open WebUI guide covers the interface. You will be working with a smaller window, but as the arithmetic above shows, so is everyone else.
Frequently asked questions
Which model has a 10 million token context window?
Only Meta’s Llama 4 Scout advertises one. It has open weights, so you host it yourself. No managed commercial API currently offers 10M.
Does Gemini have a 10M context window?
No. Google’s documentation for Gemini 3.8 Flash lists an input limit of 1,048,576 tokens. Articles citing 10M for Gemini appear to refer to Gemini 3 Pro Preview, which Google lists as shut down.
What is the largest context window in a commercial API in 2026?
Roughly one million tokens. GPT-6 Astra is 1,050,000 with a 922,000 input maximum, Gemini 3.8 Flash is 1,048,576, and Grok 4.3 is 1,000,000.
Was Llama 4 Scout really trained on 10M tokens of context?
No. Meta’s announcement states it was pre-trained and post-trained at 256K. The 10M figure comes from length generalization via the iRoPE architecture.
Can Llama 4 Scout actually use all 10 million tokens?
No published benchmark shows quality holding across the full window. Meta’s cited evidence is needle-in-a-haystack retrieval and perplexity over code, neither of which measures reasoning across the context.
How much GPU memory does a 10M token context need?
Roughly 490 GB to 2 TB of KV cache alone, depending on precision — six to twenty-five H100s for a single request. The exact figure depends on the model’s layer and head configuration.
How much text is 10 million tokens?
About 7.5 million words: seventy-five average novels, or around a million lines of code.
What should I use instead of a huge context window?
Retrieval for almost everything, prompt caching when you resend the same prefix, and chunked summarisation when the corpus genuinely will not fit. All three are cheaper and usually more accurate.
Primary sources, read 17 September 2026: ai.meta.com/blog/llama-4-multimodal-intelligence (5 Apr 2025) · ai.google.dev/gemini-api/docs/models and the Gemini 3.8 Flash model page · developers.openai.com/api/docs/models/gpt-6-astra · docs.x.ai/developers/pricing. Memory and cost calculations are our own, from the parameters and rates those pages publish.
Leave a Reply