Am I replaceable?Take the test

AI context windows compared

The same leaderboard the test places you on — in tokens, in words, and in pages.

ModelMakerTokens≈ Words≈ Pages
Llama 4 ScoutMeta10,000,0007,500,00025,000
Gemini 2.5 ProGoogle1,000,000750,0002,500
GPT-5OpenAI400,000300,0001,000
Claude Sonnet 4.5Anthropic200,000150,000500
GPT-4oOpenAI128,00096,000320
GPT-4 (2023)OpenAI8,1926,10020
GPT-2 (2019)OpenAI1,0247702.5
A typical person (this test)50–25040–2000.1–0.7

Advertised (nominal) windows as of September 2026; ≈ words uses 0.75 words per token, ≈ pages uses 300 words per page. Model numbers change often — this page is updated when the test's leaderboard is.

Advertised vs. effective

Accepting a million tokens is not the same as using them. Benchmarks such as needle-in-a-haystack and RULER hide facts in long texts and check whether the model still finds them; many models degrade well before their advertised limit, some at a tenth of it. When you compare yourself to the numbers above, remember they are the ceiling, not the working figure.

What "words" means for a person

For a model, the window is verbatim: every token is available. For a person it's retention: after reading a passage once, how much of it can you still answer questions about? Most people manage 40–200 words on a first read at normal speed. That sounds small next to the table — it is — but it is also the number that training moves, which the model column never will for you.

See where you land →

Related: What is a context window, in human terms?