The same leaderboard the test places you on — in tokens, in words, and in pages.
| Model | Maker | Tokens | ≈ Words | ≈ Pages |
|---|---|---|---|---|
| Llama 4 Scout | Meta | 10,000,000 | 7,500,000 | 25,000 |
| Gemini 2.5 Pro | 1,000,000 | 750,000 | 2,500 | |
| GPT-5 | OpenAI | 400,000 | 300,000 | 1,000 |
| Claude Sonnet 4.5 | Anthropic | 200,000 | 150,000 | 500 |
| GPT-4o | OpenAI | 128,000 | 96,000 | 320 |
| GPT-4 (2023) | OpenAI | 8,192 | 6,100 | 20 |
| GPT-2 (2019) | OpenAI | 1,024 | 770 | 2.5 |
| A typical person (this test) | — | 50–250 | 40–200 | 0.1–0.7 |
Advertised (nominal) windows as of September 2026; ≈ words uses 0.75 words per token, ≈ pages uses 300 words per page. Model numbers change often — this page is updated when the test's leaderboard is.
Accepting a million tokens is not the same as using them. Benchmarks such as needle-in-a-haystack and RULER hide facts in long texts and check whether the model still finds them; many models degrade well before their advertised limit, some at a tenth of it. When you compare yourself to the numbers above, remember they are the ceiling, not the working figure.
For a model, the window is verbatim: every token is available. For a person it's retention: after reading a passage once, how much of it can you still answer questions about? Most people manage 40–200 words on a first read at normal speed. That sounds small next to the table — it is — but it is also the number that training moves, which the model column never will for you.