What is a Context Window?
The context window refers to the maximum amount of text that a Large Language Model can process at once. It includes both your input (prompt) and the model's output.
Think of the context window as the model's working memory: everything that fits inside can be "seen" and considered by the model. What's outside doesn't exist for the model.
Context Windows of Current Models
Here's an interactive overview of context windows for over 300 LLMs from Anthropic, Google, OpenAI, Meta, and more:
Model | Developer | Context Window |
|---|---|---|
| Meta | 10.5M | |
| Alibaba | 10M | |
2M | ||
2M | ||
| xAI | 2M | |
| xAI | 2M | |
GPT-6.1 Sol | OpenAI | 1.1M |
GPT-6 Astra | OpenAI | 1.1M |
GPT-6 Sol | OpenAI | 1.1M |
GPT-6 Luna | OpenAI | 1.1M |
GPT-5.6 Sol | OpenAI | 1.1M |
GPT-5.6 Terra | OpenAI | 1.1M |
GPT-5.6 Luna | OpenAI | 1.1M |
GPT-5.5 | OpenAI | 1.1M |
GPT-5.5 Pro | OpenAI | 1.1M |
GPT-5.4 | OpenAI | 1.1M |
GPT-5.4 Pro | OpenAI | 1.1M |
Llama 4 Maverick | Meta | 1M |
Gemini 3.8 Flash | 1M | |
Gemini 3.7 Flash | 1M | |
Gemini 3.6 Flash | 1M | |
Gemini 3.5 Flash | 1M | |
Gemini 3.1 Pro Preview | 1M | |
Gemini 3.5 Flash-Lite | 1M | |
Gemini 3.1 Flash-Lite | 1M | |
Gemini 3 Flash Preview | 1M | |
Gemini 2.5 Pro | 1M | |
Gemini 2.5 Flash | 1M | |
Gemini 2.5 Flash-Lite | 1M | |
Kimi K3 | Moonshot AI | 1M |
Muse Spark 1.3 | Meta | 1M |
| Meta | 1M | |
Laguna S 2.1 | Poolside | 1M |
GPT-4.1 | OpenAI | 1M |
GPT-4.1 mini | OpenAI | 1M |
GPT-4.1 nano | OpenAI | 1M |
1M | ||
1M | ||
1M | ||
Grok 4.3 | xAI | 1M |
Grok 4.20 Reasoning | xAI | 1M |
Grok 4.20 Multi-Agent | xAI | 1M |
Claude Opus 5.5 | Anthropic | 1M |
Claude Sonnet 5.5 | Anthropic | 1M |
Claude Fable 5.1 | Anthropic | 1M |
Claude Mythos 5.1 | Anthropic | 1M |
Claude Fable 5 | Anthropic | 1M |
Claude Mythos 5 | Anthropic | 1M |
Claude Sonnet 5 | Anthropic | 1M |
Claude Opus 5 | Anthropic | 1M |
Claude Opus 4.8 | Anthropic | 1M |
Claude Opus 4.7 | Anthropic | 1M |
Claude Opus 4.6 | Anthropic | 1M |
Claude Sonnet 4.6 | Anthropic | 1M |
DeepSeek-V4.1-Flash | DeepSeek | 1M |
DeepSeek-V4-Pro | DeepSeek | 1M |
DeepSeek-V4-Flash | DeepSeek | 1M |
MiniMax M3 | MiniMax | 1M |
Qwen 3.7 Max | Alibaba | 1M |
| Alibaba | 1M | |
| Alibaba | 1M | |
LongCat 2.0 | Meituan | 1M |
Inkling | Thinking Machines Lab | 1M |
Inkling Small | Thinking Machines Lab | 1M |
Qwen 3.7 Flash | Alibaba | 1M |
Qwen 3.8 Flash | Alibaba | 1M |
Qwen 3.8 Max 0902 | Alibaba | 1M |
Qwen 3.7 Plus | Alibaba | 1M |
GLM-5.3-Flash | Z.ai | 1M |
GLM-5.2 | Z.ai | 1M |
GLM-5.3 | Z.ai | 1M |
MiMo-V2.5 | Xiaomi | 1M |
MiMo-V2.5-Pro | Xiaomi | 1M |
MiMo-V2.5-Pro-UltraSpeed | Xiaomi | 1M |
MiniMax M1 | MiniMax | 1M |
Nemotron 3 Nano | NVIDIA | 1M |
Nemotron 3 Super | NVIDIA | 1M |
Nemotron 3 Ultra | NVIDIA | 1M |
Nemotron 3.5 Lightning | NVIDIA | 1M |
Amazon Nova Premier | Amazon | 1M |
Amazon Nova 2 Lite | Amazon | 1M |
Amazon Nova 2 Sonic | Amazon | 1M |
MiniMax-01 | MiniMax | 1M |
Solar Pro 4 | Upstage | 512K |
Grok 4.7 | xAI | 500K |
Grok 4.6 | xAI | 500K |
Grok 4.5 | xAI | 500K |
| OpenAI | 400K | |
ChatGPT chat-latest | OpenAI | 400K |
GPT-5.4 mini | OpenAI | 400K |
GPT-5.4 nano | OpenAI | 400K |
GPT-5.3-Codex | OpenAI | 400K |
| OpenAI | 400K | |
| OpenAI | 400K | |
| OpenAI | 400K | |
| OpenAI | 400K | |
| OpenAI | 400K | |
| OpenAI | 400K | |
GPT-5 | OpenAI | 400K |
GPT-5 Pro | OpenAI | 400K |
GPT-5 mini | OpenAI | 400K |
GPT-5 nano | OpenAI | 400K |
Amazon Nova Pro | Amazon | 300K |
Amazon Nova Lite | Amazon | 300K |
Mistral Large 3 | Mistral AI | 262.14K |
Mistral Medium 3.5 | Mistral AI | 262.14K |
Kimi K2.6 | Moonshot AI | 262.14K |
Kimi K2.7 Code | Moonshot AI | 262.14K |
| Alibaba | 262.14K | |
| Alibaba | 262.14K | |
Laguna XS 2.1 | Poolside | 262.14K |
Nex-N2-mini | Nex AGI | 262.14K |
Nex-N2-Pro | Nex AGI | 262.14K |
Hy3 | Tencent | 262.14K |
Hunyuan A13B Instruct | Tencent | 262.14K |
Ring 2.6 1T | inclusionAI | 262.14K |
Ling 3.0 Flash | inclusionAI | 262.14K |
Trinity Large Thinking | Arcee AI | 262.14K |
Gemma 4 31B | 262.14K | |
Gemma 4 26B A4B | 262.14K | |
Ministral 3 14B | Mistral AI | 262.14K |
Ministral 3 8B | Mistral AI | 262.14K |
Ministral 3 3B | Mistral AI | 262.14K |
Qwen 3.8 2.4T A95B | Alibaba | 262.14K |
Qwen 3.8 27B | Alibaba | 262.14K |
Qwen 3.6 35B A3B | Alibaba | 262.14K |
Qwen 3.5 397B A17B | Alibaba | 262.14K |
Qwen 3.5 122B A10B | Alibaba | 262.14K |
Qwen 3.5 35B A3B | Alibaba | 262.14K |
Qwen 3.5 27B | Alibaba | 262.14K |
Qwen 3.5 9B | Alibaba | 262.14K |
Qwen 3 235B A22B Instruct 2507 | Alibaba | 262.14K |
Qwen 3 30B A3B Instruct 2507 | Alibaba | 262.14K |
Qwen 3 Coder 480B A35B | Alibaba | 262.14K |
Qwen 3 Coder Next | Alibaba | 262.14K |
Qwen 3.6 27B | Alibaba | 262.14K |
Qwen 3 Coder 30B A3B | Alibaba | 262.14K |
Qwen 3 Next 80B A3B Instruct | Alibaba | 262.14K |
Qwen 3 Next 80B A3B Thinking | Alibaba | 262.14K |
Qwen 3 235B A22B Thinking 2507 | Alibaba | 262.14K |
Qwen 3 30B A3B Thinking 2507 | Alibaba | 262.14K |
Qwen 3 VL 235B A22B | Alibaba | 262.14K |
Qwen 3 VL 30B A3B | Alibaba | 262.14K |
Qwen 3 VL 8B | Alibaba | 262.14K |
Qwen 3 VL 235B A22B Thinking | Alibaba | 262.14K |
Qwen 3 VL 30B A3B Thinking | Alibaba | 262.14K |
Qwen 3 VL 32B | Alibaba | 262.14K |
Qwen 3 VL 8B Thinking | Alibaba | 262.14K |
Kimi K2.5 | Moonshot AI | 262.14K |
Kimi K2 Thinking | Moonshot AI | 262.14K |
Kimi K2 0905 | Moonshot AI | 262.14K |
Grok Build 0.1 | xAI | 256K |
| xAI | 256K | |
| xAI | 256K | |
Mistral Small 4 | Mistral AI | 256K |
| Mistral AI | 256K | |
| Alibaba | 256K | |
Seed 1.8 | ByteDance | 256K |
Seed 2.0 Pro | ByteDance | 256K |
KAT-Coder-Air V2.5 | Kuaishou | 256K |
KAT-Coder-Pro V2.5 | Kuaishou | 256K |
Seed 1.6 | ByteDance | 256K |
Seed 1.6 Flash | ByteDance | 256K |
Seed 2.0 Lite | ByteDance | 256K |
Seed 2.0 Mini | ByteDance | 256K |
Seed 2.0 Code | ByteDance | 256K |
Seed 2.1 Turbo | ByteDance | 256K |
Step 3.5 Flash | StepFun | 256K |
Step 3.7 Flash | StepFun | 256K |
Command A | Cohere | 256K |
Command A Reasoning | Cohere | 256K |
| AI21 Labs | 256K | |
| AI21 Labs | 256K | |
| AI21 Labs | 256K | |
| MiniMax | 245.76K | |
MiniMax M2.7 | MiniMax | 204.8K |
MiniMax M2.5 | MiniMax | 204.8K |
MiniMax M2.1 | MiniMax | 204.8K |
MiniMax M2 | MiniMax | 204.8K |
GLM-4.7 | Z.ai | 204.8K |
GLM-4.6 | Z.ai | 204.8K |
Claude Opus 4.5 | Anthropic | 200K |
Claude Sonnet 4.5 | Anthropic | 200K |
Claude Haiku 4.5 | Anthropic | 200K |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
| Anthropic | 200K | |
o3 | OpenAI | 200K |
o3-pro | OpenAI | 200K |
o4-mini | OpenAI | 200K |
o3-mini | OpenAI | 200K |
o1 | OpenAI | 200K |
GLM-5.1 | Z.ai | 200K |
GLM-5 | Z.ai | 200K |
Sonar Pro | Perplexity | 200K |
GLM-4.7-Flash | Z.ai | 200K |
| 01.AI | 200K | |
| 01.AI | 200K | |
DeepSeek-V3.1 Terminus | DeepSeek | 163.84K |
DeepSeek-V3 0324 | DeepSeek | 163.84K |
DeepSeek-R1 0528 | DeepSeek | 163.84K |
Llama 3.3 70B | Meta | 131.07K |
Llama 3.2 3B | Meta | 131.07K |
Llama 3.2 1B | Meta | 131.07K |
Llama 3.1 70B | Meta | 131.07K |
Llama 3.1 8B | Meta | 131.07K |
| xAI | 131.07K | |
Mistral Nemo | Mistral AI | 131.07K |
Qwen 2.5 72B | Alibaba | 131.07K |
Qwen 2.5 7B | Alibaba | 131.07K |
Qwen 2.5 Coder 32B | Alibaba | 131.07K |
Muse Glimmer 30B | Meta | 131.07K |
Solar Pro 3 | Upstage | 131.07K |
ERNIE 4.5 VL 424B A47B | Baidu | 131.07K |
Granite 4.1 8B | IBM | 131.07K |
Granite 4.0 H Micro | IBM | 131.07K |
Virtuoso Large | Arcee AI | 131.07K |
Hermes 4 70B | Nous Research | 131.07K |
Hermes 4 405B | Nous Research | 131.07K |
GPT OSS 120B | OpenAI | 131.07K |
GPT OSS 20B | OpenAI | 131.07K |
Mistral Small 3.2 | Mistral AI | 131.07K |
Mistral Small 3.1 | Mistral AI | 131.07K |
Kimi K2 0711 | Moonshot AI | 131.07K |
GLM-4.6V | Z.ai | 131.07K |
GLM-4.5 | Z.ai | 131.07K |
GLM-4.5 Air | Z.ai | 131.07K |
| Meta | 128K | |
| Meta | 128K | |
| Meta | 128K | |
Gemma 3 27B | 128K | |
Gemma 3 12B | 128K | |
Gemma 3 4B | 128K | |
| xAI | 128K | |
| OpenAI | 128K | |
| OpenAI | 128K | |
GPT-4o | OpenAI | 128K |
GPT-4o mini | OpenAI | 128K |
| OpenAI | 128K | |
DeepSeek-V3.1 | DeepSeek | 128K |
DeepSeek-V3 | DeepSeek | 128K |
DeepSeek-R1 | DeepSeek | 128K |
DeepSeek-R1-Distill-Llama-70B | DeepSeek | 128K |
DeepSeek R1 Distill Qwen 32B | DeepSeek | 128K |
| DeepSeek | 128K | |
| DeepSeek | 128K | |
| DeepSeek | 128K | |
| DeepSeek | 128K | |
DeepSeek Coder V2 | DeepSeek | 128K |
| Mistral AI | 128K | |
Ministral 8B | Mistral AI | 128K |
| Mistral AI | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
| Alibaba | 128K | |
Nemotron Nano 2 9B | NVIDIA | 128K |
Nemotron Nano 2 VL 12B | NVIDIA | 128K |
Command R7B | Cohere | 128K |
Sonar | Perplexity | 128K |
Sonar Reasoning Pro | Perplexity | 128K |
Sonar Deep Research | Perplexity | 128K |
DeepSeek-V3.2 | DeepSeek | 128K |
DeepSeek-V3.2 Exp | DeepSeek | 128K |
Qwen 2.5 VL 72B | Alibaba | 128K |
Command R+ | Cohere | 128K |
Command R | Cohere | 128K |
Amazon Nova Micro | Amazon | 128K |
Phi-4-mini | Microsoft | 128K |
| Microsoft | 128K | |
| Microsoft | 128K | |
| Microsoft | 128K | |
| Microsoft | 128K | |
| Microsoft | 128K | |
| 01.AI | 128K | |
| 01.AI | 128K | |
| Nvidia | 128K | |
| Nvidia | 128K | |
| Nvidia | 128K | |
| Reka | 128K | |
| Reka | 128K | |
| Reka | 128K | |
| Zhipu AI | 128K | |
| Zhipu AI | 128K | |
| Baidu | 128K | |
Mixtral 8x22B | Mistral AI | 65.54K |
Solar Pro 2 | Upstage | 65.54K |
GLM-4.5V | Z.ai | 65.54K |
| Microsoft | 64K | |
Mistral Small 3 | Mistral AI | 32.77K |
Mixtral 8x7B | Mistral AI | 32.77K |
| Mistral AI | 32.77K | |
| Alibaba | 32.77K | |
| Alibaba | 32.77K | |
| Alibaba | 32.77K | |
Solar Mini | Upstage | 32.77K |
32.77K | ||
Qwen 3 32B | Alibaba | 32.77K |
Qwen 3 14B | Alibaba | 32.77K |
Qwen 3 8B | Alibaba | 32.77K |
Qwen 3 235B A22B | Alibaba | 32.77K |
Qwen 3 30B A3B | Alibaba | 32.77K |
| Microsoft | 32.77K | |
DBRX | Databricks | 32.77K |
32K | ||
Voxtral Small 24B | Mistral AI | 32K |
| 01.AI | 32K | |
Phi-4 | Microsoft | 16.38K |
| 01.AI | 16K | |
Gemma 2 27B | 8.19K | |
8.19K | ||
| OpenAI | 8.19K | |
| AI21 Labs | 8.19K | |
| Zhipu AI | 8.19K | |
| Baidu | 8K | |
| Cohere | 4.1K | |
| Nvidia | 4.1K | |
| Stability AI | 4.1K | |
| Stability AI | 4.1K |
Context window sizes of current AI language models (as of August 2026)
The table clearly shows the rapid progress: While early models like GPT-3.5 could only process 4,000 to 16,000 tokens, current models like Llama 4 Scout already reach 10 million tokens. That's equivalent to about 30 Harry Potter books or 25,000 book pages.
What Are Tokens?
Tokens are the basic units into which text is broken down for LLMs. A token isn't always a whole word. Common words are often one token, while rare words are split into multiple tokens.
Rule of thumb for English: 1 token ≈ 0.75 words. A typical blog post with 1,000 words requires about 1,300 tokens.
Why is the Context Window Important?
For Conversations
The model "forgets" earlier parts of a long conversation when they no longer fit in the context window. That's why chatbots can lose track in very long conversations.
For Document Analysis
A larger context window enables analysis of longer documents. With Gemini 3.1 Pro, Claude Opus 5.5, Claude Opus 5, or Claude Sonnet 5.5, you can analyze entire books at once. GPT-6 Sol and GPT-6 Luna go slightly further, at 1,050,000 tokens. With older models, you have to split texts.
For Code Assistants
AI code assistants like Claude Code benefit from large context windows as they can "see" and understand more files simultaneously.
Strategies for Limited Context
- Summarizing: Summarize long texts before the prompt
- Chunking: Split documents into sections and process individually
- RAG: Retrieve relevant passages via vector search instead of inserting everything
- Conversation Reset: Repeat important info in long chats
Lost in the Middle
Studies show that LLMs process information at the beginning and end of the context window better than in the middle. This phenomenon is called "Lost in the Middle." Important information should therefore be placed at the beginning or end of your prompt.
Cost Aspect
When using APIs, you pay per token (for both input and output). Using a long context window is therefore more expensive. With Claude Sonnet 5.5, processing 100,000 input tokens costs about $0.20.
Conclusion
The context window is one of the most important limitations of modern LLMs. With models like Claude Opus 5.5, Gemini 3.1 Pro, and Llama 4 Scout that can process millions of tokens, many previous workarounds become unnecessary. Still, it remains important to design prompts efficiently, both for cost reasons and because of the "Lost in the Middle" effect.
