What is a Large Language Model (LLM)?
A Large Language Model (LLM) is an artificial neural network trained on massive amounts of text data to understand and generate human language. LLMs like GPT-4, Claude, Gemini, and LLaMA can write texts, answer questions, write code, and solve complex tasks.
The term "Large" refers to the number of parameters. Modern LLMs have hundreds of billions of parameters that are optimized during training. The more parameters, the more complex patterns the model can capture.
How Do LLMs Work?
LLMs are based on the Transformer architecture, introduced by Google in 2017. The core is the "attention mechanism," which allows the model to recognize relevant relationships in text, even across large distances.
Training in Three Phases
- Pre-Training: The model learns from billions of texts (books, websites, Wikipedia) to predict the next words. This develops a deep understanding of language.
- Fine-Tuning: The model is adapted to specific tasks or formats, such as following instructions or answering questions in a dialogue format.
- RLHF (Reinforcement Learning from Human Feedback): Humans rate the model's responses, and it learns to prioritize helpful, harmless, and honest answers.
Popular LLMs Overview
GPT-5 (OpenAI)
The GPT series (Generative Pre-trained Transformer) from OpenAI is the most well-known LLM. ChatGPT is based on these models. GPT-5.5 (April 2026) introduced native multimodal inputs and a context window of up to 1 million tokens. On July 9, 2026 OpenAI made the GPT-5.6 family (Sol, Terra, Luna) generally available across ChatGPT, ChatGPT Work, Codex, and the API. On September 22, 2026, GPT-6 Sol and GPT-6 Luna followed as cheaper successors to the Sol and Luna tiers, each with a 1,050,000-token context window, 128,000 output tokens, and a knowledge cutoff of April 20, 2026 (Sol) or May 18, 2026 (Luna). OpenAI's best model overall remains GPT-6 Astra.
Claude (Anthropic)
Claude is known for particularly long context windows and a focus on safety through "Constitutional AI." The current family includes Claude Fable 5.1, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, and Claude Haiku 4.5. Fable 5.1 offers a 1 million token context window, up to 128,000 output tokens, and a June 2026 knowledge cutoff. Claude Opus 5.5 replaced Opus 5 as Anthropic's new leading model on September 22, 2026, also with a 1 million token context window and a June 2026 knowledge cutoff, but at a lower price. Opus 5 remains available in parallel. Claude Sonnet 5.5 joined on September 28, 2026 as a faster, lower-cost complement to Opus 5.5, also with a 1 million token context window and a June 2026 knowledge cutoff. It costs half as much as Opus 5.5, and Sonnet 5 remains available in parallel. Claude Mythos 5.1 uses the same model as Fable 5.1 with fewer restrictive safeguards, but remains limited to vetted Project Glasswing partners.
Gemini (Google)
Google's LLM family ranges from Gemini Nano for mobile devices to Gemini 3.1 Pro for complex reasoning and Gemini 3.8 Flash for fast coding and agentic workflows. Gemini 3.8 Flash Cyber is a separate defensive-cybersecurity variant for vetted defenders in the Fairwind program. Google has not published a separate public API price or context-window value for it. The models are natively multimodal and can process text, images, audio, and video.
Muse and Llama (Meta)
Muse Spark 1.3 is Meta's current proprietary multimodal model for long-running agentic and coding tasks. It supports a 1,048,576-token context window. Meta also publishes open-weight LLMs such as the Llama family and Muse Glimmer.
Qwen (Alibaba)
Qwen 3.8 Max 0902 is Alibaba's September 2, 2026 snapshot for coding, long-horizon autonomous tasks, agent collaboration, and visual analysis. It combines 2.4 trillion total parameters, 95 billion active parameters, and a 1 million-token context window. Qwen 3.8 Flash is a multimodal API model with a native 1-million-token context and global list prices of $0.113 for input and $0.382 for output per 1 million tokens. The open-weight Qwen 3.8 Flash Next combines 125 billion main-model parameters with 51 billion N-gram embedding parameters, activates 6 billion parameters per token, and natively supports 262,144 tokens. YaRN can extend its context window to 1 million tokens. QwenCloud uses the API name qwen3.8-flash for this variant.
Applications of LLMs
- Text Generation: Blog posts, emails, marketing copy
- Programming: Code generation, debugging, code reviews
- Customer Service: Chatbots and automated responses
- Translation: High-quality translations into dozens of languages
- Research: Summarizing documents and extracting facts
- Education: Personalized tutoring and explanations
Limitations and Challenges
Hallucinations
LLMs can generate convincing-sounding but factually incorrect information. They sometimes "invent" facts, quotes, or sources. Therefore, critical review of outputs is important.
Knowledge Cutoff
LLMs have a knowledge cutoff date, meaning they only know information up to a certain point in time. Current events are unknown to them unless they have access to external tools like web search.
Context Window Limitation
Although modern LLMs have large context windows, the amount of text they can process simultaneously is limited. With very long documents, the quality of responses may decrease.
Bias and Fairness
LLMs reflect the biases in their training data. Despite intensive efforts toward fairness, they can reproduce stereotypical or discriminatory patterns.
Using LLMs Effectively
To get the most out of LLMs, good prompts are crucial. Techniques like Chain-of-Thought Prompting can significantly improve the quality of responses.
For developers, APIs from OpenAI, Anthropic, and Google offer the ability to integrate LLMs into their own applications. Costs are typically calculated based on tokens consumed.
Comprehensive LLM Parameter Menu
The following interactive table shows over 60 well-known Large Language Models with their parameter counts. You can search by name, filter by developer, size category or model type, and sort the columns:
Legend:
Showing 289 models
Model | Developer | Parameters |
|---|---|---|
MiniMax M2.7 MoE | MiniMax | Unknown |
MiniMax M2.5 MoE | MiniMax | Unknown |
GLM-4.7 MoE | Z.ai | Unknown |
GPT-6.1 Sol | OpenAI | Unknown |
GPT-6 Astra | OpenAI | Unknown |
GPT-6 Sol | OpenAI | Unknown |
GPT-6 Luna | OpenAI | Unknown |
GPT-5.6 Sol | OpenAI | Unknown |
GPT-5.6 Terra | OpenAI | Unknown |
GPT-5.6 Luna | OpenAI | Unknown |
GPT-5.5 | OpenAI | Unknown |
GPT-5.5 Pro | OpenAI | Unknown |
GPT-5.5 Instant | OpenAI | Unknown |
ChatGPT chat-latest | OpenAI | Unknown |
GPT-5.4 | OpenAI | Unknown |
GPT-5.4 Pro | OpenAI | Unknown |
GPT-5.4 mini | OpenAI | Unknown |
GPT-5.4 nano | OpenAI | Unknown |
GPT-5.3-Codex | OpenAI | Unknown |
GPT-5.3 Instant | OpenAI | Unknown |
GPT-5.2 | OpenAI | Unknown |
GPT-5.1 Instant | OpenAI | Unknown |
GPT-5.1 Thinking | OpenAI | Unknown |
GPT-5 | OpenAI | Unknown |
GPT-5 Pro | OpenAI | Unknown |
GPT-5 mini | OpenAI | Unknown |
GPT-5 nano | OpenAI | Unknown |
GPT-4 Turbo | OpenAI | Unknown |
GPT-4.1 | OpenAI | Unknown |
GPT-4.1 mini | OpenAI | Unknown |
GPT-4.1 nano | OpenAI | Unknown |
GPT-3.5 Turbo | OpenAI | Unknown |
o3 | OpenAI | Unknown |
o3-pro | OpenAI | Unknown |
o3-mini | OpenAI | Unknown |
o4-mini | OpenAI | Unknown |
o1 | OpenAI | Unknown |
o1-mini | OpenAI | Unknown |
Claude Opus 5.5 | Anthropic | Unknown |
Claude Sonnet 5.5 | Anthropic | Unknown |
Claude Fable 5.1 | Anthropic | Unknown |
Claude Mythos 5.1 | Anthropic | Unknown |
Claude Fable 5 | Anthropic | Unknown |
Claude Mythos 5 | Anthropic | Unknown |
Claude Sonnet 5 | Anthropic | Unknown |
Claude Opus 5 | Anthropic | Unknown |
Claude Opus 4.8 | Anthropic | Unknown |
Claude Opus 4.7 | Anthropic | Unknown |
Claude Opus 4.6 | Anthropic | Unknown |
Claude Sonnet 4.6 | Anthropic | Unknown |
Claude Opus 4.5 | Anthropic | Unknown |
Claude Opus 4.1 | Anthropic | Unknown |
Claude Sonnet 4.5 | Anthropic | Unknown |
Claude Haiku 4.5 | Anthropic | Unknown |
Claude Sonnet 4 | Anthropic | Unknown |
Claude Opus 4 | Anthropic | Unknown |
Claude Sonnet 3.7 | Anthropic | Unknown |
Claude 3.5 Haiku | Anthropic | Unknown |
Gemini 3.8 Flash | Unknown | |
Gemini 3.8 Flash Cyber | Unknown | |
Gemini 3.7 Flash | Unknown | |
Gemini 3.6 Flash | Unknown | |
Gemini 3.5 Flash | Unknown | |
Gemini 3.1 Pro Preview | Unknown | |
Gemini 3 Flash Preview MoE | Unknown | |
Gemini 3.5 Flash-Lite | Unknown | |
Gemini 3.1 Flash-Lite MoE | Unknown | |
Gemini 2.5 Pro MoE | Unknown | |
Gemini 2.5 Flash MoE | Unknown | |
Gemini 2.5 Flash-Lite MoE | Unknown | |
Gemini 3 Pro MoE | Unknown | |
Gemini 2.0 Flash MoE | Unknown | |
Gemini 1.5 Pro MoE | Unknown | |
Grok 4.7 | xAI | Unknown |
Grok 4.6 | xAI | Unknown |
Grok 4.5 | xAI | Unknown |
Grok 4.3 | xAI | Unknown |
Grok 4.20 Reasoning | xAI | Unknown |
Grok 4.20 Multi-Agent | xAI | Unknown |
Grok Build 0.1 | xAI | Unknown |
Grok 4 | xAI | Unknown |
Grok 3 | xAI | Unknown |
Grok 2 | xAI | Unknown |
Qwen 3.7 Max MoE | Alibaba | Unknown |
Seed 1.8 | ByteDance | Unknown |
Seed 2.0 Pro | ByteDance | Unknown |
Muse Spark 1.3 | Meta | Unknown |
Muse Spark 1.2 | Meta | Unknown |
KAT-Coder-Air V2.5 | Kuaishou | Unknown |
KAT-Coder-Pro V2.5 | Kuaishou | Unknown |
Seed 1.6 Flash | ByteDance | Unknown |
Seed 2.0 Lite | ByteDance | Unknown |
Seed 2.0 Mini | ByteDance | Unknown |
Seed 2.0 Code | ByteDance | Unknown |
Seed 2.1 Turbo | ByteDance | Unknown |
Qwen 3.7 Flash | Alibaba | Unknown |
Qwen 3.7 Plus MoE | Alibaba | Unknown |
Amazon Nova 2 Lite | Amazon | Unknown |
Amazon Nova Premier | Amazon | Unknown |
Amazon Nova Pro | Amazon | Unknown |
Amazon Nova Lite | Amazon | Unknown |
Amazon Nova Micro | Amazon | Unknown |
Sonar | Perplexity | Unknown |
Sonar Pro | Perplexity | Unknown |
Sonar Reasoning Pro | Perplexity | Unknown |
Sonar Deep Research | Perplexity | Unknown |
Solar Pro 4 | Upstage | Unknown |
Kimi K3 MoE(104B active) | Moonshot AI | 2.8T |
Qwen 3.8 2.4T A95B MoE(95B active) | Alibaba | 2.4T |
Qwen 3.8 Max 0902 MoE | Alibaba | 2.4T |
Claude 3 Opus | Anthropic | 2T* |
Llama 4 Behemoth MoE(288B active) | Meta | 2T |
GPT-4 MoE(220B active) | OpenAI | 1.76T* |
DeepSeek-V4-Pro MoE(49B active) | DeepSeek | 1.6T |
LongCat 2.0 MoE(48B active) | Meituan | 1.6T |
Ring 2.6 1T MoE(63B active) | inclusionAI | 1T |
Kimi K2.5 MoE(32B active) | Moonshot AI | 1T |
Kimi K2 Thinking MoE(32B active) | Moonshot AI | 1T |
Kimi K2 0711 MoE(32B active) | Moonshot AI | 1T |
Kimi K2 0905 MoE(32B active) | Moonshot AI | 1T |
Kimi K2.6 MoE(32B active) | Moonshot AI | 1T |
Kimi K2.7 Code MoE(32B active) | Moonshot AI | 1T |
Qwen 3.6 Max-Preview MoE | Alibaba | 1T* |
Yi-Large MoE | 01.AI | 1T |
MiMo-V2.5-Pro MoE(42B active) | Xiaomi | 1T |
MiMo-V2.5-Pro-UltraSpeed MoE(42B active) | Xiaomi | 1T |
Inkling MoE(41B active) | Thinking Machines Lab | 975B |
GLM-5.2 MoE(40B active) | Z.ai | 744B |
GLM-5.3 MoE(40B active) | Z.ai | 744B |
GLM-5.1 MoE(40B active) | Z.ai | 744B |
GLM-5 MoE(40B active) | Z.ai | 744B |
DeepSeek-V3 0324 MoE(37B active) | DeepSeek | 685B |
Mistral Large 3 MoE(41B active) | Mistral AI | 675B |
DeepSeek-V3.1 Terminus MoE(37B active) | DeepSeek | 671B |
DeepSeek-V3.2 Exp MoE(37B active) | DeepSeek | 671B |
DeepSeek-V3.1 MoE(37B active) | DeepSeek | 671B |
DeepSeek-R1 0528 MoE(37B active) | DeepSeek | 671B |
DeepSeek-V3.2 MoE(37B active) | DeepSeek | 671B |
DeepSeek-V3 MoE(37B active) | DeepSeek | 671B |
DeepSeek-R1 MoE(37B active) | DeepSeek | 671B |
DeepSeek-V4.1-Flash MoE | DeepSeek | 552B |
Nemotron 3 Ultra MoE(55B active) | NVIDIA | 550B |
PaLM | 540B | |
Megatron-Turing NLG | NVIDIA | 530B |
Qwen 3 Coder 480B A35B MoE(35B active) | Alibaba | 480B |
MiniMax M1 MoE(45.9B active) | MiniMax | 456B |
MiniMax-01 MoE(45.9B active) | MiniMax | 456B |
MiniMax M3 MoE(23B active) | MiniMax | 428B |
ERNIE 4.5 VL 424B A47B MoE(47B active) | Baidu | 424B |
Hermes 4 405B | Nous Research | 405B |
Llama 3.1 405B | Meta | 405B |
Llama 4 Maverick MoE(17B active) | Meta | 400B |
Trinity Large Thinking MoE(13B active) | Arcee AI | 398B |
Nex-N2-Pro MoE(17B active) | Nex AGI | 397B |
Qwen 3.5 397B A17B MoE(17B active) | Alibaba | 397B |
GLM-4.6 MoE(32B active) | Z.ai | 355B |
GLM-4.5 MoE(32B active) | Z.ai | 355B |
Nemotron-4 340B | NVIDIA | 340B |
PaLM 2 | 340B* | |
GLM-5.3-Flash MoE(18B active) | Z.ai | 320B |
Grok 1 MoE(86B active) | xAI | 314B |
MiMo-V2.5 MoE(15B active) | Xiaomi | 310B |
Hy3 MoE(21B active) | Tencent | 295B |
DeepSeek-V4-Flash MoE(13B active) | DeepSeek | 284B |
Inkling Small MoE(12B active) | Thinking Machines Lab | 276B |
DeepSeek-V2 MoE(21B active) | DeepSeek | 236B |
Qwen 3 235B A22B Instruct 2507 MoE(22B active) | Alibaba | 235B |
Qwen 3 235B A22B MoE(22B active) | Alibaba | 235B |
Qwen 3 235B A22B Thinking 2507 MoE(22B active) | Alibaba | 235B |
Qwen 3 VL 235B A22B MoE(22B active) | Alibaba | 235B |
Qwen 3 VL 235B A22B Thinking MoE(22B active) | Alibaba | 235B |
MiniMax M2.1 MoE(10B active) | MiniMax | 230B |
MiniMax M2 MoE(10B active) | MiniMax | 230B |
Seed 1.6 MoE(23B active) | ByteDance | 230B |
GPT-4o | OpenAI | 200B* |
Step 3.7 Flash MoE(11B active) | StepFun | 198B |
Step 3.5 Flash MoE(11B active) | StepFun | 196.8B |
Falcon 180B | TII | 180B |
BLOOM | BigScience | 176B |
GPT-3 | OpenAI | 175B |
Claude 3.5 Sonnet | Anthropic | 175B* |
OPT-175B | Meta | 175B |
Mixtral 8x22B MoE(39B active) | Mistral AI | 141B |
LaMDA | 137B | |
DBRX MoE(36B active) | Databricks | 132B |
Mistral Medium 3.5 | Mistral AI | 128B |
Qwen 3.8 Flash MoE(6B active) | Alibaba | 125B |
Ling 3.0 Flash MoE(5.1B active) | inclusionAI | 124B |
Mistral Large 2 | Mistral AI | 123B |
Qwen 3.5 122B A10B MoE(10B active) | Alibaba | 122B |
Nemotron 3 Super MoE(12B active) | NVIDIA | 120B |
Mistral Small 4 MoE(6.5B active) | Mistral AI | 119B |
Laguna S 2.1 MoE(8B active) | Poolside | 118B |
GPT OSS 120B MoE(5.1B active) | OpenAI | 117B |
Command A | Cohere | 111B |
Llama 4 Scout MoE(17B active) | Meta | 109B |
GLM-4.6V MoE(12B active) | Z.ai | 106B |
GLM-4.5 Air MoE(12B active) | Z.ai | 106B |
GLM-4.5V MoE(12B active) | Z.ai | 106B |
Command R+ | Cohere | 104B |
Solar Pro 3 MoE(12B active) | Upstage | 102B |
Hunyuan A13B Instruct MoE(13B active) | Tencent | 80B |
Qwen 3 Coder Next MoE(3B active) | Alibaba | 80B |
Qwen 3 Next 80B A3B Instruct MoE(3B active) | Alibaba | 80B |
Qwen 3 Next 80B A3B Thinking MoE(3B active) | Alibaba | 80B |
Qwen 2.5 72B | Alibaba | 72.7B |
Virtuoso Large | Arcee AI | 72B |
Qwen 2.5 VL 72B | Alibaba | 72B |
Hermes 4 70B | Nous Research | 70B |
DeepSeek-R1-Distill-Llama-70B | DeepSeek | 70B |
Claude 3 Sonnet | Anthropic | 70B* |
Llama 3.3 70B | Meta | 70B |
Llama 3.1 70B | Meta | 70B |
Llama 3 70B | Meta | 70B |
Llama 2 70B | Meta | 70B |
Mixtral 8x7B MoE(12.9B active) | Mistral AI | 46.7B |
Falcon 40B | TII | 40B |
Nex-N2-mini MoE(3B active) | Nex AGI | 35B |
Qwen 3.6 35B A3B MoE(3B active) | Alibaba | 35B |
Qwen 3.5 35B A3B MoE(3B active) | Alibaba | 35B |
Yi-34B | 01.AI | 34B |
Laguna XS 2.1 MoE(3B active) | Poolside | 33B |
Qwen 3 32B | Alibaba | 32.8B |
Qwen 3 VL 32B | Alibaba | 32.8B |
Qwen 2.5 Coder 32B | Alibaba | 32.5B |
Qwen 2.5 32B | Alibaba | 32B |
Command R | Cohere | 32B |
Solar Pro 2 | Upstage | 31B |
Gemma 4 31B | 30.7B | |
Qwen 3 30B A3B Instruct 2507 MoE(3.3B active) | Alibaba | 30.5B |
Qwen 3 Coder 30B A3B MoE(3.3B active) | Alibaba | 30.5B |
Qwen 3 30B A3B MoE(3.3B active) | Alibaba | 30.5B |
Qwen 3 30B A3B Thinking 2507 MoE(3.3B active) | Alibaba | 30.5B |
Qwen 3 VL 30B A3B MoE(3.3B active) | Alibaba | 30.5B |
Qwen 3 VL 30B A3B Thinking MoE(3.3B active) | Alibaba | 30.5B |
Nemotron 3 Nano MoE(3.5B active) | NVIDIA | 30B |
Nemotron 3.5 Lightning MoE(3B active) | NVIDIA | 30B |
GLM-4.7-Flash MoE(3B active) | Z.ai | 30B |
Muse Glimmer 30B | Meta | 29.6B |
Gemma 3 27B | 27B | |
Qwen 3.8 27B | Alibaba | 27B |
Qwen 3.5 27B | Alibaba | 27B |
Qwen 3.6 27B | Alibaba | 27B |
Gemma 2 27B | 27B | |
Gemma 4 26B A4B MoE(3.8B active) | 25.2B | |
Mistral Small 3.2 | Mistral AI | 24B |
Mistral Small 3.1 | Mistral AI | 24B |
Voxtral Small 24B | Mistral AI | 24B |
Mistral Small 3 | Mistral AI | 24B |
GPT OSS 20B MoE(3.6B active) | OpenAI | 21B |
Claude 3 Haiku | Anthropic | 20B* |
Qwen 3 14B | Alibaba | 14.8B |
Qwen 2.5 14B | Alibaba | 14B |
Phi-4 | Microsoft | 14B |
Ministral 3 14B | Mistral AI | 13.9B |
Nemotron Nano 2 VL 12B | NVIDIA | 12.6B |
Gemma 3 12B | 12B | |
Mistral Nemo | Mistral AI | 12B |
Solar Mini | Upstage | 10.7B |
Qwen 3.5 9B | Alibaba | 9B |
Gemma 2 9B | 9B | |
Nemotron Nano 2 9B | NVIDIA | 9B |
Ministral 3 8B | Mistral AI | 8.8B |
Qwen 3 8B | Alibaba | 8.2B |
Qwen 3 VL 8B | Alibaba | 8.2B |
Qwen 3 VL 8B Thinking | Alibaba | 8.2B |
Granite 4.1 8B | IBM | 8B |
Gemma 3n E4B | 8B | |
GPT-4o mini | OpenAI | 8B* |
Llama 3.1 8B | Meta | 8B |
Llama 3 8B | Meta | 8B |
Ministral 8B | Mistral AI | 8B |
Qwen 2.5 7B | Alibaba | 7.6B |
Mistral 7B | Mistral AI | 7B |
Command R7B | Cohere | 7B |
Phi-4 Multimodal | Microsoft | 5.6B |
Gemma 3 4B | 4B | |
Ministral 3 3B | Mistral AI | 3.8B |
Phi-4 mini | Microsoft | 3.8B |
Phi-3 mini | Microsoft | 3.8B |
Gemini Nano 2 | 3.3B | |
Llama 3.2 3B | Meta | 3.2B |
Granite 4.0 H Micro | IBM | 3B |
Ministral 3B | Mistral AI | 3B |
Gemma 2 2B | 2B | |
Gemini Nano 1 | 1.8B | |
GPT-2 | OpenAI | 1.5B |
Llama 3.2 1B | Meta | 1.2B |
Qwen 2.5 0.5B | Alibaba | 0.5B |
Parameter sizes of popular Large Language Models (as of August 2026)
Conclusion
Large Language Models have fundamentally changed how we interact with computers. They are powerful tools for text processing, programming, and creative tasks, but not a replacement for human judgment and expertise. Those who understand their strengths and limitations can effectively use them for a variety of tasks.
