What is a Large Language Model (LLM)?
A Large Language Model (LLM) is an artificial neural network trained on massive amounts of text data to understand and generate human language. LLMs like GPT-4, Claude, Gemini, and LLaMA can write texts, answer questions, write code, and solve complex tasks.
The term "Large" refers to the number of parameters. Modern LLMs have hundreds of billions of parameters that are optimized during training. The more parameters, the more complex patterns the model can capture.
How Do LLMs Work?
LLMs are based on the Transformer architecture, introduced by Google in 2017. The core is the "attention mechanism," which allows the model to recognize relevant relationships in text, even across large distances.
Training in Three Phases
- Pre-Training: The model learns from billions of texts (books, websites, Wikipedia) to predict the next words. This develops a deep understanding of language.
- Fine-Tuning: The model is adapted to specific tasks or formats, such as following instructions or answering questions in a dialogue format.
- RLHF (Reinforcement Learning from Human Feedback): Humans rate the model's responses, and it learns to prioritize helpful, harmless, and honest answers.
Popular LLMs Overview
GPT-5 (OpenAI)
The GPT series (Generative Pre-trained Transformer) from OpenAI is the most well-known LLM. ChatGPT is based on these models. GPT-5.5 (April 2026) introduced native multimodal inputs and a context window of up to 1 million tokens. On July 9, 2026 OpenAI made the GPT-5.6 family (Sol, Terra, Luna) generally available across ChatGPT, ChatGPT Work, Codex, and the API.
Claude (Anthropic)
Claude is known for particularly long context windows and a focus on safety through "Constitutional AI." The current Claude family includes Claude Opus 5 and Claude Sonnet 5 with 1 million token context windows, plus Claude Haiku 4.5. In June 2026 Anthropic announced Claude Fable 5 and Claude Mythos 5, a new Mythos-class tier above Opus, priced at $10 input / $50 output per million tokens. Fable 5 has been generally available again since July 1, 2026, while Mythos 5 remains limited to vetted Project Glasswing partners.
Gemini (Google)
Google's LLM family ranges from Gemini Nano for mobile devices to Gemini 3.1 Pro for complex reasoning tasks. The models are natively multimodal and can process text, image, audio, and video.
LLaMA / Llama (Meta)
Meta's open-source LLMs gave the developer community free, adaptable models to build on. Llama 3 is freely available and forms the foundation for many specialized models.
Applications of LLMs
- Text Generation: Blog posts, emails, marketing copy
- Programming: Code generation, debugging, code reviews
- Customer Service: Chatbots and automated responses
- Translation: High-quality translations into dozens of languages
- Research: Summarizing documents and extracting facts
- Education: Personalized tutoring and explanations
Limitations and Challenges
Hallucinations
LLMs can generate convincing-sounding but factually incorrect information. They sometimes "invent" facts, quotes, or sources. Therefore, critical review of outputs is important.
Knowledge Cutoff
LLMs have a knowledge cutoff date, meaning they only know information up to a certain point in time. Current events are unknown to them unless they have access to external tools like web search.
Context Window Limitation
Although modern LLMs have large context windows, the amount of text they can process simultaneously is limited. With very long documents, the quality of responses may decrease.
Bias and Fairness
LLMs reflect the biases in their training data. Despite intensive efforts toward fairness, they can reproduce stereotypical or discriminatory patterns.
Using LLMs Effectively
To get the most out of LLMs, good prompts are crucial. Techniques like Chain-of-Thought Prompting can significantly improve the quality of responses.
For developers, APIs from OpenAI, Anthropic, and Google offer the ability to integrate LLMs into their own applications. Costs are typically calculated based on tokens consumed.
Comprehensive LLM Parameter Menu
The following interactive table shows over 60 well-known Large Language Models with their parameter counts. You can search by name, filter by developer, size category or model type, and sort the columns:
Legend:
Showing 254 models
Model | Developer | Parameters |
|---|---|---|
MiniMax M2.7 MoE | MiniMax | Unknown |
MiniMax M2.5 MoE | MiniMax | Unknown |
GLM-4.7 MoE | Z.ai | Unknown |
GLM-4.7-Flash MoE | Z.ai | Unknown |
GLM-4.6V MoE | Z.ai | Unknown |
GPT-5.6 Sol | OpenAI | Unknown |
GPT-5.6 Terra | OpenAI | Unknown |
GPT-5.6 Luna | OpenAI | Unknown |
GPT-5.5 | OpenAI | Unknown |
GPT-5.5 Pro | OpenAI | Unknown |
GPT-5.5 Instant | OpenAI | Unknown |
ChatGPT chat-latest | OpenAI | Unknown |
GPT-5.4 | OpenAI | Unknown |
GPT-5.4 Pro | OpenAI | Unknown |
GPT-5.4 mini | OpenAI | Unknown |
GPT-5.4 nano | OpenAI | Unknown |
GPT-5.3-Codex | OpenAI | Unknown |
GPT-5.3 Instant | OpenAI | Unknown |
GPT-5.2 | OpenAI | Unknown |
GPT-5.1 Instant | OpenAI | Unknown |
GPT-5.1 Thinking | OpenAI | Unknown |
GPT-5 | OpenAI | Unknown |
GPT-5 pro | OpenAI | Unknown |
GPT-5 mini | OpenAI | Unknown |
GPT-5 nano | OpenAI | Unknown |
GPT-4 Turbo | OpenAI | Unknown |
GPT-4.1 | OpenAI | Unknown |
GPT-4.1 mini | OpenAI | Unknown |
GPT-4.1 nano | OpenAI | Unknown |
GPT-3.5 Turbo | OpenAI | Unknown |
o3 | OpenAI | Unknown |
o3-pro | OpenAI | Unknown |
o3-mini | OpenAI | Unknown |
o4-mini | OpenAI | Unknown |
o1 | OpenAI | Unknown |
o1-mini | OpenAI | Unknown |
Claude Fable 5 | Anthropic | Unknown |
Claude Mythos 5 | Anthropic | Unknown |
Claude Sonnet 5 | Anthropic | Unknown |
Claude Opus 5 | Anthropic | Unknown |
Claude Opus 4.8 | Anthropic | Unknown |
Claude Opus 4.7 | Anthropic | Unknown |
Claude Opus 4.6 | Anthropic | Unknown |
Claude Sonnet 4.6 | Anthropic | Unknown |
Claude Opus 4.5 | Anthropic | Unknown |
Claude Opus 4.1 | Anthropic | Unknown |
Claude Sonnet 4.5 | Anthropic | Unknown |
Claude Haiku 4.5 | Anthropic | Unknown |
Claude Sonnet 4 | Anthropic | Unknown |
Claude Opus 4 | Anthropic | Unknown |
Claude Sonnet 3.7 | Anthropic | Unknown |
Claude 3.5 Haiku | Anthropic | Unknown |
Gemini 3.7 Flash | Unknown | |
Gemini 3.6 Flash | Unknown | |
Gemini 3.5 Flash | Unknown | |
Gemini 3.1 Pro Preview | Unknown | |
Gemini 3 Flash MoE | Unknown | |
Gemini 3.5 Flash-Lite | Unknown | |
Gemini 3.1 Flash-Lite MoE | Unknown | |
Gemini 2.5 Pro MoE | Unknown | |
Gemini 2.5 Flash MoE | Unknown | |
Gemini 2.5 Flash-Lite MoE | Unknown | |
Gemini 3 Pro MoE | Unknown | |
Gemini 2.0 Flash MoE | Unknown | |
Gemini 1.5 Pro MoE | Unknown | |
Grok 4.6 | xAI | Unknown |
Grok 4.5 | xAI | Unknown |
Grok 4.3 | xAI | Unknown |
Grok 4.20 Reasoning | xAI | Unknown |
Grok 4.20 Multi-Agent | xAI | Unknown |
Grok Build 0.1 | xAI | Unknown |
Grok 4 | xAI | Unknown |
Grok 3 | xAI | Unknown |
Grok 2 | xAI | Unknown |
Mistral Medium 3.5 | Mistral AI | Unknown |
MiniMax M3 | MiniMax | Unknown |
Qwen 3.7 Max MoE | Alibaba | Unknown |
Qwen 3.7 Plus MoE | Alibaba | Unknown |
Nova 2 Lite | Amazon | Unknown |
Nova Premier 1.0 | Amazon | Unknown |
Nova Pro 1.0 | Amazon | Unknown |
Nova Lite 1.0 | Amazon | Unknown |
Nova Micro 1.0 | Amazon | Unknown |
Sonar | Perplexity | Unknown |
Sonar Pro | Perplexity | Unknown |
Sonar Reasoning Pro | Perplexity | Unknown |
Sonar Deep Research | Perplexity | Unknown |
MiMo-V2.5 MoE | Xiaomi | Unknown |
Solar Mini | Upstage | Unknown |
Kimi K3 MoE(104B active) | Moonshot AI | 2.8T |
Qwen 3.8 2.4T A95B MoE(95B active) | Alibaba | 2.4T |
Claude 3 Opus | Anthropic | 2T* |
Llama 4 Behemoth MoE(288B active) | Meta | 2T |
GPT-4 MoE(220B active) | OpenAI | 1.76T* |
DeepSeek-V4-Pro MoE(49B active) | DeepSeek | 1.6T |
Ring 2.6 1T MoE(63B active) | inclusionAI | 1T |
Kimi K2.5 MoE(32B active) | Moonshot AI | 1T |
Kimi K2 Thinking MoE(32B active) | Moonshot AI | 1T |
Kimi K2 0711 MoE(32B active) | Moonshot AI | 1T |
Kimi K2 0905 MoE(32B active) | Moonshot AI | 1T |
Kimi K2.6 MoE(32B active) | Moonshot AI | 1T |
Kimi K2.7 Code MoE(32B active) | Moonshot AI | 1T |
Qwen 3.6 Max-Preview MoE | Alibaba | 1T* |
Yi-Large MoE | 01.AI | 1T |
MiMo-V2.5-Pro MoE(42B active) | Xiaomi | 1T |
MiMo-V2.5-Pro-UltraSpeed MoE(42B active) | Xiaomi | 1T |
GLM-5.2 MoE(40B active) | Z.ai | 744B |
GLM-5.3 MoE(40B active) | Z.ai | 744B |
GLM-5.1 MoE(40B active) | Z.ai | 744B |
GLM-5 MoE(40B active) | Z.ai | 744B |
DeepSeek-V3 0324 MoE(37B active) | DeepSeek | 685B |
DeepSeek-V3.2 Exp MoE(37B active) | DeepSeek | 685B |
DeepSeek-V3.2 MoE(37B active) | DeepSeek | 685B |
Mistral Large 3 MoE(41B active) | Mistral AI | 675B |
DeepSeek-V3.1 Terminus MoE(37B active) | DeepSeek | 671B |
DeepSeek-V3.1 MoE(37B active) | DeepSeek | 671B |
DeepSeek-R1 0528 MoE(37B active) | DeepSeek | 671B |
DeepSeek-V3 MoE(37B active) | DeepSeek | 671B |
DeepSeek-R1 MoE(37B active) | DeepSeek | 671B |
Nemotron 3 Ultra MoE(55B active) | NVIDIA | 550B |
PaLM | 540B | |
Megatron-Turing NLG | NVIDIA | 530B |
Qwen 3 Coder 480B A35B MoE(35B active) | Alibaba | 480B |
MiniMax M1 MoE(45.9B active) | MiniMax | 456B |
MiniMax-01 MoE(45.9B active) | MiniMax | 456B |
ERNIE 4.5 VL 424B A47B MoE(47B active) | Baidu | 424B |
Hermes 4 405B | Nous Research | 405B |
Llama 3.1 405B | Meta | 405B |
Llama 4 Maverick MoE(17B active) | Meta | 400B |
Trinity Large Thinking MoE(13B active) | Arcee AI | 398B |
Nex-N2-Pro MoE(17B active) | Nex AGI | 397B |
Qwen 3.5 397B A17B MoE(17B active) | Alibaba | 397B |
GLM-4.6 MoE(32B active) | Z.ai | 355B |
GLM-4.5 MoE(32B active) | Z.ai | 355B |
Nemotron-4 340B | NVIDIA | 340B |
PaLM 2 | 340B* | |
Grok 1 MoE(86B active) | xAI | 314B |
Hy3 MoE(21B active) | Tencent | 295B |
DeepSeek-V4-Flash MoE(13B active) | DeepSeek | 284B |
DeepSeek-V2 MoE(21B active) | DeepSeek | 236B |
Qwen 3 235B A22B Instruct 2507 MoE(22B active) | Alibaba | 235B |
Qwen 3 235B A22B MoE(22B active) | Alibaba | 235B |
Qwen 3 235B A22B Thinking 2507 MoE(22B active) | Alibaba | 235B |
Qwen 3 VL 235B A22B MoE(22B active) | Alibaba | 235B |
Qwen 3 VL 235B A22B Thinking MoE(22B active) | Alibaba | 235B |
MiniMax M2.1 MoE(10B active) | MiniMax | 230B |
MiniMax M2 MoE(10B active) | MiniMax | 230B |
GPT-4o | OpenAI | 200B* |
Step 3.7 Flash MoE(11B active) | StepFun | 198B |
Step 3.5 Flash MoE(11B active) | StepFun | 196.8B |
Falcon 180B | TII | 180B |
BLOOM | BigScience | 176B |
GPT-3 | OpenAI | 175B |
Claude 3.5 Sonnet | Anthropic | 175B* |
OPT-175B | Meta | 175B |
Mixtral 8x22B MoE(39B active) | Mistral AI | 141B |
LaMDA | 137B | |
DBRX MoE(36B active) | Databricks | 132B |
Ling 3.0 Flash MoE(5.1B active) | inclusionAI | 124B |
Mistral Large 2 | Mistral AI | 123B |
Qwen 3.5 122B A10B MoE(10B active) | Alibaba | 122B |
Nemotron 3 Super MoE(12B active) | NVIDIA | 120B |
Mistral Small 4 MoE(6B active) | Mistral AI | 119B |
Laguna S 2.1 MoE(8B active) | Poolside | 118B |
GPT OSS 120B MoE(5.1B active) | OpenAI | 117B |
Command A | Cohere | 111B |
Llama 4 Scout MoE(17B active) | Meta | 109B |
GLM-4.5 Air MoE(12B active) | Z.ai | 106B |
GLM-4.5V MoE(12B active) | Z.ai | 106B |
Command R+ | Cohere | 104B |
Solar Pro 3 MoE | Upstage | 102B |
Hunyuan A13B Instruct MoE(13B active) | Tencent | 80B |
Qwen 3 Coder Next MoE(3B active) | Alibaba | 80B |
Qwen 3 Next 80B A3B Instruct MoE(3B active) | Alibaba | 80B |
Qwen 3 Next 80B A3B Thinking MoE(3B active) | Alibaba | 80B |
Virtuoso Large | Arcee AI | 72B |
Qwen 2.5 VL 72B | Alibaba | 72B |
Qwen 2.5 72B | Alibaba | 72B |
Hermes 4 70B | Nous Research | 70B |
DeepSeek-R1 Distill Llama 70B | DeepSeek | 70B |
Claude 3 Sonnet | Anthropic | 70B* |
Llama 3.3 70B | Meta | 70B |
Llama 3.1 70B | Meta | 70B |
Llama 3 70B | Meta | 70B |
Llama 2 70B | Meta | 70B |
Mixtral 8x7B MoE(14B active) | Mistral AI | 56B |
Falcon 40B | TII | 40B |
Nex-N2-mini MoE(3B active) | Nex AGI | 35B |
Qwen 3.6 35B A3B MoE(3B active) | Alibaba | 35B |
Qwen 3.5 35B A3B MoE(3B active) | Alibaba | 35B |
Yi-34B | 01.AI | 34B |
Laguna XS 2.1 MoE(3B active) | Poolside | 33B |
Qwen 3 32B | Alibaba | 32B |
Qwen 2.5 Coder 32B | Alibaba | 32B |
Qwen 3 VL 32B | Alibaba | 32B |
Qwen 2.5 32B | Alibaba | 32B |
Command R | Cohere | 32B |
Gemma 4 31B | 31B | |
Solar Pro 2 | Upstage | 31B |
Nemotron 3 Nano MoE(3.5B active) | NVIDIA | 30B |
Nemotron 3.5 Lightning MoE(3B active) | NVIDIA | 30B |
Qwen 3 30B A3B Instruct 2507 MoE(3B active) | Alibaba | 30B |
Qwen 3 Coder 30B A3B MoE(3B active) | Alibaba | 30B |
Qwen 3 30B A3B MoE(3B active) | Alibaba | 30B |
Qwen 3 30B A3B Thinking 2507 MoE(3B active) | Alibaba | 30B |
Qwen 3 VL 30B A3B MoE(3B active) | Alibaba | 30B |
Qwen 3 VL 30B A3B Thinking MoE(3B active) | Alibaba | 30B |
Gemma 3 27B | 27B | |
Qwen 3.8 27B | Alibaba | 27B |
Qwen 3.5 27B | Alibaba | 27B |
Qwen 3.6 27B | Alibaba | 27B |
Gemma 2 27B | 27B | |
Gemma 4 26B A4B MoE(4B active) | 26B | |
Mistral Small 3.2 | Mistral AI | 24B |
Mistral Small 3.1 | Mistral AI | 24B |
Voxtral Small 24B | Mistral AI | 24B |
Mistral Small 3 | Mistral AI | 24B |
GPT OSS 20B MoE(3.6B active) | OpenAI | 21B |
Claude 3 Haiku | Anthropic | 20B* |
Ministral 3 14B | Mistral AI | 14B |
Qwen 3 14B | Alibaba | 14B |
Qwen 2.5 14B | Alibaba | 14B |
Phi-4 | Microsoft | 14B |
Gemma 3 12B | 12B | |
Mistral Nemo | Mistral AI | 12B |
Qwen 3.5 9B | Alibaba | 9B |
Gemma 2 9B | 9B | |
Granite 4.1 8B | IBM | 8B |
Gemma 3n E4B | 8B | |
Ministral 3 8B | Mistral AI | 8B |
Qwen 3 8B | Alibaba | 8B |
Qwen 3 VL 8B | Alibaba | 8B |
Qwen 3 VL 8B Thinking | Alibaba | 8B |
GPT-4o mini | OpenAI | 8B* |
Llama 3.1 8B | Meta | 8B |
Llama 3 8B | Meta | 8B |
Ministral 8B | Mistral AI | 8B |
Mistral 7B | Mistral AI | 7B |
Qwen 2.5 7B | Alibaba | 7B |
Command R7B | Cohere | 7B |
Phi-4 Multimodal | Microsoft | 5.6B |
Gemma 3 4B | 4B | |
Phi-4 mini | Microsoft | 3.8B |
Phi-3 mini | Microsoft | 3.8B |
Gemini Nano 2 | 3.3B | |
Granite 4.0 H Micro | IBM | 3B |
Llama 3.2 3B | Meta | 3B |
Ministral 3 3B | Mistral AI | 3B |
Ministral 3B | Mistral AI | 3B |
Gemma 2 2B | 2B | |
Gemini Nano 1 | 1.8B | |
GPT-2 | OpenAI | 1.5B |
Llama 3.2 1B | Meta | 1B |
Qwen 2.5 0.5B | Alibaba | 0.5B |
Parameter sizes of popular Large Language Models (as of August 2026)
Conclusion
Large Language Models have fundamentally changed how we interact with computers. They are powerful tools for text processing, programming, and creative tasks, but not a replacement for human judgment and expertise. Those who understand their strengths and limitations can effectively use them for a variety of tasks.
