Skip to main content

The Best Open Source LLMs: 211 Models Compared

Sortable directory of 211 open source LLMs with benchmarks, licenses, API prices, and context windows. Plus: a hardware check for your PC or Mac.

FHFinn Hillebrandt
AI Technology
The Best Open Source LLMs: 211 Models Compared
Links marked with * are affiliate links. If a purchase is made through such links, we receive a commission.

Open source LLMs are one of the most important AI trends of 2026.

And for good reason:

Open source models were long significantly weaker than proprietary models. By fall 2026 they have almost completely caught up, especially thanks to Chinese labs:

DeepSeek V4 Pro, GLM-5.3 from Z.ai, and Kimi K3 from Moonshot AI are now only a few points behind Claude Opus 5 on SWE-bench Verified, the strongest model in the independent Vals AI measurement. The small Qwen3.8 27B even beats GPT-5.5 there.

In this article, you'll find a sortable, filterable directory of 211 open source LLMs, including benchmark scores, licenses, API prices, context windows, and capabilities.

The hardware check also shows you which of these models run on your PC or Mac and how fast they respond there.

Additionally, I'll show you how to easily and freely use open LLMs on your own computer (without needing to program or use the terminal).

TL;DRKey Takeaways
  • DeepSeek V4 Pro 0813, GLM-5.3, and Kimi K3 are less than 4 points behind Claude Opus 5 on SWE-bench Verified, as measured by Vals AI
  • The best small model is Qwen3.8 27B under Apache 2.0. It fits on a 24 GB graphics card and even beats GPT-5.5 on SWE-bench Verified
  • 211 open source LLMs in a filterable, sortable directory, from MIT and Apache 2.0 through to restricted research-only licenses. Columns like prices, context, and capabilities can be toggled individually
  • Read the licenses carefully. GLM-5.3, Kimi K3, and MiniMax M3 have their own licenses with conditions, and MiniMax M2.7 is non-commercial only
  • Local usage possible with tools like Ollama, LM Studio, or GPT4All, but the new top models need serious hardware. The hardware check shows which models run on your PC or Mac and how many tokens per second to expect

All Open Source LLMs at a Glance

The directory contains every open-weights model from the models.dev catalog plus curated classics, sorted by release date by default. Use "Columns" to reveal more data, such as modalities, knowledge cutoff, max output, or the number of API providers:

Showing 50 of 211 matches (211 models in total)

Benchmark values come from different tests and test setups. Values in the same column are therefore not directly comparable and cannot be sorted together.

KnowledgeMath / ScienceCode
Xiaomi309B (15B active)1M––28.8%Terminal-BenchMIT$0.14$0.28Sep 2026
MiMo-V2.6-Pro
Xiaomi1T (42B active)1M––34.9%Terminal-BenchMIT$0.44$0.87Sep 2026
MiMo-V2.6-Pro-UltraSpeed
Xiaomi1T (42B active)1M–––MIT$4.35$8.70Sep 2026
DeepSeek V4.1 Flash
DeepSeek552B (16B active)1M––90.6%Terminal-BenchMIT$0.13$0.28Sep 2026
MiniCPM5-2B
openbmb2B128K–––Apache 2.0$0.12$0.74Sep 2026
Hy4 preview
Tencent770B (49B active)1M–––Apache 2.0$0.67$2.00Aug 2026
Qwen 3.8 Flash Next
Alibaba125B (6B active)256K––62.5%SWE-Bench ProQwen Community License 1.0$0.15$0.47Aug 2026
GLM-5.3-Flash
Z.ai320B (18B active)1M––63.4%DeepSWEMIT$0.015$0.025Aug 2026
DeepSeek V4 Flash Vision Exp
DeepSeek305B1M––83.9%Terminal-BenchMIT$0.080$0.20Aug 2026
Ornith 1.5 35B A3B
DeepReinforce35B (3B active)256K–––MIT$0.10$0.40Aug 2026
Qwen3.8 27B
Alibaba27B256K––61.7%SWE-bench ProApache 2.0$0.080$0.35Aug 2026
GLM-5.3
Z.ai753B1M––88.2%Terminal-BenchGLM-5.3 License$0.38$1.19Aug 2026
Qwen3.8 2.4T A95B
Alibaba2.4T (95B active)256K–––Qwen3.8-Max License$1.95$5.95Aug 2026
DeepSeek V4 Pro 0813
DeepSeek1.6T (49B active)1M––87.9%Terminal-BenchMIT$0.26$0.79Aug 2026
Nemotron 3.5 Lightning 30B A3B
NVIDIA30B (3B active)256K––51.6%SWE-Bench VerifiedNVIDIA Open Model License$0.050$0.15Aug 2026
Muse Glimmer 30B
Meta30B128K––76%SWE-Bench VerifiedApache 2.0$0.20$0.80Aug 2026
Motif 3
motif-technologies314B (13.2B active)256K–––MIT$0.50$2.00Aug 2026
DeepSeek V4 Flash 0731
DeepSeek284B (13B active)1M––82.7%Terminal-BenchMIT$0.021$0.070Jul 2026
Inkling Small
thinkingmachines276B (12B active)1M–––Apache 2.0$0.45$1.20Jul 2026
Laguna S 2.1
Poolside118B (8B active)1M–––OpenMDW-1.1$0.090$0.18Jul 2026
Kimi K3
Moonshot AI2.8T (104B active)1M––88.3%Terminal-BenchKimi K3 License$2.00$8.00Jul 2026
Inkling
thinkingmachines975B (41B active)1M–––Apache 2.0$0.95$4.05Jul 2026
Hy3
Tencent295B (21B active)250K––78%SWE-Bench VerifiedApache 2.0$0.066$0.26Jul 2026
Laguna XS 2.1
Poolside33B (3B active)256K––70.9%SWE-Bench VerifiedOpenMDW-1.1$0.060$0.12Jul 2026
Ornith 1.0 31B
DeepReinforce31B256K–––MIT––Jun 2026
Ornith 1.0 35B
DeepReinforce35B256K––75.6%SWE-Bench VerifiedMIT––Jun 2026
Ornith 1.0 397B
DeepReinforce397B256K––82.4%SWE-Bench VerifiedMIT––Jun 2026
Ornith 1.0 9B
DeepReinforce9B256K––69.4%SWE-Bench VerifiedMIT––Jun 2026
GLM-5.2
Z.ai753B1M–91.2%GPQA62.1%SWE-Bench ProMIT$0.30$1.05Jun 2026
Kimi K2.7 Code
Moonshot AI1T (32B active)256K–89.6%GPQA67.4%Terminal-BenchModified MIT$0.28$1.10Jun 2026
Kimi K2.7 Code Highspeed
Moonshot AI1T (32B active)256K–89.6%GPQA67.4%Terminal-BenchModified MIT$1.90$8.00Jun 2026
North Mini Code
Cohere30B (3B active)250K––61%SWE-Bench VerifiedApache 2.0––Jun 2026
Gemma 4 12B IT
Google12B256K–––Apache 2.0$0.050$0.25Jun 2026
MiMo-V2.5-Pro-UltraSpeed
Xiaomi1T (42B active)1M–––MIT$1.31$2.61Jun 2026
Nemotron 3 Ultra 550B A55B
NVIDIA550B (55B active)1M86.8%MMLU-Pro87%GPQA89%LiveCodeBenchOpenMDW-1.1$0.10$0.10Jun 2026
Nex-N2-Pro
nex-agi397B (17B active)256K–––Apache 2.0$0.50$2.50Jun 2026
MiniMax-M3
MiniMax428B (23B active)1M–92.9%GPQA80.5%SWE-Bench VerifiedMiniMax Community License$0.23$0.90Jun 2026
Step 3.7 Flash
StepFun198B (11B active)250K––76.5%SWE-Bench VerifiedApache 2.0$0.19$1.11May 2026
Command A Plus
Cohere218B (25B active)125K–––CC BY-NC-4.0$0.30$1.50May 2026
MiniCPM5-1B
openbmb1B128K–––Apache 2.0––May 2026
Mistral Medium 3.5
Mistral AI128B256K––77.6%SWE-Bench VerifiedModified MIT (Mistral)$1.50$6.90Apr 2026
Nemotron 3 Nano Omni 30B A3B Reasoning
NVIDIA30B (3B active)250K77.3%MMLU-Pro72.2%GPQA63.2%LiveCodeBenchNVIDIA Open Model License$0.20$0.80Apr 2026
Laguna M.1
Poolside225B (23B active)256K––74.6%SWE-Bench VerifiedApache 2.0––Apr 2026
Laguna XS.2
Poolside33B (3B active)256K––69.9%SWE-Bench VerifiedApache 2.0––Apr 2026
DeepSeek V4 Flash
DeepSeek284B (13B active)1M83%MMLU-Pro85%GPQA88%LiveCodeBenchMIT$0.035$0.070Apr 2026
DeepSeek V4 Pro
DeepSeek1.6T (49B active)1M87.5%MMLU-Pro90.1%GPQA93.5%LiveCodeBenchMIT$0.35$0.70Apr 2026
DeepSeek V4 Pro 0423
DeepSeek1.6T (49B active)1M–––MIT$1.32$3.30Apr 2026
Qwen3.6 27B
Alibaba27B256K86.2%MMLU-Pro87.8%GPQA77.2%SWE-Bench VerifiedApache 2.0$0.20$0.60Apr 2026
MiMo-V2.5
Xiaomi310B (15B active)1M86.3%MMLU–56.1%SWE-Bench ProMIT$0.14$0.28Apr 2026
MiMo-V2.5-Pro
Xiaomi1T (42B active)1M89.4%MMLU–78.9%SWE-Bench VerifiedMIT$0.40$0.80Apr 2026

1. Key Benchmarks Explained

A single score says little about a model. That's why the directory shows up to three values per model, one each for knowledge, math and science, and code.

The small label under each score tells you which benchmark a value comes from. These nine benchmarks appear there most often:

BenchmarkMMLU
What it measuresGeneral knowledge and problem solving across school and university subjects
Size57 subjects, about 14,000 questions with 4 choices
Directory columnMMLU
BenchmarkMMLU-Pro
What it measuresHarder successor with more reasoning
Sizeover 12,000 questions, up to 10 choices
Directory columnMMLU
BenchmarkGPQA Diamond
What it measuresGraduate-level science that can't be solved by googling
Size198 questions in biology, physics, and chemistry
Directory columnMath
BenchmarkMATH-500
What it measuresCompetition-style math with worked solutions
Size500 problems across 7 subjects
Directory columnMath
BenchmarkAIME
What it measuresUS math competition for top AMC scorers
Size30 problems per year, integer answers
Directory columnMath
BenchmarkHumanEval
What it measuresWriting Python functions from a docstring
Size164 problems with unit tests
Directory columnCode
BenchmarkLiveCodeBench
What it measuresFresh contest problems, time-filtered against contamination
Sizecontinuously new problems from LeetCode, AtCoder, and Codeforces
Directory columnCode
BenchmarkSWE-bench Verified
What it measuresResolving real GitHub issues, checked with hidden tests
Size500 human-validated tasks
Directory columnCode
BenchmarkTerminal-Bench 2.x
What it measuresAgent tasks on the command line
Size89 tasks in Docker environments
Directory columnCode

Some older tests are close to saturated by now. On HumanEval, for example, Kimi K2.5 reaches 99%. The harder benchmarks such as GPQA Diamond, LiveCodeBench, and SWE-bench Verified are more telling. If a model has no value, the directory shows "–".

How Close Open Models Are to the Top

On SWE-bench Verified, the gap to the proprietary frontier has almost disappeared. GLM-5.3 scores 95.4% at Vals AI, just ahead of Claude Fable 5. Only Claude Opus 5, GPT-5.6 Sol, and Grok 4.6 do better:

Open and proprietary models on SWE-bench Verified

Models:
Claude Opus 5
GPT-5.6 Sol
Grok 4.6
GLM-5.3
Claude Fable 5
Kimi K3
GLM-5.3-Flash
Qwen3.8 27B
Inkling Small
Gemini 3.7 Flash
DeepSeek V4 Pro (April)
Source: Vals AI, SWE-bench Verified
gradually.ai

It gets even closer with the final release of DeepSeek V4 Pro. Version 0813 reaches 96.4% in the same measurement, only 0.6 points behind Claude Opus 5. It's missing from the chart because it doesn't have its own model profile yet.

On price, the picture is more mixed than you might think. GLM-5.3-Flash reaches 92% for $0.15 per million input tokens. But OpenAI's GPT-5.6 Luna is almost level at 93% for $0.20:

Price and SWE-bench performance of open and proprietary models
Z.ai
Moonshot AI
Alibaba
DeepSeek
Anthropic
OpenAI
Google
Efficiency frontier (best price-performance)
Sources: Vals AI, official API prices
gradually.ai

The real advantage of open models lies elsewhere anyway. You can host them yourself, adapt them, and run them with any provider instead of being tied to a single API.

2. The Best Open Source LLMs in September 2026

Since July, the top of the field has been completely reshuffled. Kimi K3, GLM-5.3, and the final release of DeepSeek V4 Pro have clearly overtaken the spring models.

The six most important open models at a glance:

FeatureDeepSeek V4 Pro 0813GLM-5.3Kimi K3GLM-5.3-FlashMiMo-V2.6-ProQwen3.8 27B
DeveloperDeepSeekZ.aiMoonshot AIZ.aiXiaomiAlibaba
ReleasedAug 2026Aug 2026Jul 2026Aug 2026Sep 2026Aug 2026
Licensecheck the conditions in the license textMITGLM-5.3 LicenseKimi K3 LicenseMITMITApache 2.0
Parameterstotal (active per token)1.6T (49B)753B2.8T (104B)320B (18B)1.02T (42B)27B
Context window1M1M1M1M1M262K (up to 1M)
InputTextTextText, image, videoText, image, videoText, image, video, audioText, image, video
SWE-bench Verifiedmeasured independently by Vals AI96.4%95.4%93.4%92.0%no score yet86.0%
StrengthAgentic codingSoftware engineeringCoding, agents, visionMultimodal and cheapOmnimodal, agentsRuns on one 24 GB GPU

DeepSeek V4 Pro 0813

In mid-August, DeepSeek replaced the V4 Pro preview with the final 0813 release. The architecture stays the same. According to DeepSeek, the biggest gains are in agentic capabilities in production use.

At 96.4% on SWE-bench Verified, it's the strongest open model in the Vals AI measurement. It also comes with the MIT license and no extra conditions.

GLM-5.3

Z.ai aims GLM-5.3 at software engineering and long agentic tasks. The context window holds 1 million tokens, and a single answer can be up to 128,000 tokens long.

The license has changed. Unlike GLM-5.2 and GLM-5.3-Flash, GLM-5.3 no longer uses MIT but its own GLM-5.3 License. It requires a security review by Z.ai, but only for Model-as-a-Service providers with more than $10 billion in annual revenue.

Kimi K3

With 2.8 trillion parameters, Kimi K3 is the largest model in the directory. Of those, 104 billion work on each token, plus a vision encoder for images and video.

Moonshot is surprisingly modest about it. According to its own blog, K3 still trails Claude Fable 5 and GPT-5.6 Sol, and the Vals measurement narrowly confirms that at 93.4%.

The Kimi K3 License allows commercial use, with two conditions for very large providers. Model-as-a-Service providers with more than $20 million in revenue over twelve months need a separate agreement. Above 100 million monthly users, "Kimi K3" has to be displayed prominently in the interface.

GLM-5.3-Flash

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series and understands images and video as well as text. With 18 billion active parameters, it's far leaner than GLM-5.3.

It still reaches 92% on SWE-bench Verified. The API costs $0.15 per million input tokens, and the weights are released under the standard MIT license.

MiMo-V2.6-Pro

Xiaomi introduced MiMo-V2.6-Pro on September 22, which makes it the newest model in this selection. It handles text, images, video, and audio in a single model and uses the MIT license.

There are no independent measurements yet. Xiaomi itself reports 89.9 on Terminal-Bench 2.1, but with its own test setup.

Qwen3.8 27B

The best small model comes from Alibaba. Qwen3.8 27B is a classic dense model with 27 billion parameters and reaches 86% on SWE-bench Verified. That puts it ahead of GPT-5.5 and Gemini 3.1 Pro.

With Q5_K_M quantization, it fits on a graphics card with 24 GB, such as an RTX 4090. The hardware check in section 4 shows how fast it runs there. The license is Apache 2.0.

Other Strong Models

These models are also worth a look:

  • DeepSeek V4.1 Flash (September, MIT): a multimodal MoE with 552 billion parameters, 16 billion of which are active while generating. DeepSeek calls it the smallest model of its new architecture.
  • Qwen3.8 2.4T-A95B (August, Qwen3.8-Max License): the open version of Alibaba's flagship Qwen3.8-Max. It handles text only and always runs in thinking mode.
  • Tencent Hy4 preview (August, Apache 2.0): 770 billion parameters, 49 billion of them active, focused on software engineering and office work.
  • Inkling and Inkling Small (July, Apache 2.0): two multimodal models from Thinking Machines Lab. Inkling Small reaches 82.2% on SWE-bench Verified with only 12 billion active parameters.
  • Qwen 3.8 Flash Next (August, Qwen Community License 1.0): the architecture preview for Qwen4 with only 6 billion active parameters. On QwenCloud, it runs under the name "qwen3.8-flash".
  • MiniMax M3 (June, MiniMax Community License): multimodal with a 1 million token context. For commercial use, the license requires the notice "Built with MiniMax M3".
  • Nemotron 3 Ultra (June, OpenMDW 1.1): NVIDIA's largest Nemotron model with 550 billion parameters, built for long agent runs.
  • Gemma 4 12B (June, Apache 2.0): Google's model for local agents. According to Google, it runs on laptops with 16 GB of VRAM or unified memory.

The spring models such as DeepSeek V4 Pro (April), Kimi K2.6, and GPT-OSS-120B remain usable, but clearly trail the new frontier.

3. LLM Licenses Explained

Here's an overview of the most commonly used licenses for open source LLMs.

MIT License

A very permissive open source license, similar to Apache 2.0. It allows unrestricted use, modification, and distribution of the LLM, including in proprietary programs, as long as the copyright notice is retained. DeepSeek V3 uses MIT with some restrictions for military use.

Llama 2 Community / Llama 3 Community

Meta released Llama 2 and Llama 3 under these licenses. They allow free use of the LLMs for research and commercial applications with up to 700 million monthly active users. The source code and model weights are freely available.

Qwen License / Qianwen LICENSE

Qwen models are released under various licenses. While smaller models are often licensed under Apache 2.0, larger models like Qwen2.5-72B have special license terms that allow commercial use with certain restrictions.

Apache 2.0

A very permissive open source license with minimal restrictions. It allows use, modification, and distribution of the LLM, including in proprietary programs, as long as the copyright notice is retained. It contains no copyleft clause.

CC BY-NC-4.0

A Creative Commons license that allows editing and sharing the LLM in any form, but not for commercial purposes. The author's name must be credited.

CC BY-NC-SA-4.0

Similar to CC BY-NC-4.0, but with the additional Share-Alike condition. This means forks or modified versions of an LLM must be distributed under the same conditions.

Non-Commercial

Here, using the LLM for commercial purposes is prohibited. However, what exactly counts as "commercial" is not always clearly defined or delimited.

Usually, "non-commercial" models are only released for research purposes or private use.

4. Which Open Source LLMs Run on Your Computer?

The strongest open source LLM is useless if it doesn't fit on your computer.

And that is exactly the sticking point with local models. A model has to fit completely into the memory of your graphics card or your Mac. Otherwise it won't start at all or crawls along at a snail's pace.

The hardware check below tells you in a few seconds which models from the directory run on your hardware. You pick your device, use case, and context length, and for every model you get the right quantization, the memory it needs, the estimated speed, and a recommendation score:

Available for models: 11.2 GB of fast memory at 504 GB/s, plus 24 GB of system RAM for offloading.

109 of 198 models run on this computer, 44 of them match your filters.

Model
Fit and memoryQuantizationSpeedScore
Qwen3.5 9BAlibaba · 9B
Fits8.2 GB · 73% usedQ6_K54 tokens/sFluent90
Qwen3.6 35B-A3BAlibaba · 35B (3B active)
Fits22.1 GB · 63% usedpartly in RAMQ4_K_M27 tokens/sUsable90
Ornith 1.0 9BDeepReinforce · 9B
Fits8.2 GB · 73% usedQ6_K54 tokens/sFluent89
Qwen3.5 35B-A3BAlibaba · 35B (3B active)
Fits22.1 GB · 63% usedpartly in RAMQ4_K_M27 tokens/sUsable89
Ornith 1.5 35B A3BDeepReinforce · 35B (3B active)
Fits22.1 GB · 63% usedpartly in RAMQ4_K_M27 tokens/sUsable89
Gemma 4 12B ITGoogle · 12B
Fits8.3 GB · 74% usedQ4_K_M53 tokens/sFluent89
Laguna XS 2.1Poolside · 33B (3B active)
Fits21.1 GB · 60% usedpartly in RAMQ4_K_M27 tokens/sUsable89
Gemma 4 26B A4B ITGoogle · 26B (4B active)
Fits16.8 GB · 48% usedpartly in RAMQ4_K_M26 tokens/sUsable89
Nemotron 3.5 Lightning 30B A3BNVIDIA · 30B (3B active)
Fits21 GB · 60% usedpartly in RAMQ4_K_M25 tokens/sUsable89
GLM-4.6V-FlashZ.ai · 9B
Fits8.2 GB · 73% usedQ6_K54 tokens/sFluent89
Trendyol Asure 12Btrendyol · 12B
Fits8.7 GB · 78% usedQ4_K_M51 tokens/sFluent88
Laguna XS.2Poolside · 33B (3B active)
Fits21.1 GB · 60% usedpartly in RAMQ4_K_M27 tokens/sUsable88
North Mini CodeCohere · 30B (3B active)
Fits21 GB · 60% usedpartly in RAMQ4_K_M25 tokens/sUsable88
Apertus 8Bswiss-ai · 8B
Fits8.1 GB · 73% usedQ6_K57 tokens/sFluent88
Ministral 3 8BMistral AI · 8B
Fits8.2 GB · 73% usedQ6_K57 tokens/sFluent88
Qwen3 30B A3BAlibaba · 30B (3B active)
Fits19.7 GB · 56% usedpartly in RAMQ4_K_M28 tokens/sUsable87
Qwen3-Coder 30B-A3B InstructAlibaba · 30B (3B active)
Fits19.7 GB · 56% usedpartly in RAMQ4_K_M28 tokens/sUsable87
Nemotron Cascade 2 30B A3BNVIDIA · 30B (3B active)
Fits21 GB · 60% usedpartly in RAMQ4_K_M25 tokens/sUsable87
GLM-4.7-FlashZ.ai · 30B (3B active)
Fits19.3 GB · 55% usedpartly in RAMQ4_K_M29 tokens/sUsable87
Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA · 30B (3B active)
Fits21 GB · 60% usedpartly in RAMQ4_K_M25 tokens/sUsable87
Trinity Miniarcee-ai · 26B (3B active)
Fits16.6 GB · 47% usedpartly in RAMQ4_K_M35 tokens/sFluent87
GPT OSS 20BOpenAI · 21B (3.6B active)
Fits13.6 GB · 39% usedpartly in RAMQ4_K_M42 tokens/sFluent87
Nemotron Nano 12B v2 VLNVIDIA · 12B
Fits8.7 GB · 78% usedQ4_K_M52 tokens/sFluent87
Nemotron 3 Nano 30B A3BNVIDIA · 30B (3B active)
Fits21 GB · 60% usedpartly in RAMQ4_K_M25 tokens/sUsable86
InternLM3 8B-InstructShanghai AI Lab · 8B
Fits7.5 GB · 67% usedQ6_K60 tokens/sFluent86

The next step with more memory:

GLM-4.5-Air needs 66.9 GB at Q4_K_M, you are 32.4 GB short
GLM-4.5V needs 66.9 GB at Q4_K_M, you are 32.4 GB short
GLM-4.6V needs 66.9 GB at Q4_K_M, you are 32.4 GB short

Estimate for llama.cpp-based tools such as LM Studio and Ollama with GGUF files. Output speed is based on the average KV cache state during a chat. Full-attention layers use 50% of the context length used for the estimate, while sliding-window layers are averaged separately. The result is an average, not the speed at every moment in a chat. Reading long prompts takes extra time. Fit levels, quantization choice, and the score follow the open source tool llmfit by Alex Jones.

In Chrome and Edge, the check detects your chip or graphics card automatically. You set the memory yourself, because the browser doesn't reveal it.

The small arrow in front of each model opens the details. There you can see how the memory needs and the score come together and what to watch out for. By default, the check hides older models and everything below 10 tokens per second; you can turn off both filters.

How the Hardware Check Calculates

The principle behind it is simpler than it sounds. For every new token, an LLM has to read its weights from memory once. Speed therefore depends mainly on memory bandwidth, meaning how many gigabytes per second your chip can pull from memory.

According to its spec sheet, an RTX 4090 has 1,008 GB/s, while a Mac with an M4 gets 120 GB/s. That's why the same model runs several times faster on the graphics card.

Five more factors go into the calculation:

  • Quantization: The weights are stored with fewer bits, just under 4.9 instead of 16 bits for Q4_K_M. An 8B model shrinks from 16 to about 5 GB and gets faster accordingly. The stronger the quantization, however, the less accurate the answers become. That's why the check tries every level from Q8_0 to Q4_K_M for each model and picks the one with the best score. Q3_K_M and Q2_K only come into play when nothing else fits.
  • KV cache: For every token in the context, the model stores intermediate results, and their size varies a lot between architectures. The check calculates this exactly for 173 models from their configuration file on Hugging Face. For every new token, the model reads its currently occupied KV cache. The estimated speed therefore uses the average cache state during a chat. Full-attention layers use 50% of the context length used for the estimate, while sliding-window layers are averaged separately. A quantized KV cache (Q8_0 or Q4_0) roughly halves or quarters both the memory and the reads.
  • Mixture of Experts: In MoE models like Qwen3.6 35B-A3B, only 3 of the 35 billion parameters work on each token. The model needs room for all weights but responds almost as fast as a small model.
  • System RAM as a backup: On a PC, part of the model can spill over into regular system RAM. That works, but it slows things down a lot, because DDR5 memory has only a fraction of a graphics card's bandwidth. On a Mac, CPU and GPU share the same fast memory. By default, though, macOS assigns only about two thirds of it to the graphics unit (about three quarters from 48 GB), and that's the value the check uses.
  • Fit: Up to 60% memory use, a model counts as "Roomy", up to 85% as "Fits", and up to 98% as "Tight". The last 2% stay free, because completely full memory doesn't load in practice. If the space isn't enough, the check keeps halving the context length until the model fits (down to 2K).

How the Recommendation Score Comes Together

The score from 0 to 100 combines four sub-scores. Quality comes from a model's size, age, quantization, and benchmark results, plus speed, fit, and context.

How much each sub-score counts depends on the use case:

Use caseGeneral
Quality45%
Speed30%
Fit15%
Context10%
Use caseCoding
Quality50%
Speed20%
Fit15%
Context15%
Use caseReasoning
Quality55%
Speed15%
Fit15%
Context15%
Use caseChat
Quality40%
Speed35%
Fit15%
Context10%
Use caseImage input
Quality50%
Speed20%
Fit15%
Context15%

The basic idea comes from the open source tool llmfit by Alex Jones. In some places, the check calculates more finely. It rates size and speed on a continuous scale so that tiny models don't win automatically on small machines, and benchmarks count relative to the median of all models in the directory.

The formula is matched against 21 published measurements from the llama.cpp benchmark threads, from the RTX 3060 through several M chips to the DGX Spark. For classic models, the estimate stays within 15% of the measured value there, for MoE models within 30%.

That said:

It's still an estimate. Your actual speed also depends on the tool, your drivers, and the length of your chat. The model also has to read long prompts and documents before it answers, and that waiting time isn't part of the calculation. Models without a published parameter count are missing from the check, because their memory needs can't be calculated.

For some MoE models, the developer only publishes the total parameter count, for example MiniMax M2.5 or GLM-5.2. The check then plays it safe and calculates as if all parameters were active. In practice, these models usually respond much faster than the table shows.

5. Using Open Source LLMs Locally on Your Own Computer

Using open source LLMs locally on your own computer is easier than you might think:

1. Download LM Studio

Download LM Studio from the website. It's free and available for Mac, Windows, and Linux:

LM Studio

2. Install and Open LM Studio

Next, install LM Studio on your computer and open it.

3. Download Your Desired Open Source LLMs

Now you need to download the open source LLMs you want to use in LM Studio.

Many popular LLMs are already on the home screen. To download an LLM, simply click the blue download button:

Download open source LLMs

To find specific open source LLMs, you can also use the search function:

Search open source LLMs

4. Important: Check System Requirements Before Downloading

Before downloading an LLM, you should check the system requirements.

The hardware check in section 4 shows whether a model fits on your computer at all. LM Studio also displays the requirements right at download.

Llama 3, for example, requires more than 8 GB RAM and 4.92 GB of free storage:

Open source LLM system requirements

5. Chat with the Open Source LLM

After downloading an open source LLM, you can use it directly in LM Studio.

Simply click on the speech bubble icon (?) in the left sidebar.

The user interface and settings options are reminiscent of the OpenAI Playground:

Chat with open source LLM

Frequently Asked Questions About Open Source LLMs

Models under MIT or Apache 2.0 can be used commercially without extra conditions. MIT covers DeepSeek V4 Pro, GLM-5.3-Flash, and MiMo-V2.6-Pro, for example, while Qwen3.8 27B and the Gemma 4 family use Apache 2.0. GLM-5.3 and Kimi K3 also allow commercial use but require a separate agreement or a review from very large Model-as-a-Service providers. MiniMax M2.7, on the other hand, is licensed for non-commercial use only. Check the specific license again before every deployment.

Memory is what matters most. A model, plus a buffer, has to fit into your GPU's video memory or, on a Mac, into the share of unified memory that macOS assigns to the graphics unit. In the common Q4_K_M quantization, a model with 8 billion parameters needs about 6 GB. A 30B model needs about 21 GB, and a 70B model about 46 GB. Speed then depends mainly on memory bandwidth, which is why graphics cards are usually much faster than Macs with the same amount of memory. The hardware check in this article shows which models run on your computer and how fast they respond.

All three are large Mixture-of-Experts models for coding and agents. In the independent Vals AI measurement, DeepSeek V4 Pro 0813 leads on SWE-bench Verified with 96.4%, followed by GLM-5.3 at 95.4% and Kimi K3 at 93.4%. With 2.8 trillion parameters, Kimi K3 is the largest and the only one of the three that also understands images and video. DeepSeek V4 Pro uses the MIT license, while GLM-5.3 and Kimi K3 have their own licenses with conditions for very large providers.

Several user-friendly tools significantly simplify local LLM usage:

  • Ollama: Easiest installation, supports all common models
  • LM Studio: Graphical user interface, ideal for beginners
  • GPT4All: Lightweight solution for consumer hardware
  • Jan: Open source ChatGPT alternative with local execution
  • vLLM: High-performance solution for production environments

The models themselves are free, but operating costs can be significant. Local use is cost-free after the hardware investment, but powerful GPUs cost $1,500-$15,000+. Power consumption for training and inference should not be underestimated. Managed API providers often offer free quotas, then charge fees similar to OpenAI/Anthropic. VPS hosting starts at $20/month for CPU-only, GPU servers cost significantly more. The true costs lie in hardware, electricity, and potential cloud usage.

In September 2026, most of the strongest open-weight models still come from Chinese labs such as DeepSeek, Z.ai, Moonshot AI, Alibaba, and Xiaomi. On SWE-bench Verified, the gap to the best proprietary models has shrunk to a few points. Mixture-of-Experts dominates because only a small share of the parameters works on each token. At the same time, licenses are getting more varied, since GLM-5.3, Kimi K3, and MiniMax M3 use their own licenses instead of MIT. From the West, Gemma 4, Inkling, and Nemotron 3 Ultra add open models under Apache 2.0 and OpenMDW.
FH

Finn Hillebrandt

AI Expert & Blogger

Finn Hillebrandt is the founder of Gradually AI, an SEO and AI expert. He helps online entrepreneurs simplify and automate their processes and marketing with AI. Finn shares his knowledge here on the blog in 50+ articles as well as through the AI Business Club.

Learn more about Finn and the team, follow Finn on LinkedIn, join his Facebook group for ChatGPT, OpenAI & AI Tools or do like 17,500+ others and subscribe to his AI Newsletter with tips, news and offers about AI tools and online business. Also visit his other blog, Blogmojo, which is about WordPress, blogging and SEO.