DeepSeek-V3
Not tested yet
DeepSeek
DeepSeek explicitly describes V3-0324 as an improvement over its predecessor, DeepSeek-V3.
DeepSeek-V3-0324 model card (2026-09-09)No published category-wide ranking is available for this model.
Entries 1-12 of 121. Page 1 of 11.
Each chart compares models in the same documented test and highlights the current model.
By task type
Accuracy
What this benchmark measures: 33 automatically verifiable HAL tasks drawn from 214 realistic, time-consuming web research tasks.
Context: The published HAL subset is smaller than the full dataset. Reported costs do not account for caching benefits.
Scale: 0% - 100%
Accuracy
What this benchmark measures: 300 verified tasks across 136 dynamic live websites.
Context: SeeAct and Browser-Use are different agent scaffolds and therefore remain separate comparison cohorts.
Scale: 0% - 100%
Accuracy
What this benchmark measures: 65 scientific coding problems with 338 subproblems across 16 research fields.
Context: Zero Shot, Tool Calling, and HAL Generalist are not directly comparable scaffolds.
Scale: 0% - 100%
Accuracy
What this benchmark measures: 102 expert-validated tasks for data-driven scientific work.
Context: The scaffold, tools, and code execution are part of the result. HAL reports both 42 and 44 underlying publications in different places.
Scale: 0% - 100%
Accuracy
What this benchmark measures: A random, fully human-validated subset of 50 GitHub issues from 12 Python repositories.
Context: The small subset is cheaper but statistically less stable than the full 500 tasks. SWE-Agent and HAL Generalist remain separate.
Note (DeepSeek-V3): HAL reports min-max ranges for repeated runs. The central value is the reported mean, not a 95% confidence interval.
Note: HAL reports min-max ranges for repeated runs. The central value is the reported mean, not a 95% confidence interval.
Note: HAL reports min-max ranges for repeated runs. The central value is the reported mean, not a 95% confidence interval.
Note: HAL reports min-max ranges for repeated runs. The central value is the reported mean, not a 95% confidence interval.
Scale: 0% - 100%
Accuracy
What this benchmark measures: 307 algorithmic programming tasks from Bronze through Platinum with exhaustive tests.
Context: The benchmark measures competitive programming, not repository navigation, maintainability, or product experience.
Scale: 0% - 100%
Average across 29 languages
What this benchmark measures: 11,829 parallel questions per language compare knowledge and reasoning across 29 languages.
Context: Results use five-shot chain of thought and are not interchangeable with MMLU-Pro.
Scale: 0% - 100%
German
What this benchmark measures: 11,829 parallel questions per language compare knowledge and reasoning across 29 languages.
Context: Results use five-shot chain of thought and are not interchangeable with MMLU-Pro.
Scale: 0% - 100%
Overall
Scale: 0% - 100%
Overall
Scale: 0% - 100%
Overall
Scale: 0% - 100%
Overall
Scale: 0% - 100%
Every observation retains its source value and published test conditions.
Our model tests
Compare how the models respond to the same prompt. Each test shows the first attempt, with no subsequent fixes to the generated code. These results do not contribute to an overall score.
A community classic for free-form SVG drawing.
DeepSeek-V3
Not tested yet
Generate an SVG of a pelican riding a bicycle
Token limit including reasoning: 8,192 tokens
Task origin (Simon Willison)Profile
Published information about this model. Existing estimates are explicitly labeled.
Dated documentation and model-card observations. Provider limits, native context and extended context can differ. Configuration notes retain the source wording.
Pricing
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
Cost example
$0.077175
100 requests with 1,000 input and 500 output tokens each calculate to $0.077175. Input accounts for $0.02574; output accounts for $0.051435.
The calculation uses documented token prices and no cache discount. It excludes extra tools, tax, reasoning tokens, and further hidden output tokens.
Open sourceHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
There are no published direct comparisons for this model yet.
Evidence
Every statement links to its underlying documentation or leaderboard.
Base API prices without caching or batch discounts. Higher context tiers and other rates are listed under Costs.
Documented modalities. Retrieved September 8, 2026.
Documented modalities. Retrieved September 8, 2026.
Research date September 8, 2026. 2 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication.
Unresolved
release-date. No exact release-date field in checked card.
Not found in the inspected sources
knowledge-cutoff. No cutoff stated in card.