Skip to main content

Head-to-head comparison

Gemini 3.5 Flash vs. DeepSeek-V4-Flash

Compare published benchmark results from matching versions and cohorts, API pricing, and technical specifications. Any documented test setup differences remain visible.

shared result series
31
distinct benchmarks
2
latest retrieval
2026-09-29

Change model selection

Specifications

General specifications

Release date, availability, context window, architecture, and published parameter counts at a glance. Undisclosed details remain marked as unavailable.

Specifications for Gemini 3.5 Flash and DeepSeek-V4-Flash
AttributeGemini 3.5 FlashDeepSeek-V4-Flash
ProviderGoogleDeepSeek
ReleasedMay 19, 2026Apr 24, 2026
AvailabilityActiveLimited access
Model typeProprietaryOpen Weights
ParametersNot published284B, 13B active
ArchitectureNot publishedMixture of Experts
Context window1,048,576 tokens1,000,000 tokens
Knowledge cutoffNot publishedNot published
Technical sourcesModelModel

Pricing

API pricing compared

The table compares official standard API rates per 1M tokens, including published cache pricing. Batch, fast, priority, regional surcharges, tool calls, and cloud platform pricing are excluded.

API pricing for Gemini 3.5 Flash and DeepSeek-V4-Flash
Price typeGemini 3.5 FlashDeepSeek-V4-Flash
API input$1.5$0.14
API output$9$0.4
Cache read$0.15Not listed
Cache write (5 min.)Not listedNot listed
Cache write (1 hr.)Not listedNot listed
Price verified09/22/2026 Source10/02/2026 Source
Gemini 3.5 FlashDeepSeek-V4-Flash
API inputUSD per 1M tokens, base tier
Gemini 3.5 Flash$1.5
DeepSeek-V4-FlashLower price$0.142
API outputUSD per 1M tokens, base tier
Gemini 3.5 Flash$9
DeepSeek-V4-FlashLower price$0.4

Key to the marks:better valuebetter, but not conclusivetie

Pricing by provider

Direct developer pricing and provider endpoint tariffs routed through OpenRouter are listed separately. Prices are in USD per 1M tokens. Context-dependent direct tiers plus regional, Flex, Priority, and other OpenRouter tariffs remain individually identifiable.

Provider pricing for Gemini 3.5 Flash and DeepSeek-V4-Flash
ModelProviderPurchase route and tariffInput per 1M tokensOutput per 1M tokensSource
Gemini 3.5 FlashGoogleDirect from developerStandard API$1.5$9Source09/22/2026
GoogleVia OpenRoutergoogle-vertex/global$1.5$9Source10/02/2026
GoogleVia OpenRoutergoogle-vertex/global/flexLowest listed price: $0.75Lowest listed price: $4.5Source10/02/2026
GoogleVia OpenRoutergoogle-vertex/global/priority$2.7$16.2Source10/02/2026
GoogleVia OpenRoutergoogle-vertex/us$1.65$9.9Source10/02/2026
Google AI StudioVia OpenRoutergoogle-ai-studio$1.5$9Source10/02/2026
Google AI StudioVia OpenRoutergoogle-ai-studio/flexLowest listed price: $0.75Lowest listed price: $4.5Source10/02/2026
Google AI StudioVia OpenRoutergoogle-ai-studio/priority$2.7$16.2Source10/02/2026
DeepSeek-V4-FlashAlibabaVia OpenRouteralibaba$0.176$0.528Source10/02/2026
AtlasCloudVia OpenRouteratlas-cloud/fp4$0.44$1.32Source10/02/2026
BaiduVia OpenRouterbaidu/fp8$0.44$1.32Source10/02/2026
BaseTenVia OpenRouterbaseten/fp8#1$0.13$0.26Source10/02/2026
BaseTenVia OpenRouterbaseten/fp8#2$0.13$0.26Source10/02/2026
CloudflareVia OpenRoutercloudflare$0.44$1.32Source10/02/2026
CohereVia OpenRoutercohere$0.14$0.28Source10/02/2026
CoreWeaveVia OpenRoutercoreweave/fp8$0.13$0.28Source10/02/2026
DeepInfraVia OpenRouterdeepinfra/fp8$0.06$0.18Source10/02/2026
DigitalOceanVia OpenRouterdigitalocean$0.119$0.238Source10/02/2026
GMICloudVia OpenRoutergmicloud/fp8$0.286$0.858Source10/02/2026
InceptronVia OpenRouterinceptron/fp4$0.05$0.65Source10/02/2026
MakoraVia OpenRoutermakora$0.09$0.195Source10/02/2026
Mancer 2Via OpenRoutermancer/fp8$0.2$0.6Source10/02/2026
MorphVia OpenRoutermorph/bf16$0.142$0.4Source10/02/2026
NextBitVia OpenRouternextbit/fp8$0.352$1.056Source10/02/2026
NovitaVia OpenRouternovita/fp8$0.4092$1.2276Source10/02/2026
OpenInferenceVia OpenRouteropen-inference/fp8Lowest listed price: $0.0038$1.6Source10/02/2026
ParasailVia OpenRouterparasail/fp8$0.14$0.28Source10/02/2026
PhalaVia OpenRouterphala$0.308$0.924Source10/02/2026
RekaVia OpenRouterreka$0.021$0.528Source10/02/2026
RelaceVia OpenRouterrelace/fp4$0.0051$1.28Source10/02/2026
Sail ResearchVia OpenRoutersail-research/fp4$0.019$0.3Source10/02/2026
Sail ResearchVia OpenRoutersail-research/us$0.019$0.42Source10/02/2026
SiliconFlowVia OpenRoutersiliconflow/fp8$0.22$0.66Source10/02/2026
StreamLakeVia OpenRouterstreamlake/fp8$0.044Lowest listed price: $0.132Source10/02/2026
TogetherVia OpenRoutertogether$0.14$0.28Source10/02/2026
VeniceVia OpenRoutervenice$0.175$0.35Source10/02/2026
WaferVia OpenRouterwafer/fast$0.12$0.7Source10/02/2026

Performance

Shared benchmarks

Only published results with the same benchmark version, task, metric, and comparison cohort are paired. Different reasoning levels, agents, harnesses, or output limits appear directly in the table.

Capability profile from matched benchmarks

Each axis averages directly matched benchmark families. A value of 100 means the stronger value within this pair, not a universal quality score.

255075100CodingReasoningMathematicsKnowledgeFinanceLegalCybersecurityHealth
Gemini 3.5 Flash
DeepSeek-V4-Flash
Coding: 100 / 98.2 (1 benchmark family)
Reasoning: 100 / 98 (1 benchmark family)
Mathematics: 100 / 96.9 (1 benchmark family)
Knowledge: 100 / 97.8 (1 benchmark family)
Finance: 100 / 97.7 (1 benchmark family)
Legal: 100 / 98.7 (1 benchmark family)
Cybersecurity: 97.2 / 94.7 (1 benchmark family)
Health: 100 / 98.5 (1 benchmark family)

31 matched rows from 2 benchmark families were considered. The radar shows 8 of 9 comparable categories.

Gemini 3.5 FlashDeepSeek-V4-Flash
CyberBench v1.1Patch. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner81.034 %
DeepSeek-V4-Flash72.414 %
LMArena Text, Style ControlOverall. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,476.958
DeepSeek-V4-Flash1,438.511
CyberBench v1.1PoC. Higher is better. No documented setup difference
Gemini 3.5 Flash57.627 %
DeepSeek-V4-FlashWinner61.017 %
LMArena Text, Style ControlBusiness, management and finance. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,467.855
DeepSeek-V4-Flash1,434.189
LMArena Text, Style ControlChinese. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,514.086
DeepSeek-V4-Flash1,479.924
LMArena Text, Style ControlCoding. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,507.378
DeepSeek-V4-Flash1,479.555
LMArena Text, Style ControlCreative writing. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,462.846
DeepSeek-V4-Flash1,410.703
LMArena Text, Style ControlEnglish. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,478.707
DeepSeek-V4-Flash1,449.94
LMArena Text, Style ControlEntertainment, sports and media. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,452.16
DeepSeek-V4-Flash1,408.063
LMArena Text, Style ControlExcluding ties. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,486.83
DeepSeek-V4-Flash1,436.008
LMArena Text, Style ControlExpert prompts. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,490.882
DeepSeek-V4-Flash1,466.306
LMArena Text, Style ControlFrench. Higher is better. No documented setup difference
Gemini 3.5 FlashWinner1,494.334
DeepSeek-V4-Flash1,459.263

Key to the marks:better valuebetter, but not conclusivetie

Benchmark scores for Gemini 3.5 Flash and DeepSeek-V4-Flash
Benchmark and taskGemini 3.5 FlashDeepSeek-V4-FlashSource
CyberBench v1.1Patch
Test details
No documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better
Winner81.034 %
72.414 %
Vals AI
Source details
Data as of: 08/03/2026
LMArena Text, Style ControlOverall
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 46,190 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 51,747 anonymous pairwise votes.
Winner1,476.958
95% confidence interval: ±3.954The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,438.511
95% confidence interval: ±4.115The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
CyberBench v1.1PoC
Test details
No documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better
57.627 %
Winner61.017 %
Vals AI
Source details
Data as of: 08/03/2026
LMArena Text, Style ControlBusiness, management and finance
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 9,088 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,641 anonymous pairwise votes.
Winner1,467.855
95% confidence interval: ±7.031The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,434.189
95% confidence interval: ±6.982The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlChinese
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 3,491 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,992 anonymous pairwise votes.
Winner1,514.086
95% confidence interval: ±10.738The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,479.924
95% confidence interval: ±11.42The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlCoding
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 13,289 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,239 anonymous pairwise votes.
Winner1,507.378
95% confidence interval: ±6.21The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,479.555
95% confidence interval: ±6.265The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlCreative writing
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 8,808 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,638 anonymous pairwise votes.
Winner1,462.846
95% confidence interval: ±7.404The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,410.703
95% confidence interval: ±7.575The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlEnglish
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 19,425 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 22,908 anonymous pairwise votes.
Winner1,478.707
95% confidence interval: ±5.278The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,449.94
95% confidence interval: ±5.323The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlEntertainment, sports and media
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 11,195 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,428 anonymous pairwise votes.
Winner1,452.16
95% confidence interval: ±6.734The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,408.063
95% confidence interval: ±6.896The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlExcluding ties
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 34,820 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 39,391 anonymous pairwise votes.
Winner1,486.83
95% confidence interval: ±5.234The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,436.008
95% confidence interval: ±5.353The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlExpert prompts
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 5,340 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,311 anonymous pairwise votes.
Winner1,490.882
95% confidence interval: ±8.777The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,466.306
95% confidence interval: ±8.867The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlFrench
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 1,472 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,966 anonymous pairwise votes.
Winner1,494.334
95% confidence interval: ±17.019The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,459.263
95% confidence interval: ±15.286The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlGerman
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 726 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 906 anonymous pairwise votes.
Winner1,485.2
95% confidence interval: ±22.822The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,433.595
95% confidence interval: ±20.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlHard prompts
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 31,069 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 34,700 anonymous pairwise votes.
Winner1,494.691
95% confidence interval: ±4.663The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,458.083
95% confidence interval: ±4.832The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlHard prompts, English
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 13,168 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,039 anonymous pairwise votes.
Winner1,490.87
95% confidence interval: ±6.176The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,464.61
95% confidence interval: ±6.133The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlInstruction following
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 16,713 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 18,051 anonymous pairwise votes.
Winner1,464.775
95% confidence interval: ±5.696The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,433.413
95% confidence interval: ±5.881The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlJapanese
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 639 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 635 anonymous pairwise votes.
Winner1,474.61
95% confidence interval: ±25.099The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,424.242
95% confidence interval: ±25.166The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlKorean
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 929 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 951 anonymous pairwise votes.
Winner1,435.275
95% confidence interval: ±20.203The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,378.153
95% confidence interval: ±20.872The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlLegal and government
Test details
No documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 3,917 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,107 anonymous pairwise votes.
No clear lead1,476.168
95% confidence interval: ±10.17The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,457.491
95% confidence interval: ±9.927The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlLife, physical and social science
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 7,683 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,509 anonymous pairwise votes.
Winner1,488.692
95% confidence interval: ±7.391The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,459.724
95% confidence interval: ±7.418The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlLonger queries
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 21,881 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,501 anonymous pairwise votes.
Winner1,480.524
95% confidence interval: ±5.441The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,448.636
95% confidence interval: ±5.577The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlMath
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 2,261 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,619 anonymous pairwise votes.
Winner1,495.757
95% confidence interval: ±12.768The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,441.412
95% confidence interval: ±11.893The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlMathematical professions
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 2,730 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,828 anonymous pairwise votes.
Winner1,482.844
95% confidence interval: ±11.978The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,445.003
95% confidence interval: ±11.761The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlMedicine and healthcare
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 3,429 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,774 anonymous pairwise votes.
Winner1,484.201
95% confidence interval: ±10.913The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,461.307
95% confidence interval: ±10.453The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlMulti-turn
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 7,746 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,233 anonymous pairwise votes.
Winner1,481.739
95% confidence interval: ±7.534The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,446.332
95% confidence interval: ±7.34The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlNon-English
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 26,754 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 28,837 anonymous pairwise votes.
Winner1,468.08
95% confidence interval: ±4.816The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,423.125
95% confidence interval: ±5.023The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlPolish
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 765 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,037 anonymous pairwise votes.
Winner1,510.58
95% confidence interval: ±22.117The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,434.977
95% confidence interval: ±18.457The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlRussian
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 5,188 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,790 anonymous pairwise votes.
Winner1,488.319
95% confidence interval: ±8.821The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,437.327
95% confidence interval: ±8.593The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlSoftware and IT services
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 18,620 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,329 anonymous pairwise votes.
Winner1,501.513
95% confidence interval: ±5.465The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,472.722
95% confidence interval: ±5.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlSpanish
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 1,388 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,497 anonymous pairwise votes.
Winner1,470.651
95% confidence interval: ±16.935The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,429.802
95% confidence interval: ±16.45The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026
LMArena Text, Style ControlWriting, literature and language
Test details
No documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control
DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style control
Gemini 3.5 Flash: Style-controlled Arena rating from 11,905 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 12,763 anonymous pairwise votes.
Winner1,472.184
95% confidence interval: ±6.479The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
1,422.277
95% confidence interval: ±6.549The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result.
LMArena
Source details
Data as of: 09/25/2026

Key to the marks:better valuebetter, but not conclusivetie

Our model tests

Our tests

Compare how the models respond to the same prompt. Each test shows the first attempt, with no subsequent fixes to the generated code. These results do not contribute to an overall score.

Pelican on a bicycle

A community classic for free-form SVG drawing.

Gemini 3.5 Flash

API access was unavailable for this run. The model was not evaluated.

Reasoning (requested): high

First attempt
Run date
7 Sept 2026
API provider
openrouter
Requested API model
google/gemini-3.5-flash
Reasoning setting
high
Pinned endpoint
Google AI Studio | google/gemini-3.5-flash-20260519
Temperature
Provider default
Token limit including reasoning
8,192 tokens
Test protocol
community-visual-tests-v1
Technical checks
Not run

Not yet visually reviewed. A successful recording does not confirm correct physics or full compliance with the prompt.

DeepSeek-V4-Flash

API access was unavailable for this run. The model was not evaluated.

Reasoning (requested): high

First attempt
Run date
7 Sept 2026
API provider
openrouter
Requested API model
deepseek/deepseek-v4-flash-0731
Reasoning setting
high
Pinned endpoint
DeepSeek | deepseek/deepseek-v4-flash-20260731
Temperature
Provider default
Token limit including reasoning
8,192 tokens
Test protocol
community-visual-tests-v1
Technical checks
Not run

Not yet visually reviewed. A successful recording does not confirm correct physics or full compliance with the prompt.

Prompt and test conditions

Generate an SVG of a pelican riding a bicycle

Token limit including reasoning: 8,192 tokens

Task origin (Simon Willison)

How to read this comparison

Winning one benchmark is not an overall verdict

Your workload, budget, and required context length matter most. A coding benchmark says little about visual understanding or agent performance.

The standard price table uses direct developer rates. The provider comparison labels OpenRouter endpoints, regions, and special tariffs separately.

Every result links to its measurement source and retrieval date. Labels distinguish vendor reports, official benchmark leaderboards, and independent evaluations. Disputed or archived results are excluded.

More matchups

Compare Gemini 3.5 Flash and DeepSeek-V4-Flash with other leading models using the same data and benchmark logic.

View the complete LLM comparison