Skip to main content

Gemini Models: All Google Models at a Glance

All Google Gemini models compared: From Gemini 1.0 to Gemini 3.8 Flash Cyber with prices, context windows, and use cases.

FHFinn Hillebrandt
AI Technology
Gemini Models: All Google Models at a Glance
Links marked with * are affiliate links. If a purchase is made through such links, we receive a commission.

15 models compared across 4 generations. Some of them free. Some of them genuinely excellent. And some already discontinued.

The Gemini model lineup has gotten confusing fast since Google launched the first version in December 2023. Pro, Flash, Flash-Lite, Nano, Ultra. What does each one do? Which one should you actually use?

I use Gemini models regularly (mostly via the API and Antigravity CLI), and I keep this overview updated as new models drop. Here is everything you need to know about features, prices, and availability of all ChatGPT and Claude competitor models from Google.

TL;DRKey Takeaways
  • Gemini 3.8 Flash is Google's current Flash generation, with a 1,048,576-token context window, up to 65,536 output tokens, and pricing of $0.75 / $3.75 per million tokens through the end of 2026
  • According to Google, Gemini 3.8 Flash improves Terminal-Bench 2.1 to 90.8% and SWE-Bench Pro to 61.6%; on Humanity's Last Exam, its 45.4% trails 3.7 Flash slightly
  • Gemini 3.8 Flash Cyber is a separate defensive-cybersecurity variant, available only to vetted defenders through the Fairwind program
  • In the independent Vals AI comparison, Gemini 3.1 Pro remains ahead on reasoning; that chart uses an older, consistent model cohort and is not interchangeable with Google's vendor figures for 3.8 Flash
  • Gemini 2.5 Flash-Lite remains the cheapest powerful LLM on the market; the newer Gemini 3.5 Flash-Lite offers more quality for high-throughput tasks at a somewhat higher price
  • All modern Gemini models (from 1.5) are natively multimodal and process text, images, audio, and video simultaneously with up to 1 million token context

What are Gemini Models?

Gemini models are Google's advanced Large Language Models developed by DeepMind and Google Research.

What makes Gemini different? A few things stand out immediately:

First: Native multimodality from the start. Google trained Gemini with text, images, audio, and video, not like other providers who patched that in later. This gives Gemini a much deeper understanding of all these modalities together.

Then there's the context window: Gemini 3.1 Pro processes up to 1 million tokens. That's approximately 700,000 words or over 1,400 book pages. In a single request. That's... very large.

Google also doesn't have a one-model strategy. Instead: Nano for smartphones, Flash for most standard tasks, Pro for demanding stuff. Each has its place. And because Gemini is deeply integrated into Google Search, Workspace, and Android, it works particularly well there.

Google has taken a different approach with Gemini than OpenAI: Instead of focusing on maximum benchmark performance, the focus is on practical versatility, multimodality, and integration into the Google ecosystem.

Before we look at the individual models, here are the key milestones of Gemini's evolution from 2023 to today.

December 2023
Gemini 1.0 and Gemini Nano
Google's first attempt with a 32,000 token context, still text-only; Nano brings on-device AI to smartphones
February 2024
Gemini 1.5 Pro
First model with a 2 million token context, a world record at the time
May 2024
Gemini 1.5 Flash
The cheaper, faster variant of 1.5 Pro
December 2024
Gemini 2.0 Flash
1 million token context and a free API tier with rate limits
March 2025
Gemini 2.5 Pro
Premium model with a 1 million token context, experimentally up to 2 million
April 2025
Gemini 2.5 Flash
90% of Pro performance at a fraction of the cost
June 2025
Gemini 2.5 Flash-Lite
The cheapest powerful LLM on the market ($0.10 / $0.40 per million tokens)
November 2025
Gemini 3 Pro
Third generation of the premium model with frontier intelligence and Deep Research
December 2025
Gemini 3 Flash
Frontier intelligence at Flash pricing, 3x faster than 2.5 Pro and the default model in the Gemini app
February 2026
Gemini 3.1 Pro
Strongest reasoning model with 94.3% on GPQA Diamond and 80.6% on SWE-bench Verified
May 2026
Gemini 3.5 Flash
The newest Flash generation at the time, wins 11 of 15 benchmarks against 3.1 Pro and runs roughly 4x faster
July 2026
Gemini 3.6 Flash & Flash-Lite
More token-efficient Flash generation (17% fewer output tokens) plus a new high-throughput Lite model
August 2026
Gemini 3.7 Flash
Stable multimodal Flash model with a 1M context window, 64K output, and introductory pricing of $0.75 / $3.75
September 2026
Gemini 3.8 Flash
New Flash generation with stronger coding and agentic performance, a 1M context window, and 64K output
September 2026
Gemini 3.8 Flash Cyber
Specialized cybersecurity variant for vetted defenders in the Fairwind program

Comparison of All Gemini Models

Here's a detailed overview of all Gemini models with their key properties:

Column groups:
Model
Release
Status
Input
Output
Cache Input
Gemini 2.0 Flash12/2024Discontinued$0.1$0.4$0.03
Gemini 2.5 Flash-Lite06/2025Active$0.1$0.4$0.01
Gemini 2.5 Flash04/2025Active$0.3$2.5$0.03
Gemini 2.5 Pro03/2025Active$1.25 ≤200K / $2.5 >200K$10 ≤200K / $15 >200K$0.13 ≤200K / $0.25 >200K
Gemini 3 Flash12/2025Preview$0.5$3$0.05
Gemini 3 Pro11/2025Discontinued$2 ≤200K / $4 >200K$12 ≤200K / $18 >200K$0.2 ≤200K / $0.4 >200K
Gemini 3.1 Pro02/2026Preview$2 ≤200K / $4 >200K$12 ≤200K / $18 >200K$0.2 ≤200K / $0.4 >200K
Gemini 3.5 Flash05/2026Active$1.5$9$0.15
Gemini 3.6 Flash07/2026Active$0.75$3.75$0.07
Gemini 3.7 Flash08/2026Active$0.75$3.75$0.07
Gemini 3.8 Flash09/2026Active$0.75$3.75$0.07
Gemini 3.8 Flash Cyber09/2026Limited access———
Gemini 3.5 Flash-Lite07/2026Active$0.3$2.5$0.03
Gemini Nano-112/2023Active———
Gemini Nano-205/2024Active———

Gemini 3.8 Flash

Released: September 2026

Gemini 3.8 Flash has been available as a stable API model since September 2, 2026. Google positions it for coding, agents, and knowledge-intensive workflows. It accepts text, images, video, audio, and PDF files and produces text.

The model keeps the 1,048,576-token context window and maximum output of 65,536 tokens from 3.7 Flash. Its documented knowledge cutoff is March 2026, although Google still lists January 2025 for some knowledge domains. Reasoning levels include low, medium, and high. It supports function calling, structured outputs, code execution, Google Search, URL context, file search, and computer use in preview.

In Google's own matched measurements, 3.8 Flash scores 90.8% on Terminal-Bench 2.1, 61.6% on SWE-Bench Pro, and 45.4% on Humanity's Last Exam. Under the same setup, 3.7 Flash scores 81.6%, 60.4%, and 45.7%. The largest published gain is therefore in terminal work, not in the broad knowledge benchmark. SWE-Bench Pro remains a secondary signal because some tasks are disputed.

  • Price through December 31, 2026: $0.75 input / $3.75 output per million tokens
  • Price from January 1, 2027: $1.50 input / $7.50 output per million tokens
  • Context caching through 2026: $0.075 per million cached input tokens
  • API string: gemini-3.8-flash

Gemini 3.8 Flash is the first Flash model to test when coding or agentic work is the priority. Google also notes that it may use more tokens than 3.7 Flash on some tasks, so existing production workflows should be evaluated with their own prompts.

Gemini 3.8 Flash Cyber

Released: September 2026

Gemini 3.8 Flash Cyber is Google's defensive-cybersecurity variant of Gemini 3.8 Flash. Google describes it as a model for vulnerability discovery and automated patching. It shares the foundational intelligence of Gemini 3.8 Flash, while its deployment and safety boundaries are tailored to cybersecurity work.

Access is limited to vetted defenders through the Fairwind program. Google has not published a separate public API price or context-window value for Flash Cyber, so it should not be compared with the regular API models in the pricing table.

Google reports 86.2% on CyberGym Pass@1 and 47.2% on CWE-Bench Pass@1 for Flash Cyber. It also cites more than 70% on an internal real-world vulnerability benchmark. These are vendor-reported figures from the Fairwind context and are not directly comparable with the Gemini Flash scores above because the evaluation setups differ.

Gemini 3.7 Flash

Released: August 2026

Gemini 3.7 Flash has been available as a stable API model since August 13, 2026. Google positions it for fast agentic workflows, coding, and complex multi-step tasks. It accepts text, images, video, audio, and PDF files and produces text.

The model provides a 1,048,576-token context window and up to 65,536 output tokens. Reasoning levels include low, medium, and high. It also supports function calling, structured outputs, code execution, Google Search, URL context, and computer use in preview.

Key features:

  • Context window: 1,048,576 tokens
  • Maximum output: 65,536 tokens
  • Knowledge cutoff: March 2026
  • Price through December 31, 2026: $0.75 input / $3.75 output per million tokens
  • Context caching: $0.075 per million cached input tokens
  • API string: gemini-3.7-flash

One cost-planning detail matters:

Google labels this as time-limited introductory pricing. According to the current pricing table, standard rates rise to $1.50 for input and $7.50 for output on January 1, 2027. The calculator and price index always show the currently valid price.

Gemini 3.6 Flash

Released: July 2026

Gemini 3.6 Flash arrived on July 21, 2026. It did not replace Gemini 3.5 Flash, but was tuned for greater token efficiency: 17% fewer output tokens according to the Artificial Analysis Index, plus fewer reasoning steps and tool calls in multi-step workflows. Gemini 3.7 Flash has sat above it as the newer Flash generation since August 2026.

Google has since reduced the price. Through the end of 2026, Gemini 3.6 Flash costs $0.75 for input and $3.75 for output per million tokens. Google lists $1.50 / $7.50 from January 1, 2027. Its knowledge cutoff also moved forward significantly from January 2025 to March 2026.

Key Features:

  • More token-efficient: 17% fewer output tokens than 3.5 Flash, fewer reasoning steps and tool calls
  • Agentic coding: 58.7% on SWE-Bench Pro (3.5 Flash 55.1%, 3.1 Pro 54.2%)
  • Computer use: 83% on OSWorld-Verified (up from 78.4%)
  • DeepSWE: 49% (up from 37%), MLE-Bench: 63.9%
  • 1 million token context, knowledge cutoff March 2026 (previously January 2025)
  • Price through the end of 2026: $0.75 input / $3.75 output per million tokens
  • Context caching: $0.075 for cached inputs
  • API string: gemini-3.6-flash

What makes Gemini 3.6 Flash special?

Gemini 3.6 Flash is not a pure benchmark jump, it is an efficiency generation: at comparable or slightly better quality, the model needs noticeably fewer output tokens and tool calls to reach a result. In agentic workflows with many intermediate steps, that shows up directly in the price.

Availability: Gemini 3.6 Flash has been available since July 21, 2026 in the Gemini app, for developers through Google Antigravity, AI Studio, Android Studio, and via the Gemini API and Vertex AI.

Gemini 3.5 Flash-Lite

Released: July 2026

Gemini 3.5 Flash-Lite is Google's newest Lite model, announced on July 21, 2026 and a step up from its predecessor 3.1 Flash-Lite (February 2026) as well as the older 2.5 Flash-Lite. It is built for high-throughput, low-latency tasks such as agentic search and document processing.

The quality gain over 3.1 Flash-Lite is substantial: on Terminal-Bench 2.1, the model reaches 54% instead of 31%, a 23-point jump. It also outpaces several larger predecessor models on coding and computer-use tasks.

Key Features:

  • Terminal-Bench 2.1: 54% (3.1 Flash-Lite: 31%)
  • Agentic coding: 54.2% on SWE-Bench Pro (Gemini 3 Flash: 49.6%)
  • Computer use: 74.0% on OSWorld-Verified (up from 65.1%)
  • Long context: 72.2% on GDM-MRCR v2
  • 1 million token context
  • Price: $0.30 input / $2.50 output per million tokens
  • Context caching: $0.03 for cached inputs
  • API string: gemini-3.5-flash-lite

What makes Gemini 3.5 Flash-Lite special?

Gemini 3.5 Flash-Lite shows how fast the Lite lineup keeps moving: for high-throughput use cases like bulk classification or agents with many small intermediate steps, it delivers noticeably more quality than 3.1 Flash-Lite at a similarly low price. For the absolute cheapest option, Gemini 2.5 Flash-Lite remains the right pick.

Availability: Gemini 3.5 Flash-Lite has been available since July 21, 2026 in the Gemini app and through the Gemini API. Integration into Google Search has been announced but hasn't rolled out yet.

Gemini 3.5 Flash

Released: May 2026

Gemini 3.5 Flash is Google's agentic Flash generation, unveiled at Google I/O 2026 (May 19, 2026). Gemini 3.6 Flash has since joined it as an even more token-efficient Flash model, but 3.5 Flash remains fully available. It brings frontier performance on agentic tasks and beats Gemini 3.1 Pro on several coding and agent benchmarks, while costing a fraction.

Key Features:

  • Agentic coding: 76.2% on Terminal-Bench 2.1 (beats 3.1 Pro's 70.3%)
  • Tool use: 83.6% on MCP Atlas
  • Multimodal reasoning: 84.2% on CharXiv Reasoning
  • ~4x faster on output tokens than comparable frontier models
  • 1 million token input, up to 64,000 tokens output
  • Multimodal: text, images, audio, video, and PDF
  • Price: $1.50 input / $9 output per million tokens
  • Context caching: $0.15 for cached inputs (90% discount)
  • API string: gemini-3.5-flash

What makes Gemini 3.5 Flash special?

Gemini 3.5 Flash is the first Flash model that competes with Pro-tier models in agentic workflows. It wins 11 of 15 published benchmarks against Gemini 3.1 Pro, especially on tool calls, long-running tasks, and browser agents. On pure reasoning (Humanity's Last Exam, ARC-AGI-2), 3.1 Pro still leads, but for most coding and agent applications, 3.5 Flash is the better and significantly cheaper choice.

The API has also changed: instead of an integer `thinking_budget`, you now set a string enum `thinking_level` with values minimal, low, medium (the new default), and high. If you migrate from 3 Flash, set it explicitly to high or the model will think less than before.

Availability: Gemini 3.5 Flash has been generally available since May 19, 2026 through the Gemini API, Google AI Studio, Vertex AI, Google Antigravity, AI Mode in Google Search, and the Gemini app. Computer Use is not yet supported on 3.5 Flash, so stay on gemini-3-flash-preview for that. Gemini 3.5 Pro has been announced, but Google hasn't confirmed a release date yet.

Gemini 3.1 Pro

Released: February 2026

Gemini 3.1 Pro is Google's strongest pure-reasoning model and marks a massive leap over all predecessors. Since February 2026, it has been generally available through the Gemini API, Google AI Studio, and Vertex AI.

Key Features:

  • 2x reasoning leap over Gemini 3 Pro on complex tasks
  • 94.3% on GPQA Diamond (PhD-level reasoning), a new best score
  • 80.6% on SWE-bench Verified (agentic coding)
  • 1 million token input, up to 64,000 tokens output
  • Multimodal: text, images, audio, video, and PDF
  • Tiered API pricing: $2.00 / $12.00 per million tokens (under 200K context), $4.00 / $18.00 (over 200K context)
  • Context caching: 90% discount on cached input tokens ($0.20 / $0.40)
  • API string: gemini-3.1-pro

Availability: Gemini 3.1 Pro is generally available since February 2026 through the Gemini API in Google AI Studio, Vertex AI, and Gemini Enterprise. It is positioned as the premium model for demanding research, code, and analysis tasks.

Gemini 3 Flash

Released: December 2025

Gemini 3 Flash is Google's balanced model that combines frontier intelligence with high speed and low cost. Generally available since December 2025, it is the default model in the Gemini app.

Key Features:

  • Frontier performance: 90.4% on GPQA Diamond, 81.2% on MMMU Pro
  • Agentic coding: 78% on SWE-bench Verified
  • 3x faster than Gemini 2.5 Pro at comparable quality
  • 15% better than Gemini 2.5 Flash in overall accuracy
  • 1 million token input, up to 64,000 tokens output
  • Multimodal: text, images, audio, video, and PDF
  • Price: $0.50 input / $3.00 output per million tokens
  • Context caching: 90% cost reduction on cached tokens ($0.05 input)
  • API string: gemini-3-flash

Availability: Gemini 3 Flash is generally available through the Gemini API in Google AI Studio, Vertex AI, and Gemini Enterprise. It is the default model in the Gemini app and AI Mode in Google Search. Companies like JetBrains, Bridgewater Associates, and Figma use it in production.

Gemini 3 Pro

Released: November 2025

Gemini 3 Pro was the third generation of Google's premium AI model with frontier intelligence, Deep Research, and premium performance. Generally available from November 2025 to March 2026.

Key Features:

  • Frontier intelligence from Google DeepMind
  • Improved reasoning capabilities over Gemini 2.5 Pro
  • Deep Research for complex, multi-step analyses
  • Multimodal improvements especially in video understanding
  • 1 million token context window (input), up to 64,000 tokens output
  • Tiered API pricing: $2.00 / $12.00 per million tokens (under 200K context), $4.00 / $18.00 (over 200K context)
  • API string: gemini-3-pro

Availability: Gemini 3 Pro was superseded by Gemini 3.1 Pro and retired in March 2026. If you need premium performance today, go straight to Gemini 3.1 Pro.

Gemini 2.5 Pro

Released: March 2025

Gemini 2.5 Pro was Google's premium variant through late 2025 and remains a solid choice for demanding tasks. (More about the Gemini API in our separate guide.)

What does Pro offer specifically?

  • Top-tier performance on complex reasoning and code tasks
  • 1 million token context window (experimentally also 2 million)
  • Tiered pricing: $1.25 / $10 for standard prompts (≤ 200K tokens), $2.50 / $15 for longer ones
  • Native multimodality (process text, images, audio, video together)
  • Prompt caching with 90% discount on cached inputs ($0.125 / $0.25 instead of $1.25-2.50)
  • API model string: gemini-2.5-pro

What makes Gemini 2.5 Pro special?

Gemini 2.5 Pro is Google's answer to the then-current Claude 4 Opus and GPT-4o. It offers comparable performance on complex reasoning tasks and surpasses both competitors in processing very long contexts.

The 1 million token window enables analysis of complete books, large codebases, or hours of video transcripts in a single API call.

The tiered pricing structure is competitive: for most standard prompts you pay only $1.25 / $10, below Claude Opus 5.5 ($4 / $20) on both input and output.

Where can you get it? The Google AI API, Google AI Studio, Vertex AI, or Google Cloud.

When do you need Pro? When you want to analyze entire codebases, comb through long research papers, summarize thick contracts, or process hours of video in one shot. This isn't meant for chatbots. That's what Flash is for and it's cheaper.

Gemini 2.5 Flash

Released: April 2025

Gemini 2.5 Flash is the balanced variant, the model evergreen of the 2.5 series. It delivers 90% of Pro performance but costs a fraction and is significantly faster.

The key specs:

  • 90% of Pro performance at a fraction of the cost
  • 2-3x faster than Pro (inference speed)
  • 1 million token context
  • $0.30 input / $2.50 output per million tokens
  • Prompt caching: $0.03 for cached inputs
  • Multimodal: text, images, audio, video
  • API string: gemini-2.5-flash

What makes Gemini 2.5 Flash special?

Gemini 2.5 Flash is the ideal production model for 90% of all use cases. It offers nearly the same quality as Pro (90% performance) at 80% lower cost and 2-3x faster response time. This makes it perfect for chatbots, content generation, and automation workflows where fast responses matter more than absolute highest precision.

Compared to ChatGPT GPT-4o ($2.50 / $10 per million tokens), Gemini 2.5 Flash offers 75-88% cost savings at similar quality, a strong price-performance ratio.

You can find Flash via Google AI API, Google AI Studio, Vertex AI, Google Cloud, and it's also the backend model for many Google products.

Specific use cases: Chatbots that need to respond quickly. Content generation (articles, marketing copy, social posts). Data extraction from unstructured sources. Email classification, sentiment analysis, summaries. Screenshot understanding and OCR. For all this, you don't need Pro, Flash is sufficient and saves money.

Gemini 2.5 Flash-Lite

Released: June 2025

Gemini 2.5 Flash-Lite is what it says: The cheapest usable LLM on the market. And extremely fast at the same time.

The key numbers:

  • $0.10 input / $0.40 output per million tokens (cheapest on the market)
  • 5x faster than Pro models
  • Still 70-80% of Flash performance
  • 1 million token context
  • Prompt caching: $0.01 for cached inputs
  • Multimodal: text, images, audio, video
  • API string: gemini-2.5-flash-lite

Why is this so interesting? It's roughly a third cheaper than GPT-4o-mini and a fraction of the price of Claude 4.5 Haiku. And it's not slow. Rather the opposite.

The quality? 70-80% of Flash performance for chatbot responses, simple text generation, and classification. If you need millions of API calls daily, the cost savings are enormous.

Where can you find it? Google AI API, Google AI Studio, Vertex AI.

Use cases: Chatbots with millions daily. Large-scale content moderation. Sentiment analysis, categorization, tags. Real-time applications where low latency matters. Massive batch processing on a small budget.

Gemini 2.0 Flash

Released: December 2024

Gemini 2.0 Flash is the older version of Flash and was discontinued in June 2026. Its big advantage was the free API tier with rate limits.

Quick info (as of its discontinuation):

  • Free tier (rate limits: 15 req./min, 1,500/day, 1M/month) or paid tier: $0.10 / $0.40 per million tokens
  • ~80% of 2.5 Flash performance
  • 1 million token context
  • Multimodal: text, images, audio
  • API string: gemini-2.0-flash

If you still use 2.0 Flash in legacy projects, migrate to 2.5 Flash. For free prototyping, the free tier in Google AI Studio remains available.

Gemini 1.5 Pro

Released: February 2024

Gemini 1.5 Pro was a big deal in 2024: First model with 2 million token context. That was a world record at the time.

Today: It was shut down on April 30, 2025. If you're still using 1.5 Pro, migrate to 2.5 Pro or newer. Better performance, less hassle.

What 1.5 had: 2 million tokens (impressive back then). Native multimodality. Strong video and document analysis. But that was 2024.

Gemini 1.5 Flash

Released: May 2024

Gemini 1.5 Flash was basically the cheaper, faster version of 1.5 Pro. Also deprecated.

The facts: 1 million token context. Fast, low cost. Multimodal. But it went offline on April 30, 2025. Users should switch to 2.5 Flash or newer.

Gemini 1.0 Pro and Ultra

Released: December 2023

Gemini 1.0 was the first attempt. Today: No longer relevant.

What was it? 32,000 token context. Text-only, no images/videos. Pro was standard, Ultra was premium. Both are long gone. Google quickly replaced them with 1.5 and 2.x, much better models.

Gemini Nano

Released: December 2023 / May 2024

Gemini Nano is different: On-device AI for smartphones. Runs locally, no cloud.

What's important:

  • On-device: Directly on smartphones, no cloud call
  • Two variants: Nano-1 (text-only) and Nano-2 (multimodal)
  • 4,000 token context (small, but sufficient for smartphone tasks)
  • Privacy: Everything stays local
  • Hardware: Pixel smartphones, Samsung Galaxy S24+, other Android devices
  • Use cases: Smart reply, live transcription, offline translation, photo editing

Availability: Already integrated in various Android phones. Google rolls it out via system updates. Developers can use the AICore API.

Price Comparison of All Gemini Models

The following table shows a detailed overview of all Gemini prices (all figures in $ per million tokens). For a detailed analysis, we recommend our API cost calculator:

StatusActive
Input (Standard)$0.75
Output (Standard)$3.75
Input (Cached)$0.07
Output (Cached)—
ModelGemini 3.7 Flash
StatusActive
Input (Standard)$0.75
Output (Standard)$3.75
Input (Cached)$0.07
Output (Cached)—
ModelGemini 3.6 Flash
StatusActive
Input (Standard)$0.75
Output (Standard)$3.75
Input (Cached)$0.07
Output (Cached)—
ModelGemini 3.5 Flash-Lite
StatusActive
Input (Standard)$0.3
Output (Standard)$2.5
Input (Cached)$0.03
Output (Cached)—
ModelGemini 3.5 Flash
StatusActive
Input (Standard)$1.5
Output (Standard)$9
Input (Cached)$0.15
Output (Cached)—
ModelGemini 3.1 Pro Preview
StatusPreview
Input (Standard)$2 ≤200K / $4 >200K
Output (Standard)$12 ≤200K / $18 >200K
Input (Cached)$0.2 ≤200K / $0.4 >200K
Output (Cached)—
ModelGemini 3 Pro
StatusDiscontinued · 03/2026
Input (Standard)$2 ≤200K / $4 >200K
Output (Standard)$12 ≤200K / $18 >200K
Input (Cached)$0.2 ≤200K / $0.4 >200K
Output (Cached)—
ModelGemini 3 Flash Preview
StatusPreview
Input (Standard)$0.5
Output (Standard)$3
Input (Cached)$0.05
Output (Cached)—
ModelGemini 2.5 Pro
StatusActive
Input (Standard)$1.25 ≤200K / $2.5 >200K
Output (Standard)$10 ≤200K / $15 >200K
Input (Cached)$0.13 ≤200K / $0.25 >200K
Output (Cached)—
ModelGemini 2.5 Flash
StatusActive
Input (Standard)$0.3
Output (Standard)$2.5
Input (Cached)$0.03
Output (Cached)—
ModelGemini 2.5 Flash-Lite
StatusActive
Input (Standard)$0.1
Output (Standard)$0.4
Input (Cached)$0.01
Output (Cached)—
ModelGemini 2.0 Flash
StatusDiscontinued · 06/2026
Input (Standard)$0.1
Output (Standard)$0.4
Input (Cached)$0.03
Output (Cached)—

Important notes on the price table:

  • Gemini 2.5 Pro has tiered pricing: Lower prices for prompts ≤ 200,000 tokens ($1.25 / $10), higher prices for longer prompts (greater than 200,000 tokens: $2.50 / $15)
  • Gemini 3.1 Pro uses a two-tier model: $2.00 / $12.00 per million tokens under 200K context, $4.00 / $18.00 above 200K
  • Context caching (prompt caching) enables up to 90% discount on cached input tokens with repeated use. Example: Gemini 2.5 Flash input normally costs $0.30, cached only $0.03
  • Output prices for cached prompts remain the same as standard (no discount on output)

Gemini models in a consistent Vals AI evaluation setup

Models:
Gemini 3.1 Pro
Gemini 3 Pro
Gemini 3 Flash
Gemini 2.5 Pro
Gemini 2.5 Flash
Source: Vals AI, uniform cohort
gradually.ai

It gets even more interesting when you plot coding performance against API price and compare Gemini with Claude.

Price and performance in the same SWE-bench cohort
Ideal: strong + cheap
Google
Anthropic
OpenAI
DeepSeek
Moonshot AI
Efficiency frontier (best price-performance)
Sources: Vals AI, vendor price lists
gradually.ai

Frequently Asked Questions About Gemini Models

Gemini was developed by Google and offers native multimodal capabilities for text, images, audio, and video. Gemini 3.8 Flash accepts up to 1,048,576 input tokens and integrates closely with Google Search, Workspace, and Android. Gemini 2.5 Flash-Lite remains the family's cheapest API option at $0.10 / $0.40 per million tokens. ChatGPT has a more mature user community and a broader app ecosystem.

Gemini 3.1 Pro is the choice for particularly complex programming tasks, large codebases, and deep reasoning. For fast coding, agentic workflows, and frequent tool calls, Gemini 3.8 Flash is the current alternative. It offers roughly 1 million context tokens and costs $0.75 / $3.75 per million tokens through the end of 2026. For simple code formatting or syntax checks, Gemini 2.5 Flash-Lite is sufficient at $0.10 / $0.40.

Gemini 3.8 Flash, Gemini 3.7 Flash, and Gemini 3.6 Flash each cost $0.75 for input and $3.75 for output per million tokens through December 31, 2026. Cached input costs $0.075. Google lists $1.50 / $7.50 for Gemini 3.8 Flash from January 1, 2027. Google has not published a separate public API price for Gemini 3.8 Flash Cyber. The other current models are listed in the pricing table below. Gemini 3.1 Pro uses tiered pricing.

Gemini Nano is Google's on-device AI model that runs directly on smartphones and edge devices without a cloud connection. There are two variants: Nano-1 (4,000 token context, text-only) and Nano-2 (4,000 tokens, multimodal). Gemini Nano is used for privacy-sensitive tasks like offline translation, smart reply in messaging apps, live transcription, and photo editing directly on the device. It's already integrated in Pixel smartphones, Samsung Galaxy S24+, and other Android devices.

Yes, partially. Developers can test current Gemini models through Google AI Studio and the free API tier with rate limits (Gemini 2.0 Flash, which used to be the go-to free model, was discontinued in June 2026). For production applications, however, Google recommends the paid tiers without strict rate limits. The web version via google.com/gemini also offers free access with limitations.

Gemini Pro (currently 3.1 Pro) is the premium option for complex reasoning, coding, and analysis. Gemini Flash (currently 3.8 Flash) targets faster responses, agent workflows, and lower costs. Through the end of 2026, 3.8 Flash costs $0.75 / $3.75 per million tokens, while 3.1 Pro costs $2 / $12 below 200,000 tokens. Pro suits particularly demanding analysis, while Flash covers chatbots, coding, content, and automation.

Gemini 3.8 Flash supports 1,048,576 input tokens and up to 65,536 output tokens. Gemini 3.1 Pro, Gemini 3.7 Flash, Gemini 3.6 Flash, and the 2.5 family also process roughly 1 million context tokens. Gemini Nano for on-device use has only 4,000 tokens. Among older, retired models, Gemini 1.5 Pro supported up to 2 million tokens.

Current models include Gemini 3.8 Flash as the newest general release and Gemini 3.8 Flash Cyber for vetted defenders, plus 3.7 Flash, 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash, 3.1 Pro, 3 Flash, and the 2.5 family. Retired models include Gemini 3 Pro, 2.0 Flash, 1.5 Pro, and 1.5 Flash. Gemini Nano-1 and Nano-2 remain available for on-device use. The exact release dates are in the timeline further up this article.

Yes, all Gemini models from version 1.5 onwards are natively multimodal and can process images, videos, audio, and text simultaneously. Gemini can analyze images, recognize objects, understand screenshots, interpret diagrams, summarize videos, and even transcribe audio. This native multimodality (trained from the start with all modalities) distinguishes Gemini from competitors like GPT-4, which added text and vision separately. Gemini is particularly strong at analyzing long videos (up to 60 minutes).

Gemini supports four modalities simultaneously: text, images (PNG, JPEG, WEBP), audio (MP3, WAV, FLAC), and video (MP4, MOV). The model can: analyze and explain screenshots, extract code from images, convert diagrams to text, summarize long videos (up to 60 min.), transcribe and translate audio, compare multiple documents simultaneously, and process complex multimodal prompts (e.g., image + text + audio). This native multimodality makes Gemini ideal for content analysis, accessibility tools, and EdTech applications.

Gemini 3 is already available. Gemini 3 Flash arrived in December 2025, followed by Gemini 3.1 Pro in February 2026, Gemini 3.5 Flash in May, Gemini 3.6 Flash in July, Gemini 3.7 Flash in August, and Gemini 3.8 Flash on September 2, 2026. The Gemini 3 Pro model that launched alongside Gemini 3 Flash was retired in March 2026.

Gemini 2.5, released in March 2025, brought a clear jump in reasoning performance and inference speed over 1.5, along with improved code generation, more efficient context usage, and significantly better multimodality for video and audio. The pricing structure was simplified too, with prompt caching support. Gemini 1.5 was fully discontinued on April 30, 2025 and replaced by 2.5.

Gemini 2.5 Flash-Lite is the cheapest available model at $0.10 for input and $0.40 for output per million tokens, making it one of the most affordable powerful LLMs on the market. The newer Gemini 3.5 Flash-Lite (July 2026) delivers noticeably more quality but also costs more at $0.30 / $2.50, so 2.5 Flash-Lite remains the cheapest option. For free access, the free tier in Google AI Studio remains available (with rate limits). For comparison: GPT-4o-mini costs $0.15 / $0.60, Claude 4.5 Haiku $1 / $5. That makes Gemini 2.5 Flash-Lite roughly a third cheaper than GPT-4o-mini and a fraction of the price of Claude 4.5 Haiku.

Yes. Current Gemini models support context caching for reusable prompt content. With Gemini 3.8 Flash, cached input costs $0.075 instead of $0.75 per million tokens through the end of 2026. This is most useful for long system prompts, large documents, and repeated requests with the same context. Retention and storage charges depend on the cache configuration.

According to Google, Gemini 3.8 Flash scores 90.8% on Terminal-Bench 2.1 and 61.6% on SWE-Bench Pro. That puts it above 3.7 Flash at 81.6% and 60.4%, respectively. On Humanity's Last Exam, 3.8 Flash scores 45.4%, just below 3.7 Flash at 45.7%. Google also reports 86.2% on CyberGym Pass@1 and 47.2% on CWE-Bench Pass@1 for Gemini 3.8 Flash Cyber. These are vendor-reported figures from different evaluation setups, not one shared ranking. The chart in this article deliberately uses the independent Vals AI comparison with an older, consistent model cohort.
FH

Finn Hillebrandt

AI Expert & Blogger

Finn Hillebrandt is the founder of Gradually AI, an SEO and AI expert. He helps online entrepreneurs simplify and automate their processes and marketing with AI. Finn shares his knowledge here on the blog in 50+ articles as well as through the AI Business Club.

Learn more about Finn and the team, follow Finn on LinkedIn, join his Facebook group for ChatGPT, OpenAI & AI Tools or do like 17,500+ others and subscribe to his AI Newsletter with tips, news and offers about AI tools and online business. Also visit his other blog, Blogmojo, which is about WordPress, blogging and SEO.