Gemini 3.8 Flash Cyber
Google DeepMind
- Released
- September 2, 2026
- Data date
- October 3, 2026
Gemini 3.8 Flash Cyber is Google’s cyber specialist variant for autonomous vulnerability discovery and automated patching. Google describes a dedicated focus on trusted defenders and security-critical codebases. It should not be treated as the generally available Gemini 3.8 Flash.
Access is provided through Google’s Fairwind Program for trusted defenders, including government authorities, critical infrastructure operators, and software maintainers. A general Gemini API release for the cyber variant is not evidenced.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Cyber specialist model for authorized security workSource |
| Access | Fairwind Program with prioritized access for trusted defendersSource |
| Capabilities | Autonomous vulnerability discovery and automated patchingSource |
| Availability | No general Gemini API release for this cyber variant is evidencedSource |
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
CyberGym final-submission pass@1 (Show measurement, test conditions, and source)
- Source value
- 86.2
- Score
- 86.2
- Metric
- CyberGym final-submission pass@1
- Unit
- %
- Category
- cybersecurity
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Gemini 3.8 Flash Cyber evaluation methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. September 2026; C/C++ vulnerability discovery, final-submission pass@1; single attempt, internal non-cyber-specialized Antigravity harness; Gemini API default sampling.
- Context
- Google's self-computed agent-harness result, not an isolated unassisted model score or a shared comparison cohort. Exact run date, sample count, and thinking level are unspecified.
CWE-Bench pass@1 (Show measurement, test conditions, and source)
- Source value
- 47.2
- Score
- 47.2
- Metric
- CWE-Bench pass@1
- Unit
- %
- Category
- cybersecurity
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Gemini 3.8 Flash Cyber evaluation methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. September 2026; Collinear AI evaluation with Antigravity agent harness and high thinking; pass@1, Gemini API default sampling.
- Context
- Independent benchmark-owner evaluation quoted by Google; the original public leaderboard record was not retrieved directly. Exact run date and sample count are unspecified.
Real-world vulnerability discovery recall (Show measurement, test conditions, and source)
- Source value
- 71
- Score
- 71
- Metric
- Real-world vulnerability discovery recall
- Unit
- %
- Category
- cybersecurity
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Gemini 3.8 Flash Cyber evaluation methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. September 2026; internal recall benchmark with over 1,200 confirmed historical vulnerabilities in open-source projects across 20 languages; internal non-cyber-specialized Antigravity harness, Gemini API default sampling.
- Context
- Vendor-reported internal dataset, not a public reproducible benchmark. Exact sample count, run date, and thinking level are unspecified; the rounded lower bound is not stored as a sample count.
Gray Swan IPI ASR@15 (Show measurement, test conditions, and source)
- Source value
- 6
- Score
- 6
- Metric
- Gray Swan IPI ASR@15
- Unit
- %
- Category
- cybersecurity
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Gemini 3.8 Flash Cyber evaluation methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. September 2026; Gray Swan independent evaluation, combined attack set, transfer-only, no computer use; attack success within 15 attempts, Gemini API default sampling.
- Context
- Google quotes an independently conducted Gray Swan evaluation; the evaluator's original result was not retrieved directly. This is attack success, so lower is better, and it is not pass@1.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 3 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Not published | wiz-penetration-test-protocol Original wording The current model page shows a Wiz Penetration Test value for the exact model without metric definition, direction, harness, or sampling details. No benchmark value is assigned from that chart. |
| Additional source | Google Gemini 3.8 Flash Cyber announcement (retrieved October 3, 2026; October 4, 2026) · Google Gemini 3.8 Flash Cyber announcement · Editorial description reviewed October 3, 2026 |
| Additional source | Gemini 3.8 Flash Cyber evaluation methodology (retrieved October 4, 2026) |
| Additional source | Model metadata (retrieved October 4, 2026) |