Language modelOpen weights
Naive-N0.5-Flash
NaiveAI
- Released
- -
- Data date
- October 3, 2026
Naive-N0.5-Flash combines sliding-window attention with DeepSeek sparse attention for a 1M-token context window. The open MoE checkpoint targets coding and AI research, activating 15.5 of its 309 billion parameters per token.
NaiveAI releases the weights under MIT. Its model card documents FP8 deployment on NVIDIA hardware.
Specifications and access
| Specification | Value and source |
|---|---|
| Model ID | NaiveAI/Naive-N0.5-FlashSource |
| Model class | Mixture-of-experts model for coding and AI researchSource |
| Parameters | 309 billion total, 15.5 billion activeSource |
| Context window | 1M tokens according to the model cardSource |
| Architecture | Hybrid of sliding-window attention and DeepSeek sparse attentionSource |
| License | MITSource |
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
MLE-bench-30 (Show measurement, test conditions, and source)
- Source value
- 73.7
- Score
- 73.7
- Metric
- MLE-bench-30
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort.
DeepSWE v1.1 (Show measurement, test conditions, and source)
- Source value
- 67.8
- Score
- 67.8
- Metric
- DeepSWE v1.1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Agents Last Exam (Show measurement, test conditions, and source)
- Source value
- 32.4
- Score
- 32.4
- Metric
- Agents Last Exam reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
TerminalBench 2.1 (Show measurement, test conditions, and source)
- Source value
- 86.7
- Score
- 86.7
- Metric
- TerminalBench 2.1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Bench Pro (Show measurement, test conditions, and source)
- Source value
- 73.6
- Score
- 73.6
- Metric
- SWE-Bench Pro reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
ProgramBench (Show measurement, test conditions, and source)
- Source value
- 17.5
- Score
- 17.5
- Metric
- ProgramBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
NL2Repo (Show measurement, test conditions, and source)
- Source value
- 71.9
- Score
- 71.9
- Metric
- NL2Repo reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
FrontierSWE v1 (Show measurement, test conditions, and source)
- Source value
- 78.2
- Score
- 78.2
- Metric
- FrontierSWE v1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
PostTrainBench (Show measurement, test conditions, and source)
- Source value
- 37.5
- Score
- 37.5
- Metric
- PostTrainBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
PaperBench (Show measurement, test conditions, and source)
- Source value
- 63.2
- Score
- 63.2
- Metric
- PaperBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SOL-ExecBench meanSOL (Show measurement, test conditions, and source)
- Source value
- 72.81
- Score
- 72.81
- Metric
- SOL-ExecBench meanSOL reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
NanoChat AutoResearch validation BPB (Show measurement, test conditions, and source)
- Source value
- 0.9051
- Score
- 0.9051
- Metric
- NanoChat AutoResearch validation BPB reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Official image-only card, visually inspected. Default Claude Code 2.1.207, 1M context, tool environment; proprietary agent/R&D harnesses.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Naive-N0.5-Flash model card (retrieved October 3, 2026; October 4, 2026) · Naive-N0.5-Flash model card · Editorial description reviewed October 4, 2026 |