Language modelPreview
MAI-Thinking-1
Microsoft AI
- Released
- August 12, 2026
- Data date
- October 3, 2026
MAI-Thinking-1 was trained without distillation from other providers’ models, according to Microsoft. The company uses its own traceable training data and executable coding environments where real tests grade outcomes. The MoE model activates 35 billion of roughly 1 trillion parameters.
It has been available in public preview on Microsoft Foundry since August 12, 2026. Microsoft specifies a 256K-token context window.
Specifications and access
| Specification | Value and source |
|---|---|
| Release | August 12, 2026Source |
| Access | Public preview in Microsoft FoundrySource |
| Focus | Reasoning for mathematics, knowledge, and codingSource |
| Parameters | About 1 trillion total according to Microsoft, 35 billion active (mixture of experts)Source |
| Context window | 256K tokens according to MicrosoftSource |
| Capabilities | Function calling, developer instructions, and a Chat Completions-compatible APISource |
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
AIME 2025 (Show measurement, test conditions, and source)
- Source value
- 97
- Score
- 97
- Metric
- AIME 2025
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
AIME 2026 (Show measurement, test conditions, and source)
- Source value
- 94.5
- Score
- 94.5
- Metric
- AIME 2026
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
HMMT Feb 2026 (Show measurement, test conditions, and source)
- Source value
- 84.9
- Score
- 84.9
- Metric
- HMMT Feb 2026
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
GPQA Diamond (Show measurement, test conditions, and source)
- Source value
- 84.2
- Score
- 84.2
- Metric
- GPQA Diamond
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
LCB v6 (Show measurement, test conditions, and source)
- Source value
- 87.7
- Score
- 87.7
- Metric
- LCB v6
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
Terminal Bench 2.0 (Show measurement, test conditions, and source)
- Source value
- 46
- Score
- 46
- Metric
- Terminal Bench 2.0
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
SWE-Bench Verified (Show measurement, test conditions, and source)
- Source value
- 73.5
- Score
- 73.5
- Metric
- SWE-Bench Verified
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
SWE-Bench Pro (Show measurement, test conditions, and source)
- Source value
- 52.8
- Score
- 52.8
- Metric
- SWE-Bench Pro
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft announcement Table 1 pixels
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Post-trained model evaluation; percentages. Agentic coding uses 256K total context; others max output 256K. Separate blind human preference study has 1,276 tasks.
- Context
- Source-specific observation; it is not a shared comparison cohort. Microsoft self-reported table; peer numbers taken from respective official cards
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Microsoft AI MAI-Thinking-1 announcement (retrieved October 3, 2026; October 4, 2026) · Microsoft AI MAI-Thinking-1 announcement · Editorial description reviewed October 4, 2026 |