Language modelResearch preview
Ember-1
Fireworks Research
- Released
- September 23, 2026
- Data date
- October 3, 2026
Ember-1 targets shorter reasoning from Kimi K3. Fireworks post-trained the checkpoint using task feedback and reports 35-50% shorter reasoning without loss of accuracy across seven benchmarks and two customer A/B tests. Fireworks serves Ember-1 through its own API as a time-limited research preview.
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
Terminal-Bench 2.1 (Show measurement, test conditions, and source)
- Source value
- 82
- Score
- 82
- Metric
- Terminal-Bench 2.1
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- fireworks.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Fireworks vendor table; agent-task setup details not sufficient for cross-source comparison.
- Context
- Source-specific observation; it is not a shared comparison cohort.
SWE-bench Verified (Show measurement, test conditions, and source)
- Source value
- 92.2
- Score
- 92.2
- Metric
- SWE-bench Verified
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- fireworks.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Fireworks vendor table; agent-task setup details not sufficient for cross-source comparison.
- Context
- Source-specific observation; it is not a shared comparison cohort.
SWE-Interact (Show measurement, test conditions, and source)
- Source value
- 20
- Score
- 20
- Metric
- SWE-Interact
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- fireworks.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Fireworks vendor table; agent-task setup details not sufficient for cross-source comparison.
- Context
- Source-specific observation; it is not a shared comparison cohort.
DeepSWE 1.1 (Show measurement, test conditions, and source)
- Source value
- 75.2
- Score
- 75.2
- Metric
- DeepSWE 1.1
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- fireworks.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Fireworks vendor table; agent-task setup details not sufficient for cross-source comparison.
- Context
- Source-specific observation; it is not a shared comparison cohort.
tau2 Airline (Show measurement, test conditions, and source)
- Source value
- 66
- Score
- 66
- Metric
- tau2 Airline
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- fireworks.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Fireworks vendor table; agent-task setup details not sufficient for cross-source comparison.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Fireworks Ember-1 announcement (retrieved October 3, 2026; October 4, 2026) · Fireworks Ember-1 announcement · Editorial description reviewed October 4, 2026 |