Language modelOpen weights
Nex-N2.5-mini
Nex-AGI
- Released
- -
- Data date
- October 3, 2026
Nex-N2.5-mini uses visual feedback to work through longer computer and browser tasks. Nex-AGI trains this compact multimodal variant on a broader set of environments and tasks than the previous Nex-N2 series.
Mini is available through OpenRouter alongside Pro. Nex-AGI also publishes local weights under Apache 2.0.
Page 1 of 2
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
TerminalBench 2.1 (Show measurement, test conditions, and source)
- Source value
- 73.4
- Score
- 73.4
- Metric
- TerminalBench 2.1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Pro (Show measurement, test conditions, and source)
- Source value
- 43.8
- Score
- 43.8
- Metric
- SWE-Pro reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
DeepSWE (Show measurement, test conditions, and source)
- Source value
- 36.1
- Score
- 36.1
- Metric
- DeepSWE reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Automation (Show measurement, test conditions, and source)
- Source value
- 32.3
- Score
- 32.3
- Metric
- Automation reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Toolathlon (Show measurement, test conditions, and source)
- Source value
- 54.6
- Score
- 54.6
- Metric
- Toolathlon reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GDPval (Show measurement, test conditions, and source)
- Source value
- 1446
- Score
- 1446
- Metric
- GDPval reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
JobBench (Show measurement, test conditions, and source)
- Source value
- 28.5
- Score
- 28.5
- Metric
- JobBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BrowseComp (Show measurement, test conditions, and source)
- Source value
- 83.4
- Score
- 83.4
- Metric
- BrowseComp reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
OSWorld Verified (Show measurement, test conditions, and source)
- Source value
- 71.2
- Score
- 71.2
- Metric
- OSWorld Verified reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
OSWorld 2 (Show measurement, test conditions, and source)
- Source value
- 30.5
- Score
- 30.5
- Metric
- OSWorld 2 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
WebTest (Show measurement, test conditions, and source)
- Source value
- 48.6
- Score
- 48.6
- Metric
- WebTest reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
WebArena (Show measurement, test conditions, and source)
- Source value
- 63.4
- Score
- 63.4
- Metric
- WebArena reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Nex-N2.5-mini model card (retrieved October 3, 2026; October 4, 2026) · Nex-N2.5-mini model card · Editorial description reviewed October 4, 2026 |
| Additional source | Nex-N2.5-Pro model card (retrieved October 3, 2026) · Nex-N2.5-Pro model card · Editorial description reviewed October 4, 2026 |
| Additional source | Nex-N2.5-Max model card (retrieved October 3, 2026) · Nex-N2.5-Max model card · Editorial description reviewed October 4, 2026 |