Language modelOpen weights
Nex-N2.5-Max
Nex-AGI
- Released
- -
- Data date
- October 3, 2026
Nex-N2.5-Max marks Nex-AGI’s first complete post-training effort at trillion-parameter scale, according to the company. Its foundation is a text-only MoE model with 1.6 trillion parameters. Unlike Mini and Pro, Max does not accept visual input.
Nex-AGI publishes the weights under Apache 2.0. The release does not list a dedicated hosted Max endpoint.
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
TerminalBench 2.1 (Show measurement, test conditions, and source)
- Source value
- 86.1
- Score
- 86.1
- Metric
- TerminalBench 2.1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Pro (Show measurement, test conditions, and source)
- Source value
- 65.7
- Score
- 65.7
- Metric
- SWE-Pro reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
DeepSWE (Show measurement, test conditions, and source)
- Source value
- 65.6
- Score
- 65.6
- Metric
- DeepSWE reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Automation (Show measurement, test conditions, and source)
- Source value
- 50.2
- Score
- 50.2
- Metric
- Automation reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Toolathlon (Show measurement, test conditions, and source)
- Source value
- 74.7
- Score
- 74.7
- Metric
- Toolathlon reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GDPval (Show measurement, test conditions, and source)
- Source value
- 1713
- Score
- 1713
- Metric
- GDPval reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
JobBench (Show measurement, test conditions, and source)
- Source value
- 53.6
- Score
- 53.6
- Metric
- JobBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BrowseComp (Show measurement, test conditions, and source)
- Source value
- 92.6
- Score
- 92.6
- Metric
- BrowseComp reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Vendor table. Sampling temperature 0.7, top_p 0.95 and top_k 40; coding uses NexAU, computer-use tasks use NexCUA, and WebTest is oracle-mode defect detection.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Nex-N2.5-Max model card (retrieved October 3, 2026; October 4, 2026) · Nex-N2.5-Max model card · Editorial description reviewed October 4, 2026 |
| Additional source | Nex-N2.5-Pro model card (retrieved October 3, 2026) · Nex-N2.5-Pro model card · Editorial description reviewed October 4, 2026 |
| Additional source | Nex-N2.5-mini model card (retrieved October 3, 2026) · Nex-N2.5-mini model card · Editorial description reviewed October 4, 2026 |