Language modelOpen weights
AREX-2
BAAI
- Released
- -
- Data date
- October 3, 2026
AREX-2 learns to improve a solution over multiple rounds. BAAI trains the 27B checkpoint to propose, measure, reflect, and revise. Verifiable feedback from machine-learning and programming tasks is intended to transfer this behavior to deep research.
The multimodal model is available as open weights under Apache 2.0. BAAI documents local deployment through Transformers.
Specifications and access
| Specification | Value and source |
|---|---|
| Model ID | BAAI/AREX-2Source |
| Model class | Multimodal agent model for long-horizon tasksSource |
| Parameters | 27 billionSource |
| Context window | 262,144 tokensSource |
| Architecture | Dense multimodal model with a Qwen3.8-compatible architectureSource |
| Access | Local weights with TransformersSource |
| License | Apache 2.0Source |
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
Frontier-CS (Show measurement, test conditions, and source)
- Source value
- 70.7
- Score
- 70.7
- Metric
- Frontier-CS reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. AREX technical report. Agent loop, tool/skills context, 5-12 hour task budgets; MLE score is mean over three seeds.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
MLE-Lite (Show measurement, test conditions, and source)
- Source value
- 81.8
- Score
- 81.8
- Metric
- MLE-Lite reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. AREX technical report. Agent loop, tool/skills context, 5-12 hour task budgets; MLE score is mean over three seeds.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BrowseComp (Show measurement, test conditions, and source)
- Source value
- 84
- Score
- 84
- Metric
- BrowseComp reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. AREX technical report. Agent loop, tool/skills context, 5-12 hour task budgets; MLE score is mean over three seeds.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
HLE text-only (Show measurement, test conditions, and source)
- Source value
- 52.6
- Score
- 52.6
- Metric
- HLE text-only reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. AREX technical report. Agent loop, tool/skills context, 5-12 hour task budgets; MLE score is mean over three seeds.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GAIA (Show measurement, test conditions, and source)
- Source value
- 92.2
- Score
- 92.2
- Metric
- GAIA reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. AREX technical report. Agent loop, tool/skills context, 5-12 hour task budgets; MLE score is mean over three seeds.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
DeepSearchQA F1 (Show measurement, test conditions, and source)
- Source value
- 93.8
- Score
- 93.8
- Metric
- DeepSearchQA F1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. AREX technical report. Agent loop, tool/skills context, 5-12 hour task budgets; MLE score is mean over three seeds.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | AREX-2 model card (retrieved October 3, 2026) · AREX-2 model card · Editorial description reviewed October 4, 2026 |
| Additional source | Technical report (retrieved October 4, 2026) |