Language modelOpen weights
Granite 4.2 8B
IBM
- Released
- August 25, 2026
- Data date
- October 3, 2026
For Granite 4.2 8B, IBM adds agent tasks in real coding, terminal, and search environments to foundational training. This block distinguishes 8B from Granite 4.2 3B. It trains multi-step tool use against the task’s actual outcome.
The dense 8-billion-parameter checkpoint sits between 3B and 30B. Its weights use Apache 2.0.
Page 1 of 2
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
SWE-bench Multilingual (Show measurement, test conditions, and source)
- Source value
- 30.78
- Score
- 30.78
- Metric
- SWE-bench Multilingual reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Pro (Show measurement, test conditions, and source)
- Source value
- 19.11
- Score
- 19.11
- Metric
- SWE-Pro reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Verified (Show measurement, test conditions, and source)
- Source value
- 47.67
- Score
- 47.67
- Metric
- SWE-Verified reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Terminal-Bench 2.1 (Show measurement, test conditions, and source)
- Source value
- 20.56
- Score
- 20.56
- Metric
- Terminal-Bench 2.1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
tau3-bench (Show measurement, test conditions, and source)
- Source value
- 58.06
- Score
- 58.06
- Metric
- tau3-bench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BFCL v4 (Show measurement, test conditions, and source)
- Source value
- 50.29
- Score
- 50.29
- Metric
- BFCL v4 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
ProfBench (Show measurement, test conditions, and source)
- Source value
- 41.2
- Score
- 41.2
- Metric
- ProfBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BirdBench (Show measurement, test conditions, and source)
- Source value
- 41.07
- Score
- 41.07
- Metric
- BirdBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GDPval (Show measurement, test conditions, and source)
- Source value
- 1189
- Score
- 1189
- Metric
- GDPval reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
AIME25 (Show measurement, test conditions, and source)
- Source value
- 86.67
- Score
- 86.67
- Metric
- AIME25 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GPQA (Show measurement, test conditions, and source)
- Source value
- 64.14
- Score
- 64.14
- Metric
- GPQA reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
LCBench v6 (Show measurement, test conditions, and source)
- Source value
- 73.24
- Score
- 73.24
- Metric
- LCBench v6 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | IBM Granite 4.2 announcement (retrieved October 3, 2026) · IBM Granite 4.2 announcement · Editorial description reviewed October 4, 2026 |
| Additional source | IBM Granite 4.2 technical report (retrieved October 4, 2026) · IBM Granite 4.2 technical report · Editorial description reviewed October 4, 2026 |