Language modelOpen weights
Granite 4.2 30B
IBM
- Released
- August 25, 2026
- Data date
- October 3, 2026
Granite 4.2 30B receives a second supervised fine-tuning phase specifically for agentic coding. IBM increases the share of code and software-engineering data while retaining part of the original training mix. Like 8B, it then completes the agent block for software, terminal, and search tasks.
At 30 billion parameters, it is the largest of the three Granite 4.2 variants. IBM releases the weights under Apache 2.0.
Page 1 of 2
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
SWE-bench Multilingual (Show measurement, test conditions, and source)
- Source value
- 41.89
- Score
- 41.89
- Metric
- SWE-bench Multilingual reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Pro (Show measurement, test conditions, and source)
- Source value
- 33.29
- Score
- 33.29
- Metric
- SWE-Pro reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SWE-Verified (Show measurement, test conditions, and source)
- Source value
- 57
- Score
- 57
- Metric
- SWE-Verified reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Terminal-Bench 2.1 (Show measurement, test conditions, and source)
- Source value
- 29.24
- Score
- 29.24
- Metric
- Terminal-Bench 2.1 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
tau3-bench (Show measurement, test conditions, and source)
- Source value
- 62
- Score
- 62
- Metric
- tau3-bench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BFCL v4 (Show measurement, test conditions, and source)
- Source value
- 61.39
- Score
- 61.39
- Metric
- BFCL v4 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
ProfBench (Show measurement, test conditions, and source)
- Source value
- 42.9
- Score
- 42.9
- Metric
- ProfBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BirdBench (Show measurement, test conditions, and source)
- Source value
- 41.85
- Score
- 41.85
- Metric
- BirdBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GDPval (Show measurement, test conditions, and source)
- Source value
- 1225
- Score
- 1225
- Metric
- GDPval reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
AIME25 (Show measurement, test conditions, and source)
- Source value
- 89.17
- Score
- 89.17
- Metric
- AIME25 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GPQA (Show measurement, test conditions, and source)
- Source value
- 66.41
- Score
- 66.41
- Metric
- GPQA reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
LCBench v6 (Show measurement, test conditions, and source)
- Source value
- 75.77
- Score
- 75.77
- Metric
- LCBench v6 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | IBM Granite 4.2 announcement (retrieved October 3, 2026) · IBM Granite 4.2 announcement · Editorial description reviewed October 4, 2026 |
| Additional source | IBM Granite 4.2 technical report (retrieved October 4, 2026) · IBM Granite 4.2 technical report · Editorial description reviewed October 4, 2026 |