Language modelOpen weights
Granite 4.2 3B
IBM
- Released
- August 25, 2026
- Data date
- October 3, 2026
Granite 4.2 3B follows a shorter training schedule than 8B and 30B in IBM’s new series. The compact version receives reasoning and coding training but skips their additional agent block covering software, terminal, and search tasks.
IBM releases this dense 3-billion-parameter model under Apache 2.0. Like the larger variants, it can switch between answers with and without a thinking process.
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
tau3-bench (Show measurement, test conditions, and source)
- Source value
- 45.78
- Score
- 45.78
- Metric
- tau3-bench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BFCL v4 (Show measurement, test conditions, and source)
- Source value
- 52.41
- Score
- 52.41
- Metric
- BFCL v4 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
ProfBench (Show measurement, test conditions, and source)
- Source value
- 32.1
- Score
- 32.1
- Metric
- ProfBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
AIME25 (Show measurement, test conditions, and source)
- Source value
- 78.33
- Score
- 78.33
- Metric
- AIME25 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
GPQA (Show measurement, test conditions, and source)
- Source value
- 54.8
- Score
- 54.8
- Metric
- GPQA reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
LCBench v6 (Show measurement, test conditions, and source)
- Source value
- 69.71
- Score
- 69.71
- Metric
- LCBench v6 reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
SciCode (Show measurement, test conditions, and source)
- Source value
- 24.11
- Score
- 24.11
- Metric
- SciCode reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
MMLU-Pro (Show measurement, test conditions, and source)
- Source value
- 67.84
- Score
- 67.84
- Metric
- MMLU-Pro reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
IFBench Prompt (Show measurement, test conditions, and source)
- Source value
- 74.33
- Score
- 74.33
- Metric
- IFBench Prompt reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
RULER 64K (Show measurement, test conditions, and source)
- Source value
- 67.52
- Score
- 67.52
- Metric
- RULER 64K reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
RULER 128K (Show measurement, test conditions, and source)
- Source value
- 55.3
- Score
- 55.3
- Metric
- RULER 128K reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Provider model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. IBM official technical breakdown; agent/coding values and missing sampling details are not cross-source comparable.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | IBM Granite 4.2 announcement (retrieved October 3, 2026) · IBM Granite 4.2 announcement · Editorial description reviewed October 4, 2026 |
| Additional source | IBM Granite 4.2 technical report (retrieved October 4, 2026) · IBM Granite 4.2 technical report · Editorial description reviewed October 4, 2026 |