Ternary Bonsai 2 27B
PrismML
- Released
- September 17, 2026
- Data date
- October 3, 2026
Ternary Bonsai 2 27B compresses Qwen3.8-27B into ternary weights. According to PrismML, the densely packed PTQ1_0 files need 5.95 GB for the language model and store 1.75 bits per weight. A separately supplied vision component is added for image input.
The full checkpoint has 27.36 billion parameters and handles 262,144 context tokens. It retains the base model’s architecture. PrismML releases the weights under Apache 2.0, with the idealized ternary precision of 1.72 bits differing from the shipped PTQ1_0 format.
Page 1 of 2
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
MMLU-Redux (Show measurement, test conditions, and source)
- Source value
- 89.09
- Score
- 89.09
- Metric
- MMLU-Redux
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
GPQA Diamond (Show measurement, test conditions, and source)
- Source value
- 85.76
- Score
- 85.76
- Metric
- GPQA Diamond
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
GSM8K (Show measurement, test conditions, and source)
- Source value
- 96.66
- Score
- 96.66
- Metric
- GSM8K
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
AIME25 (Show measurement, test conditions, and source)
- Source value
- 95
- Score
- 95
- Metric
- AIME25
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
AIME26 (Show measurement, test conditions, and source)
- Source value
- 95.83
- Score
- 95.83
- Metric
- AIME26
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
HumanEval+ (Show measurement, test conditions, and source)
- Source value
- 95.12
- Score
- 95.12
- Metric
- HumanEval+
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
LiveCodeBench v6 (Show measurement, test conditions, and source)
- Source value
- 90.07
- Score
- 90.07
- Metric
- LiveCodeBench v6
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
IFEval (Show measurement, test conditions, and source)
- Source value
- 91.31
- Score
- 91.31
- Metric
- IFEval
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
IFBench (Show measurement, test conditions, and source)
- Source value
- 74
- Score
- 74
- Metric
- IFBench
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
tau2-Bench (Show measurement, test conditions, and source)
- Source value
- 80.22
- Score
- 80.22
- Metric
- tau2-Bench
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
BFCL v3 (Show measurement, test conditions, and source)
- Source value
- 74.92
- Score
- 74.92
- Metric
- BFCL v3
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
OCR Bench v2 (Show measurement, test conditions, and source)
- Source value
- 56.88
- Score
- 56.88
- Metric
- OCR Bench v2
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- PrismML whitepaper Tables 7, 8, 10 and methodology
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. NVIDIA H100, EvalScope + vLLM, TP2/DP4, thinking mode, xhigh, top-p 0.95/top-k 20, temperature 1.0, reasoning parser strips thought block. Most tasks single-pass; GPQA mean 5 samples/question; AIME25/26 mean 8/problem; vision tower loaded for vision tasks. Twenty-benchmark suite plus Terminal-Bench 2.1 and SWE-Bench Verified.
- Context
- Source-specific observation; it is not a shared comparison cohort. PrismML-run EvalScope/vLLM comparison; common harness across baselines and Bonsai, not external leaderboard
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | PrismML Bonsai 2 27B announcement (retrieved October 3, 2026) · PrismML Bonsai 2 27B announcement · Editorial description reviewed October 4, 2026 |
| Additional source | Ternary Bonsai 2 model card (retrieved October 3, 2026) · Ternary Bonsai 2 model card · Editorial description reviewed October 4, 2026 |
| Additional source | PrismML whitepaper Tables 7, 8, 10 and methodology (retrieved October 4, 2026) |