Language modelOpen source
Ling-3.0-tiny
InclusionAI
- Released
- -
- Data date
- October 3, 2026
Ling-3.0-tiny brings InclusionAI’s hybrid attention architecture to a 7.9B checkpoint with 1.3 billion active parameters. Three Kimi Delta Attention layers alternate with one Multi-Head Latent Attention layer. The provider documents local tests on DGX Spark and Apple Silicon Macs.
InclusionAI releases BF16, FP8, and INT4 weights under MIT. The model handles 262,144 context tokens and is also available through the OpenRouter route linked in its model card.
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
TAU3-Banking-AA (Show measurement, test conditions, and source)
- Source value
- 20.8
- Score
- 20.8
- Metric
- TAU3-Banking-AA
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
BFCL-v4 (FC) (Show measurement, test conditions, and source)
- Source value
- 62.72
- Score
- 62.72
- Metric
- BFCL-v4 (FC)
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
Terminal-Bench 2.1 (Show measurement, test conditions, and source)
- Source value
- 27.7
- Score
- 27.7
- Metric
- Terminal-Bench 2.1
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
ArtifactsBench (Show measurement, test conditions, and source)
- Source value
- 47.93
- Score
- 47.93
- Metric
- ArtifactsBench
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
SciCode (Show measurement, test conditions, and source)
- Source value
- 24.2
- Score
- 24.2
- Metric
- SciCode
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
GPQA Diamond (Show measurement, test conditions, and source)
- Source value
- 73.4
- Score
- 73.4
- Metric
- GPQA Diamond
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
HLE (Show measurement, test conditions, and source)
- Source value
- 9.3
- Score
- 9.3
- Metric
- HLE
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
IFBench (Show measurement, test conditions, and source)
- Source value
- 63.61
- Score
- 63.61
- Metric
- IFBench
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
Multi-IF (Show measurement, test conditions, and source)
- Source value
- 83.15
- Score
- 83.15
- Metric
- Multi-IF
- Unit
- %
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
GDPval v2-AA (Show measurement, test conditions, and source)
- Source value
- 772
- Score
- 772
- Metric
- GDPval v2-AA
- Unit
- elo
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- HF model-card table reproducing Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Terminal-Bench 2.1 uses AA protocol, Terminus 2 default harness, 2-hour timeout, JSON parser, preserve-thinking, 3 runs/task mean, temperature 1.0, max_new_tokens 32K, 256K context.
- Context
- Source-specific observation; it is not a shared comparison cohort. AA-protocol/provider-card evidence, not an exact committed independent record
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Ling-3.0-tiny model card (retrieved October 3, 2026; October 4, 2026) · Ling-3.0-tiny model card · Editorial description reviewed October 4, 2026 |
| Additional source | Ling-3.0-flash-VL research card (retrieved October 3, 2026) · Ling-3.0-flash-VL research card · Editorial description reviewed October 4, 2026 |
| Additional source | Ling-3.0-flash-Fin research card (retrieved October 3, 2026) · Ling-3.0-flash-Fin research card · Editorial description reviewed October 4, 2026 |