Language modelActive
Ling 3.1 Flash
InclusionAI
- Released
- -
- Data date
- October 3, 2026
Ling 3.1 Flash activates 25 billion of its 560 billion parameters per token. InclusionAI’s hybrid reasoning model handles 262K context tokens on Vercel AI Gateway and targets extended coding, document, and agent tasks.
Vercel offers the model free through October 13, 2026. The regular model identifier starts billing afterward, while the separate inclusionai/ling-3.1-flash-free identifier stops serving requests.
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
Artificial Analysis Intelligence Index (Show measurement, test conditions, and source)
- Source value
- 41.0906200397391
- Score
- 41.0906200397391
- Metric
- Artificial Analysis Intelligence Index
- Unit
- points
- Benchmark version
- Artificial Analysis Intelligence Index v4.3.2
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis independent evaluation on dedicated hardware; the exact page identifies Ling 3.1 Flash as reasoning-enabled with 2,000 reasoning tokens.
- Context
- Source-specific observation; it is not a shared comparison cohort.
AA-Briefcase (Show measurement, test conditions, and source)
- Source value
- 1400.35
- Score
- 1400.35
- Metric
- AA-Briefcase
- Unit
- elo
- Benchmark version
- AA-Briefcase v1.1
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 91-task agentic knowledge-work evaluation; the AA page describes a combined Elo over rubric pass rate, analytical-quality Elo and presentation Elo.
- Context
- Source-specific observation; it is not a shared comparison cohort.
GDPval-AA (Show measurement, test conditions, and source)
- Source value
- 1621.75
- Score
- 1621.75
- Metric
- GDPval-AA
- Unit
- elo
- Benchmark version
- GDPval-AA v2.1
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 220 real-world work tasks, agent loop with shell/web access via Stirrup, blind pairwise comparisons; AA evaluator result for the exact Ling 3.1 Flash slug.
- Context
- Source-specific observation; it is not a shared comparison cohort.
AutomationBench-AA (Show measurement, test conditions, and source)
- Source value
- 0.6174373352506193
- Score
- 0.6174373352506193
- Metric
- AutomationBench-AA
- Unit
- ratio
- Benchmark version
- AutomationBench-AA (dataset version 1.0.6 noted on evaluator page)
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 657 simulated SaaS workflow tasks across six business domains; exact model result from AA's dedicated-hardware evaluator.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Terminal-Bench (Show measurement, test conditions, and source)
- Source value
- 0.333333333333333
- Score
- 0.333333333333333
- Metric
- Terminal-Bench
- Unit
- ratio
- Benchmark version
- Terminal-Bench 4.0
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. AA page states pass@1 averaged over three repeats per task with the mini-swe-agent harness; exact Ling result.
- Context
- Source-specific observation; it is not a shared comparison cohort.
SciCode (Show measurement, test conditions, and source)
- Source value
- 0.540509259259259
- Score
- 0.540509259259259
- Metric
- SciCode
- Unit
- ratio
- Benchmark version
- SciCode current AA implementation
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- disputed
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 288 test-set subproblems from scientist-curated laboratory problems; independent AA evaluation of the exact checkpoint.
- Context
- Source-specific observation; it is not a shared comparison cohort. The evaluator currently labels this result Under review.
Humanity's Last Exam (Show measurement, test conditions, and source)
- Source value
- 0.394346617238183
- Score
- 0.394346617238183
- Metric
- Humanity's Last Exam
- Unit
- ratio
- Benchmark version
- Humanity's Last Exam current AA implementation
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Independent AA evaluation; exact Ling 3.1 Flash result, with reasoning enabled as reported on the exact model page.
- Context
- Source-specific observation; it is not a shared comparison cohort.
GDP.pdf (Show measurement, test conditions, and source)
- Source value
- 0.1
- Score
- 0.1
- Metric
- GDP.pdf
- Unit
- ratio
- Benchmark version
- GDP.pdf current AA implementation
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 100 professional-document tasks across ten domains; all-pass means every atomic criterion passes.
- Context
- Source-specific observation; it is not a shared comparison cohort.
CritPt (Show measurement, test conditions, and source)
- Source value
- 0.18
- Score
- 0.18
- Metric
- CritPt
- Unit
- ratio
- Benchmark version
- CritPt current AA implementation
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- disputed
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 70 test challenges (71 including one example), run without an agent harness against the authors' grading server.
- Context
- Source-specific observation; it is not a shared comparison cohort. The evaluator currently labels this result Under review.
AA-Omniscience Index (Show measurement, test conditions, and source)
- Source value
- 2.216666666666667
- Score
- 2.216666666666667
- Metric
- AA-Omniscience Index
- Unit
- points
- Benchmark version
- AA-Omniscience current AA implementation
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 6,000 factuality/hallucination questions across six domains; independent AA evaluation of exact Ling 3.1 Flash.
- Context
- Source-specific observation; it is not a shared comparison cohort.
AA-LCR (Show measurement, test conditions, and source)
- Source value
- 0.83
- Score
- 0.83
- Metric
- AA-LCR
- Unit
- ratio
- Benchmark version
- AA-LCR v1.1
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. 100 long-context questions over 10k-100k-token documents; score is the average pass rate from the AA equality-checker evaluation.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Sources and data date
Every statement links to its underlying documentation or leaderboard.