o3-pro
Not tested yet
OpenAI
No published category-wide ranking is available for this model.
Entries 1-12 of 18. Page 1 of 2.
Each chart compares models in the same documented test and highlights the current model.
By task type
Overall
What this benchmark measures: 225 code-editing tasks from Exercism across C++, Go, Java, JavaScript, Python, and Rust.
Context: Aider, the edit format, reasoning budget, and two attempts are part of the result.
Scale: 0% - 100%
Overall
What this benchmark measures: Multi-turn conversations test memory, instruction retention, versioned editing, and self-coherence.
Context: Tasks were selected from shared failures of six older frontier models and may disadvantage those systems.
Scale: 0% - 100%
Overall
What this benchmark measures: Native reasoning tasks in French, Spanish, and Chinese that were not translated from English templates.
Context: Three languages counter English-only evaluation, but they do not represent global language coverage.
Scale: 0% - 100%
Overall
What this benchmark measures: 600 realistic finance questions with weighted criteria, including a 300-question hard subset.
Context: An o4-mini judge scores the responses, so domain correctness and operational reliability remain separate questions.
Scale: 0% - 100%
Overall
What this benchmark measures: 500 realistic legal questions with weighted criteria across jurisdictions and practice areas.
Context: The benchmark measures professional response quality but cannot replace legal review or human accountability.
Scale: 0% - 100%
Overall
What this benchmark measures: 1,490 tasks test adaptive explanations, feedback, and hints across six STEM subjects.
Context: A Claude 4 Sonnet judge scores weighted rubrics, so strong results do not prove safe tutoring of real learners.
Scale: 0% - 100%
Overall
What this benchmark measures: 758 image-text tasks combine OCR, spatial understanding, object recognition, and open-ended visual reasoning.
Context: The dataset is private and uses task-specific rubrics, making independent reproduction difficult.
Scale: 0% - 100%
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
Every observation retains its source value and published test conditions.
Our model tests
Compare how the models respond to the same prompt. Each test shows the first attempt, with no subsequent fixes to the generated code. These results do not contribute to an overall score.
A community classic for free-form SVG drawing.
o3-pro
Not tested yet
Generate an SVG of a pelican riding a bicycle
Token limit including reasoning: 8,192 tokens
Task origin (Simon Willison)Profile
Published information about this model. Existing estimates are explicitly labeled.
Dated documentation and model-card observations. Provider limits, native context and extended context can differ. Configuration notes retain the source wording.
Pricing
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
Cost example
$6
100 requests with 1,000 input and 500 output tokens each calculate to $6. Input accounts for $2; output accounts for $4.
The calculation uses documented token prices and no cache discount. It excludes extra tools, tax, reasoning tokens, and further hidden output tokens.
Open sourceHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
There are no published direct comparisons for this model yet.
Evidence
Every statement links to its underlying documentation or leaderboard.
Base API prices without caching or batch discounts. Higher context tiers and other rates are listed under Costs.
Documented modalities. Retrieved September 8, 2026.
Documented modalities. Retrieved September 8, 2026.
Research date September 8, 2026. 7 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication.
Not found in the inspected sources
independent-benchmarks. The recorded o3-pro benchmark rows are official vendor evaluations; no independent evaluator row was merged.
Not found in the inspected sources
parameters-billions. OpenAI does not publish parameter counts.
Not found in the inspected sources
AA benchmark rows. No exact-model row in the selected public JSON-LD datasets (some pages only expose N/A or unknown metric summaries).