AI model comparison / The decision desk

Better at what?

Choose a task. Inspect the evidence. Compare the tradeoffs.
A practical model shortlist starts with the work you need done.

Source snapshot
Sep 25, 2026

01 / Start with the task

Different tests.
Different strengths.

Highest observed means in six Epoch AI benchmark tables. These are source results, not Rubrex tests or universal winners. Coverage and evaluation settings vary.

01 / Long-horizon coding

Claude Fable 5.1 (high)

73.3%MirrorCode
Mean score · SE not reported

Reimplement complete programs from their behavior and documentation.

Run started: .

Uses a large agent budget: up to seven days and 10 billion tokens per attempt. This is not a quick coding-assistant test.

Source & test conditions ↗
02 / Scientific questions

GPT-6 Astra (max)

95.8%GPQA Diamond
Mean score · ± 1.4 pp SE

Answer difficult multiple-choice questions in biology, chemistry, and physics.

Run started: .

Measures a specific question set. It does not establish scientific research ability or accuracy on your documents.

Source & test conditions ↗
03 / Advanced mathematics

GPT-6 Astra (max)

93.7%FrontierMath Tiers 1–3 v2
Mean score · ± 1.4 pp SE

Solve expert-written mathematics problems with access to Python.

Run started: .

These are v2 results with tool access. Do not combine them with the original question set or tool-free results.

Source & test conditions ↗
04 / Factual questions

GPT-6 Astra (max)

75.6%SimpleQA Verified
Mean score · ± 1.4 pp SE

Answer short factual questions, scored using Epoch AI’s proportion-correct measure.

Run started: .

Epoch encourages a best guess in this setup. This is neither a hallucination rate nor a test of useful abstention or private-document RAG.

Source & test conditions ↗
05 / Long-horizon game agents

GPT-6 Astra (max)

76.2%EBR-bench
Mean score · ± 8.4 pp SE

Track progress in the Earthborne Rangers game environment.

Run started: .

Game-agent performance is a narrow signal. It does not directly measure business workflow completion or tool authorization.

Source & test conditions ↗
06 / Repository issue fixes

Claude Opus 4.7 (max)

83.5%SWE-bench Verified
Mean score · ± 1.7 pp SE

Resolve software issues under Epoch AI’s evaluation setup.

Run started: .

This source has limited recent-model coverage. Agent scaffolds, budgets, and evaluation subsets can differ from other SWE-bench leaderboards.

Source & test conditions ↗
Which AI model is best?

In this snapshot, Claude Fable 5.1 (high) has the highest observed mean on MirrorCode; GPT-6 Astra (max) leads the included GPQA Diamond runs. Use the task-specific results below to select candidates, then test them on your own examples at a fixed budget.

02 / Explore the evidence

Compare within the same test.

MirrorCode

Reimplement complete programs from their behavior and documentation. Uses a large agent budget: up to seven days and 10 billion tokens per attempt. This is not a quick coding-assistant test.

Read the original methodology ↗

Showing 8 of 8 matching evaluation runs. Ordered by mean score, highest first. ± is one standard error in percentage points, where supplied.

Scroll the table horizontally on smaller screens.

MirrorCode: Epoch AI internal evaluations
Model and settingMean scoreStandard errorRun started
Claude Fable 5.1 (high)Anthropic73.3%Not reportedSep 10, 2026
Claude Fable 5 (high)Anthropic63.9%± 10.4 ppAug 10, 2026
GPT-6 Astra (high)OpenAI46.7%Not reportedAug 30, 2026
Claude Opus 4.7 (high)Anthropic31.1%± 9.0 ppAug 12, 2026
GPT-5.6 Sol (high)OpenAI20.0%± 9.2 ppAug 10, 2026
GPT-5.4 (high)OpenAI15.6%± 7.7 ppAug 10, 2026
GPT-5.5 (high)OpenAI10.0%± 6.0 ppAug 10, 2026
Gemini 3.1 Pro PreviewGoogle DeepMind8.9%± 4.6 ppAug 10, 2026

03 / Build your shortlist

Two models. Six different questions.

Compare model families using their highest observed mean score in each test. The exact setting appears below every result. This view does not equalize reasoning effort or compute budgets.

Long-horizon coding

MirrorCode ↗
Claude Fable 5.173.3%Claude Fable 5.1 (high)Standard error not reported
GPT-6 Astra46.7%GPT-6 Astra (high)Standard error not reported

Scientific questions

GPQA Diamond ↗
Claude Fable 5.1No resultMissing from this source snapshot
GPT-6 Astra95.8%GPT-6 Astra (max)± 1.4 pp standard error
Claude Fable 5.190.2%Claude Fable 5.1 (max)± 1.8 pp standard error
GPT-6 Astra93.7%GPT-6 Astra (max)± 1.4 pp standard error

Factual questions

SimpleQA Verified ↗
Claude Fable 5.170.8%Claude Fable 5.1 (max)± 1.4 pp standard error
GPT-6 Astra75.6%GPT-6 Astra (max)± 1.4 pp standard error

Long-horizon game agents

EBR-bench ↗
Claude Fable 5.157.1%Claude Fable 5.1 (max)± 6.0 pp standard error
GPT-6 Astra76.2%GPT-6 Astra (max)± 8.4 pp standard error

Repository issue fixes

SWE-bench Verified ↗
Claude Fable 5.1No resultMissing from this source snapshot
GPT-6 AstraNo resultMissing from this source snapshot

A missing score is unknown, not zero. Higher means can have overlapping uncertainty. These results are a shortlist signal, not proof of superiority on your workload.

04 / Make the budget explicit

What would the tokens cost?

A dated sample from OpenRouter’s public catalog. Prices are USD per million tokens at the listed base tier. The example uses 10,000 input tokens and 2,000 total billed output tokens for one request.

Scroll the table horizontally on smaller screens.

OpenRouter catalog prices, independent of the benchmark runs above
ModelInput / 1MOutput / 1MExample request
Z.ai: GLM 5.3 Flash ↗$0.045$0.14$0.0007
OpenAI: GPT-6 Luna ↗Higher-context price tiers apply$0.1$0.5$0.0020
DeepSeek: DeepSeek V4.1 Flash ↗Higher-context price tiers apply$0.3$1.2$0.0054
Google: Gemini 3.8 Flash ↗$0.75$3.75$0.0150
Meta: Muse Spark 1.3 ↗$1.25$4.25$0.0210
Z.ai: GLM 5.3 ↗$1.4$4.4$0.0228
Qwen: Qwen3.8 Max (0902) ↗$2$6$0.0320
OpenAI: GPT-6 Sol ↗Higher-context price tiers apply$2$10$0.0400
Anthropic: Claude Sonnet 5 ↗$2$10$0.0400
Google: Gemini 3.1 Pro Preview ↗Higher-context price tiers apply$2$12$0.0440
MoonshotAI: Kimi K3 ↗$3$15$0.0600
Anthropic: Claude Opus 5.5 ↗$4$20$0.0800
OpenAI: GPT-6 Astra ↗Higher-context price tiers apply$10$50$0.2000
Anthropic: Claude Fable 5.1 ↗$10$50$0.2000

This is a token estimate, not cost per successful task. Extra reasoning tokens, retries, tools, search, media, routing/provider differences, and higher-context tiers can change the bill. Cache and batch discounts are excluded. Verify the linked model listing before purchase.

05 / Evidence you can trace

Know what the numbers mean.

One source for benchmark scores

215 model families appear across these six tables. We import Epoch AI’s internal-run mean scores, retain exact settings, and show supplied standard errors. We convert fractions to percentages and sort results. We do not average unrelated benchmarks into a new rating.

Epoch AI data & attribution ↗

Freshness is visible

Sources retrieved Sep 25, 2026. This is a saved snapshot, not a live feed. The run date in each row can be older than the download date. New models enter different tests at different times; a missing result does not mean poor performance.

Download the current source export ↗

Context changes the decision

Reasoning effort, tools, scaffolds, and compute budgets affect results. A small mean-score gap is not a statistical significance test. Before switching models, evaluate task success, tail latency, failure severity, and total cost under your own operating conditions.

Read the model migration guide →

Speed and preferences need another lens

These benchmark tables do not measure API speed or user preference. Artificial Analysis provides performance comparisons; Arena provides preference-based comparisons. Their tests answer different questions, so their numbers are not merged into this table.

Explore speed at Artificial Analysis ↗
Explore preferences at Arena ↗

Benchmark data: Epoch AI, Capabilities & Benchmarking, used under CC BY 4.0. Selected, reformatted, and converted to percentages by Rubrex. Original benchmark creators and methods are linked in each category. No endorsement is implied. API price data: OpenRouter public model catalog, retrieved on the same date.

The source export can change after this snapshot. We publish no provider-sponsored ranking and make no claim that these scores establish production reliability.

Before you choose

A few useful answers.

Why is a newer model missing from some tests?

Independent evaluation takes time, and each benchmark covers a different set of models. We leave missing results empty. We do not infer a score from another model in the same family or from a launch claim.

Does the cheapest model offer the best value?

Only if it meets your quality bar. Compare cost per successful task, including retries and human review. A low token price can still produce an expensive workflow.

Can these results choose a model for RAG or customer support?

They can help build a shortlist. You still need representative documents, answerability cases, citation checks, tool boundaries, and tests of the complete application.

Build a RAG evaluation →

From a shortlist to a decision

Which model works
for your product?

Bring one use case. We can define the quality bar, establish a baseline, and investigate the failures.

Explore the AI Reliability Sprint →