International benchmarks
Our own tests show how the models solve three concrete tasks. Here is what independent, international benchmarks say about the same models. The figures are measured by others, with other tasks and settings, and have not been checked by us.
← Back to our testsOur models
The models we have tested
Each column is one benchmark, with those covering the most models first. The variant the model was measured with is shown below the figure, because the sources often measure a different reasoning level from the one we use (“high”). The list price stays fixed on the far right, so it can always be read against the benchmarks. It is the provider's own price, not a benchmark.
International benchmarks of the models we have tested, with the providers' list prices at the end.
| Model | Arena textArena · 02/10/2026 | Arena codeArena · 02/10/2026 | Arena agentArena · 02/10/2026 | ECIEpoch AI · 04/10/2026 | GPQAEpoch AI · 04/10/2026 | ARC-AGI-2Epoch AI · 04/10/2026 | HLEEpoch AI · 04/10/2026 | METREpoch AI · 04/10/2026 | List priceUSD per 1M tokens, in / out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5Anthropic | 1,504high · ±9 · 4,552 votes | 1,815max · ±16 · 2,062 votes | 0.138high · ±0.022 · 5,359 sessions | 167.4±4.0 | 90.6%max | 91.7%max | not measured | not measured | 4 / 20 |
| GPT-6 SolOpenAI | 1,457max · ±7 · 8,787 votes | 1,689max · ±11 · 3,411 votes | 0.097max · ±0.024 · 6,106 sessions | not measured | 94.3%max | 89.6%max | not measured | not measured | 2 / 10 |
| Claude Fable 5.1Anthropic | 1,501max · ±6 · 11,800 votes | 1,749max · ±10 · 6,318 votes | 0.143max · ±0.019 · 15,130 sessions | 164.8±3.6 | not measured | 88.8%high | 46.5%xhigh | not measured | 10 / 50 |
| Gemini 3.8 FlashGoogle | 1,495high · ±5 · 26,298 votes | 1,583high · ±8 · 8,524 votes | 0.030high · ±0.008 · 34,396 sessions | 156.9±2.9 | 95.4%high | not measured | 44.5%level unknown | not measured | 0.75 / 3.75 |
| Gemini 3.1 ProGoogle | 1,487±3 · 121,806 votes | 1,446±5 · 24,347 votes | -0.077±0.010 · 86,483 sessions | 154.8±2.4 | 94.4%high | 77.1%0.96 USD | 46.4%high | 6 h 24 min | 2 / 12 |
| Mistral Medium 3.5Mistral AI | 1,427±6 · 11,757 votes | 1,263±15 · 2,322 votes | -0.124±0.015 · 11,156 sessions | 141.4±2.6 | not measured | not measured | not measured | not measured | 1.65 / 8.25 |
None of our models has been measured yet in SWE-bench Verified. The benchmark is in the top lists and will be added to the table when the source has figures.
Higher is better in all the benchmarks, and for list prices lower is better. The variant the model was measured with is shown below the figure, because the sources often measure a different reasoning level from the one we use (“high”). Arena scores with few votes are uncertain.
Last updated: Epoch AI 4 October 2026 · Arena 2 October 2026
How to read it
The table shows how the models fare outside our three tasks. Arena is based on user votes, and Epoch AI combines many tests with answer keys into one index.
The benchmarks use other tasks, in English and often at a different reasoning level from ours. Use them alongside our results, not instead of them.
Explanation
What is measured
Arena text
Users ask the same question to two anonymous models and pick the better answer. The result is a ranking in Arena scores, calculated with the Bradley–Terry model and with style control (answer length and formatting do not count), as Arena itself shows it.
ArenaArena score (Bradley–Terry) · Arena (LMArena)
Arena code (WebDev)
As above, but users ask for websites and apps and judge what is built. There is no style-controlled variant here.
ArenaArena score (Bradley–Terry) · Arena (LMArena)
Arena agent
Users judge the model as an agent in longer working sessions with tools. Measured in sessions, not votes.
Arenascore from −1 to 1 · Arena (LMArena)
Epoch Capabilities Index (ECI)
Epoch's composite index. It weights many benchmarks by how hard they are, so models can be compared even if they have not taken exactly the same tests.
Epoch AIindex points · Epoch AI
GPQA Diamond
198 multiple-choice questions in biology, physics and chemistry at PhD level. The questions are written so that they cannot be googled.
Epoch AIshare correct · Epoch AI
ARC-AGI-2
Visual pattern tasks that are easy for people and hard for machines. Measures the ability to learn a new rule from a few examples.
Epoch AIshare correct · Epoch AI after ARC Prize Foundation
Humanity's Last Exam
“Humanity's Last Exam”: around 2,500 very hard questions from experts in many fields.
Epoch AIshare correct · Epoch AI after Center for AI Safety and Scale AI (lastexam.ai)
METR time horizon
How long programming tasks (measured in human working time) the model solves alone with 50% probability.
SWE-bench Verified
Real bug fixes in Python projects on GitHub. The model has to write a change that makes the project's own tests pass.
Epoch AIshare correct · Epoch AI
List price
The provider's own price per million tokens in and out, read from their pricing pages. Not a benchmark.
USD per 1M tokens · from the provider
Top lists
The top of each benchmark
The ten best in each benchmark, with the source's own names. Arena gives the rank itself; the rank in the Epoch lists is calculated by us, where each reasoning level counts as its own row and equal values share a place. The models we have tested are highlighted. In Arena, only models with at least 1,000 votes (sessions in Arena agent) are included.
Arena text
- gemini-4-argon-high 1,525 · 4,932 votes
- claude-opus-4-6-high 1,505 · 77,636 votes
- claude-fable-5-high 1,504 · 38,387 votes
- claude-opus-5.5-high 1,504 tested by us · 4,552 votes
- claude-opus-4-7-high 1,501 · 64,946 votes
- claude-fable-5.1-max 1,501 tested by us · 11,800 votes
- claude-opus-4-6 1,497 · 82,189 votes
- gemini-3.8-flash-high 1,495 tested by us · 26,298 votes
- muse-spark-1.3-max 1,494 · 12,343 votes
- claude-opus-4-7 1,494 · 66,058 votes
Arena code (WebDev)
- claude-opus-5.5-max 1,815 tested by us · 2,062 votes
- gpt-6-astra-max 1,788 · 6,123 votes
- claude-sonnet-5.5-xhigh 1,786 · 1,531 votes
- gpt-6.1-sol-max 1,758 · 1,620 votes
- claude-fable-5.1-max 1,749 tested by us · 6,318 votes
- claude-sonnet-5.5-high 1,715 · 2,527 votes
- claude-opus-5-max 1,695 · 16,954 votes
- gpt-6-sol-max 1,689 tested by us · 3,411 votes
- gemini-4-argon-high 1,680 · 2,422 votes
- qwen3.8-max 1,671 · 3,454 votes
Arena agent
- Claude Fable 5.1 (Max) 0.143 tested by us · 15,130 sessions
- Claude Opus 5.5 (High) 0.138 tested by us · 5,359 sessions
- Claude Sonnet 5.5 (Max) 0.125 · 5,219 sessions
- GPT 6 Astra (Max) 0.123 · 11,717 sessions
- GPT 6.1 Sol (Max) 0.112 · 6,278 sessions
- GPT 6 Sol (Max) 0.097 tested by us · 6,106 sessions
- Claude Opus 5 (High) 0.087 · 27,560 sessions
- Claude Fable 5 (High) 0.082 · 41,832 sessions
- Claude Opus 5 (Max) 0.079 · 22,781 sessions
- Gemini 4 Argon (High) 0.076 · 10,030 sessions
Epoch Capabilities Index (ECI)
- Claude Opus 5.5 167.4 tested by us
- GPT-6 Astra 166.5
- Claude Sonnet 5.5 165.2
- Claude Fable 5.1 164.8 tested by us
- Claude Opus 5 162.9
- GPT-5.5 Pro 162.4
- Claude Fable 5 162.2
- GPT-5.6 Sol 161.8
- GPT-5.6 Terra 159.8
- GPT-5.5 159.2
GPQA Diamond
- GPT-6 Astra (max) 95.8%
- Claude Sonnet 5.5 (max) 95.6%
- GPT-6.1 Sol (max) 95.4%
- Gemini 3.8 Flash (high) 95.4% tested by us
- Gemini 3.7 Flash (high) 94.8%
- GPT-5.4 Pro (xhigh) 94.6%
- Gemini 3.1 Pro Preview (high) 94.4% tested by us
- GPT-6 Sol (max) 94.3% tested by us
- Gemini 3.6 Flash (high) 94.1%
- Gemini 3.1 Pro Preview 94.1% tested by us
ARC-AGI-2
- GPT-6 Astra (max) 95.0%
- GPT-6 Astra (xhigh) 93.3%
- GPT-5.6 Sol (max) 92.5%
- Claude Opus 5.5 (xhigh) 92.5%
- GPT-6 Astra (high) 92.1%
- GPT-6 Astra (medium) 92.1%
- Claude Opus 5.5 (max) 91.7% tested by us
- Claude Opus 5 (max) 90.4%
- GPT-5.6 Sol (xhigh) 90.0%
- Claude Fable 5.1 (max) 90.0% tested by us
Humanity's Last Exam
- GPT-6 Astra (unknown thinking) 54.8%
- Claude Fable 5.1 (xhigh) 46.5% tested by us
- Gemini 3.1 Pro Preview 46.4% tested by us
- Gemini 3.8 Flash (unknown) 44.5% tested by us
- GPT-5.4 Pro 44.3%
- Muse Spark 40.6%
- Gemini 3 Pro Preview 37.5%
- GPT-5.4 (xhigh) 36.2%
- Claude Opus 4.7 (unknown) 36.2%
- Claude Opus 4.6 (max) 34.4%
METR time horizon
- Claude Mythos Preview (Early) 17 h 25 min
- Claude Opus 4.6 (unknown thinking) 11 h 59 min
- Gemini 3.1 Pro Preview 6 h 24 min tested by us
- GPT-5.2 (high) 5 h 52 min
- GPT-5.3 Codex 5 h 50 min
- GPT-5.4 (xhigh) 5 h 42 min
- Claude Opus 4.5 (no thinking) 4 h 53 min
- Claude Opus 4.5 (16k thinking) 4 h 49 min
- Gemini 3 Pro Preview 3 h 44 min
- GPT-5.1-Codex-Max 3 h 44 min
SWE-bench Verified
- Claude Opus 4.7 (max) 83.5%
- GPT-5.5 (xhigh) 80.6%
- Gemini 3.5 Flash (high) 79.3%
- Claude Opus 4.6 (no thinking) 78.7%
- GLM-5.2 (max) 78.7%
- DeepSeek v4 Pro (max) 77.6%
- Qwen3.7 Max 77.3%
- GPT-5.4 (high) 76.9%
- Qwen 3.6 Max (Preview) 76.7%
- Kimi K2.6 76.7%
Sources
Sources and licence
The figures are fetched on our server and shown without any content from the sources' own websites. Figures from Epoch and Arena are shared under CC BY 4.0 and are reproduced with attribution. We have selected benchmarks and models, rounded the figures and calculated the ranks in Epoch's figures ourselves.
Epoch AI
Licence: CC BY 4.0
Epoch AI, ‘Capabilities & benchmarking’. Published online at epoch.ai. Retrieved from ‘https://epoch.ai/benchmarks’ [online resource].
Original source: Center for AI Safety and Scale AI (lastexam.ai) (Humanity's Last Exam); ARC Prize Foundation (ARC-AGI-2); METR (METR time horizon). The figures are collected by Epoch AI.
Last fetched 4 October 2026 · fetched directly
Arena (LMArena)
Licence: CC BY 4.0
Arena (LMArena), leaderboard-dataset, CC BY 4.0. https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset
Last fetched 4 October 2026 · fetched directly
Other sources we have considered
- Artificial Analysis: Its terms do not allow the figures to be shown in tables for others without a written agreement.
- Scale SEAL: Has no open data licence or machine-readable interface, so we link instead of reproducing.
- Vals AI: Has no open data licence or machine-readable interface, so we link instead of reproducing.
- LiveBench: The results have no licence, and the latest public table lacks the models we test.