Hopp til hovedinnhold

Model tests

International benchmarks

Our own tests show how the models solve three concrete tasks. Here is what independent, international benchmarks say about the same models. The figures are measured by others, with other tasks and settings, and have not been checked by us.

← Back to our tests

Our models

The models we have tested

Each column is one benchmark, with those covering the most models first. The variant the model was measured with is shown below the figure, because the sources often measure a different reasoning level from the one we use (“high”). The list price stays fixed on the far right, so it can always be read against the benchmarks. It is the provider's own price, not a benchmark.

International benchmarks of the models we have tested, with the providers' list prices at the end.

International benchmarks of the models we have tested, with the providers' list prices at the end.
ModelArena textArena · 02/10/2026Arena codeArena · 02/10/2026Arena agentArena · 02/10/2026ECIEpoch AI · 04/10/2026GPQAEpoch AI · 04/10/2026ARC-AGI-2Epoch AI · 04/10/2026HLEEpoch AI · 04/10/2026METREpoch AI · 04/10/2026List priceUSD per 1M tokens, in / out
Claude Opus 5.5Anthropic1,504high · ±9 · 4,552 votes1,815max · ±16 · 2,062 votes0.138high · ±0.022 · 5,359 sessions167.4±4.090.6%max91.7%maxnot measurednot measured4 / 20
GPT-6 SolOpenAI1,457max · ±7 · 8,787 votes1,689max · ±11 · 3,411 votes0.097max · ±0.024 · 6,106 sessionsnot measured94.3%max89.6%maxnot measurednot measured2 / 10
Claude Fable 5.1Anthropic1,501max · ±6 · 11,800 votes1,749max · ±10 · 6,318 votes0.143max · ±0.019 · 15,130 sessions164.8±3.6not measured88.8%high46.5%xhighnot measured10 / 50
Gemini 3.8 FlashGoogle1,495high · ±5 · 26,298 votes1,583high · ±8 · 8,524 votes0.030high · ±0.008 · 34,396 sessions156.9±2.995.4%highnot measured44.5%level unknownnot measured0.75 / 3.75
Gemini 3.1 ProGoogle1,487±3 · 121,806 votes1,446±5 · 24,347 votes-0.077±0.010 · 86,483 sessions154.8±2.494.4%high77.1%0.96 USD46.4%high6 h 24 min2 / 12
Mistral Medium 3.5Mistral AI1,427±6 · 11,757 votes1,263±15 · 2,322 votes-0.124±0.015 · 11,156 sessions141.4±2.6not measurednot measurednot measurednot measured1.65 / 8.25

None of our models has been measured yet in SWE-bench Verified. The benchmark is in the top lists and will be added to the table when the source has figures.

Higher is better in all the benchmarks, and for list prices lower is better. The variant the model was measured with is shown below the figure, because the sources often measure a different reasoning level from the one we use (“high”). Arena scores with few votes are uncertain.

Last updated: Epoch AI 4 October 2026 · Arena 2 October 2026

How to read it

The table shows how the models fare outside our three tasks. Arena is based on user votes, and Epoch AI combines many tests with answer keys into one index.

The benchmarks use other tasks, in English and often at a different reasoning level from ours. Use them alongside our results, not instead of them.

Explanation

What is measured

  • Arena text

    Users ask the same question to two anonymous models and pick the better answer. The result is a ranking in Arena scores, calculated with the Bradley–Terry model and with style control (answer length and formatting do not count), as Arena itself shows it.

    ArenaArena score (Bradley–Terry) · Arena (LMArena)

  • Arena code (WebDev)

    As above, but users ask for websites and apps and judge what is built. There is no style-controlled variant here.

    ArenaArena score (Bradley–Terry) · Arena (LMArena)

  • Arena agent

    Users judge the model as an agent in longer working sessions with tools. Measured in sessions, not votes.

    Arenascore from −1 to 1 · Arena (LMArena)

  • Epoch Capabilities Index (ECI)

    Epoch's composite index. It weights many benchmarks by how hard they are, so models can be compared even if they have not taken exactly the same tests.

    Epoch AIindex points · Epoch AI

  • GPQA Diamond

    198 multiple-choice questions in biology, physics and chemistry at PhD level. The questions are written so that they cannot be googled.

    Epoch AIshare correct · Epoch AI

  • ARC-AGI-2

    Visual pattern tasks that are easy for people and hard for machines. Measures the ability to learn a new rule from a few examples.

    Epoch AIshare correct · Epoch AI after ARC Prize Foundation

  • Humanity's Last Exam

    “Humanity's Last Exam”: around 2,500 very hard questions from experts in many fields.

    Epoch AIshare correct · Epoch AI after Center for AI Safety and Scale AI (lastexam.ai)

  • METR time horizon

    How long programming tasks (measured in human working time) the model solves alone with 50% probability.

    Epoch AIminutes of human working time · Epoch AI after METR

  • SWE-bench Verified

    Real bug fixes in Python projects on GitHub. The model has to write a change that makes the project's own tests pass.

    Epoch AIshare correct · Epoch AI

  • List price

    The provider's own price per million tokens in and out, read from their pricing pages. Not a benchmark.

    USD per 1M tokens · from the provider

Top lists

The top of each benchmark

The ten best in each benchmark, with the source's own names. Arena gives the rank itself; the rank in the Epoch lists is calculated by us, where each reasoning level counts as its own row and equal values share a place. The models we have tested are highlighted. In Arena, only models with at least 1,000 votes (sessions in Arena agent) are included.

Arena text

  1. gemini-4-argon-high 1,525 · 4,932 votes
  2. claude-opus-4-6-high 1,505 · 77,636 votes
  3. claude-fable-5-high 1,504 · 38,387 votes
  4. claude-opus-5.5-high 1,504 tested by us · 4,552 votes
  5. claude-opus-4-7-high 1,501 · 64,946 votes
  6. claude-fable-5.1-max 1,501 tested by us · 11,800 votes
  7. claude-opus-4-6 1,497 · 82,189 votes
  8. gemini-3.8-flash-high 1,495 tested by us · 26,298 votes
  9. muse-spark-1.3-max 1,494 · 12,343 votes
  10. claude-opus-4-7 1,494 · 66,058 votes

Arena code (WebDev)

  1. claude-opus-5.5-max 1,815 tested by us · 2,062 votes
  2. gpt-6-astra-max 1,788 · 6,123 votes
  3. claude-sonnet-5.5-xhigh 1,786 · 1,531 votes
  4. gpt-6.1-sol-max 1,758 · 1,620 votes
  5. claude-fable-5.1-max 1,749 tested by us · 6,318 votes
  6. claude-sonnet-5.5-high 1,715 · 2,527 votes
  7. claude-opus-5-max 1,695 · 16,954 votes
  8. gpt-6-sol-max 1,689 tested by us · 3,411 votes
  9. gemini-4-argon-high 1,680 · 2,422 votes
  10. qwen3.8-max 1,671 · 3,454 votes

Arena agent

  1. Claude Fable 5.1 (Max) 0.143 tested by us · 15,130 sessions
  2. Claude Opus 5.5 (High) 0.138 tested by us · 5,359 sessions
  3. Claude Sonnet 5.5 (Max) 0.125 · 5,219 sessions
  4. GPT 6 Astra (Max) 0.123 · 11,717 sessions
  5. GPT 6.1 Sol (Max) 0.112 · 6,278 sessions
  6. GPT 6 Sol (Max) 0.097 tested by us · 6,106 sessions
  7. Claude Opus 5 (High) 0.087 · 27,560 sessions
  8. Claude Fable 5 (High) 0.082 · 41,832 sessions
  9. Claude Opus 5 (Max) 0.079 · 22,781 sessions
  10. Gemini 4 Argon (High) 0.076 · 10,030 sessions

Epoch Capabilities Index (ECI)

  1. Claude Opus 5.5 167.4 tested by us
  2. GPT-6 Astra 166.5
  3. Claude Sonnet 5.5 165.2
  4. Claude Fable 5.1 164.8 tested by us
  5. Claude Opus 5 162.9
  6. GPT-5.5 Pro 162.4
  7. Claude Fable 5 162.2
  8. GPT-5.6 Sol 161.8
  9. GPT-5.6 Terra 159.8
  10. GPT-5.5 159.2

GPQA Diamond

  1. GPT-6 Astra (max) 95.8%
  2. Claude Sonnet 5.5 (max) 95.6%
  3. GPT-6.1 Sol (max) 95.4%
  4. Gemini 3.8 Flash (high) 95.4% tested by us
  5. Gemini 3.7 Flash (high) 94.8%
  6. GPT-5.4 Pro (xhigh) 94.6%
  7. Gemini 3.1 Pro Preview (high) 94.4% tested by us
  8. GPT-6 Sol (max) 94.3% tested by us
  9. Gemini 3.6 Flash (high) 94.1%
  10. Gemini 3.1 Pro Preview 94.1% tested by us

ARC-AGI-2

  1. GPT-6 Astra (max) 95.0%
  2. GPT-6 Astra (xhigh) 93.3%
  3. GPT-5.6 Sol (max) 92.5%
  4. Claude Opus 5.5 (xhigh) 92.5%
  5. GPT-6 Astra (high) 92.1%
  6. GPT-6 Astra (medium) 92.1%
  7. Claude Opus 5.5 (max) 91.7% tested by us
  8. Claude Opus 5 (max) 90.4%
  9. GPT-5.6 Sol (xhigh) 90.0%
  10. Claude Fable 5.1 (max) 90.0% tested by us

Humanity's Last Exam

  1. GPT-6 Astra (unknown thinking) 54.8%
  2. Claude Fable 5.1 (xhigh) 46.5% tested by us
  3. Gemini 3.1 Pro Preview 46.4% tested by us
  4. Gemini 3.8 Flash (unknown) 44.5% tested by us
  5. GPT-5.4 Pro 44.3%
  6. Muse Spark 40.6%
  7. Gemini 3 Pro Preview 37.5%
  8. GPT-5.4 (xhigh) 36.2%
  9. Claude Opus 4.7 (unknown) 36.2%
  10. Claude Opus 4.6 (max) 34.4%

METR time horizon

  1. Claude Mythos Preview (Early) 17 h 25 min
  2. Claude Opus 4.6 (unknown thinking) 11 h 59 min
  3. Gemini 3.1 Pro Preview 6 h 24 min tested by us
  4. GPT-5.2 (high) 5 h 52 min
  5. GPT-5.3 Codex 5 h 50 min
  6. GPT-5.4 (xhigh) 5 h 42 min
  7. Claude Opus 4.5 (no thinking) 4 h 53 min
  8. Claude Opus 4.5 (16k thinking) 4 h 49 min
  9. Gemini 3 Pro Preview 3 h 44 min
  10. GPT-5.1-Codex-Max 3 h 44 min

SWE-bench Verified

  1. Claude Opus 4.7 (max) 83.5%
  2. GPT-5.5 (xhigh) 80.6%
  3. Gemini 3.5 Flash (high) 79.3%
  4. Claude Opus 4.6 (no thinking) 78.7%
  5. GLM-5.2 (max) 78.7%
  6. DeepSeek v4 Pro (max) 77.6%
  7. Qwen3.7 Max 77.3%
  8. GPT-5.4 (high) 76.9%
  9. Qwen 3.6 Max (Preview) 76.7%
  10. Kimi K2.6 76.7%

Sources

Sources and licence

The figures are fetched on our server and shown without any content from the sources' own websites. Figures from Epoch and Arena are shared under CC BY 4.0 and are reproduced with attribution. We have selected benchmarks and models, rounded the figures and calculated the ranks in Epoch's figures ourselves.

  • Epoch AI

    Licence: CC BY 4.0

    Epoch AI, ‘Capabilities & benchmarking’. Published online at epoch.ai. Retrieved from ‘https://epoch.ai/benchmarks’ [online resource].

    Original source: Center for AI Safety and Scale AI (lastexam.ai) (Humanity's Last Exam); ARC Prize Foundation (ARC-AGI-2); METR (METR time horizon). The figures are collected by Epoch AI.

    Last fetched 4 October 2026 · fetched directly

  • Arena (LMArena)

    Licence: CC BY 4.0

    Arena (LMArena), leaderboard-dataset, CC BY 4.0. https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset

    Last fetched 4 October 2026 · fetched directly

Other sources we have considered

  • Artificial Analysis: Its terms do not allow the figures to be shown in tables for others without a written agreement.
  • Scale SEAL: Has no open data licence or machine-readable interface, so we link instead of reproducing.
  • Vals AI: Has no open data licence or machine-readable interface, so we link instead of reproducing.
  • LiveBench: The results have no licence, and the latest public table lacks the models we test.