CorX Labs

Leaderboard

The top models, field by field.

One ranking across reasoning, maths, coding, knowledge, multimodal, instruction following and human preference — built from published figures, and honest about what that can and cannot tell you.

Across the 7 capability fields tracked here, Claude Opus 4.5 holds the highest average percentile among models reporting at least 3 fields. Rankings are computed per field and then averaged, never rolled into a single invented score. 58 of the 159 models in the index clear the reporting threshold.

Where these numbers come from

Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.

A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.

Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.

Overall

Best across every field

Ranked by average percentile across the fields each model reports. Fields shows how many of the 7 it reports at all — a high average over three fields is a narrower claim than the same average over six, and the column is there so you can see which you are looking at.

AI models ranked by average percentile across the capability fields they report.
# Model Fields Avg. percentile Top-3 finishes Strongest field
1Claude Opus 4.5Anthropic4 / 7984Maths #1 of 56
2GPT-5OpenAI4 / 7962Multimodal #1 of 39
3Gemini 3 ProGoogle DeepMind4 / 7953Reasoning #1 of 83
4Grok 4xAI3 / 7901Reasoning #2 of 83
5gpt-oss-120bOpenAI3 / 7880Knowledge #5 of 77
6Kimi K2 ThinkingMoonshot AI3 / 7870Maths #4 of 56
7DeepSeek-V3.2DeepSeek3 / 7841Knowledge #1 of 77
8Claude Sonnet 4.5Anthropic4 / 7831Coding #2 of 33
9Qwen3-MaxAlibaba Qwen3 / 7820Reasoning #6 of 83
10GPT-5 miniOpenAI3 / 7790Coding #10 of 33
11o4-miniOpenAI4 / 7790Multimodal #5 of 39
12Kimi K2 InstructMoonshot AI3 / 7781Human preference #2 of 24
13o3OpenAI4 / 7781Multimodal #2 of 39
14Llama 4 MaverickMeta AI4 / 7771Human preference #3 of 24
15GLM-4.6Z.ai (Zhipu)3 / 7760Maths #6 of 56
16Gemini 2.5 ProGoogle DeepMind5 / 7761Human preference #1 of 24
17DeepSeek-V3.1DeepSeek4 / 7741Knowledge #2 of 77
18gpt-oss-20bOpenAI3 / 7720Maths #15 of 56
19Grok 3xAI3 / 7700Human preference #4 of 24
20Gemini 2.5 FlashGoogle DeepMind4 / 7690Human preference #5 of 24
21Claude Opus 4.1Anthropic3 / 7690Coding #5 of 33
22Command ACohere3 / 7681Instruction following #3 of 15
23Llama-3.1-Nemotron-Ultra-253BNVIDIA3 / 7660Instruction following #4 of 15
24GLM-4.5Z.ai (Zhipu)3 / 7660Maths #13 of 56
25Gemini 2.0 FlashGoogle DeepMind4 / 7640Human preference #7 of 24
26Qwen3-235B-A22BAlibaba Qwen4 / 7630Human preference #8 of 24
27Claude Sonnet 4Anthropic4 / 7600Coding #7 of 33
28DeepSeek-R1DeepSeek5 / 7601Knowledge #3 of 77
29Claude Haiku 4.5Anthropic3 / 7590Coding #6 of 33
30Mistral Small 3.2 24BMistral AI3 / 7581Instruction following #1 of 15
31DeepSeek-V3DeepSeek3 / 7580Human preference #10 of 24
32Amazon Nova PremierAmazon3 / 7580Knowledge #8 of 77
33Mistral Medium 3Mistral AI3 / 7560Knowledge #11 of 77
34GPT-4.1OpenAI5 / 7560Knowledge #6 of 77
35GLM-4.5-AirZ.ai (Zhipu)3 / 7560Maths #16 of 56
36MiniMax-M2MiniMax3 / 7540Coding #14 of 33
37o3-miniOpenAI3 / 7530Reasoning #18 of 83
38Llama 3.3 70BMeta AI4 / 7531Instruction following #2 of 15
39Llama 3.1 405BMeta AI4 / 7520Instruction following #5 of 15
40Llama 4 ScoutMeta AI3 / 7520Knowledge #17 of 77

Showing the top 40 of 58 ranked models. A model needs at least 3 of 7 fields to be ranked at all.

Field leaders

Who leads what

The top five in each field, with the score that put them there. These lists are the more reliable half of this page — no averaging across fields, no threshold, just the published numbers in order.

Reasoning

Multi-step logic on problems that cannot be looked up

Ranked on GPQA Diamond · 83 models reporting

  1. 1Gemini 3 ProGoogle DeepMind91.9%
  2. 2Grok 4xAI87.5%
  3. 3Claude Opus 4.5Anthropic87.0%
  4. 4GPT-5OpenAI85.7%
  5. 5Grok 4 FastxAI85.7%

Also in this field, on a different test: Claude Opus 5 reports 43.3% on Frontier-Bench v0.1. Its maker published no GPQA Diamond figure, so there is no like-for-like way to place it in the list above — these are not ranked against it or against each other.

Maths

Competition mathematics, graded on the final answer

Ranked on AIME 2025 · 56 models reporting

  1. 1Claude Opus 4.5Anthropic96.0%
  2. 2Gemini 3 ProGoogle DeepMind95.0%
  3. 3GPT-5OpenAI94.6%
  4. 4Kimi K2 ThinkingMoonshot AI94.5%
  5. 5Grok 4xAI94.0%

Coding

Writing and repairing real code

Ranked on SWE-bench Verified · 33 models reporting

  1. 1Claude Opus 4.5Anthropic80.9%
  2. 2Claude Sonnet 4.5Anthropic77.2%
  3. 3Gemini 3 ProGoogle DeepMind76.2%
  4. 4GPT-5OpenAI74.9%
  5. 5Claude Opus 4.1Anthropic74.5%

Also in this field, on a different test: Claude Fable 5.1 reports 81.2% on SWE-bench Pro; Claude Mythos 5.1 reports 60.9% on Terminal-Bench 4.0; Claude Opus 5 reports 89.1% on Terminal-Bench 2.1. Their makers published no SWE-bench Verified figure, so there is no like-for-like way to place them in the list above — these are not ranked against it or against each other.

Knowledge

Breadth of factual recall under exam conditions

Ranked on MMLU-Pro · 77 models reporting

  1. 1DeepSeek-V3.2DeepSeek85.0%
  2. 2DeepSeek-V3.1DeepSeek84.8%
  3. 3DeepSeek-R1DeepSeek84.0%
  4. 4Kimi K2 InstructMoonshot AI81.1%
  5. 5gpt-oss-120bOpenAI80.9%

Multimodal

Reading charts, diagrams and photographs

Ranked on MMMU · 39 models reporting

  1. 1GPT-5OpenAI84.2%
  2. 2o3OpenAI82.9%
  3. 3Claude Opus 4.5Anthropic82.0%
  4. 4Gemini 2.5 ProGoogle DeepMind81.7%
  5. 5o4-miniOpenAI81.6%

Instruction following

Obeying an exact, checkable format

Ranked on IFEval · 15 models reporting

  1. 1Mistral Small 3.2 24BMistral AI92.9%
  2. 2Llama 3.3 70BMeta AI92.1%
  3. 3Command ACohere90.9%
  4. 4Llama-3.1-Nemotron-Ultra-253BNVIDIA89.5%
  5. 5Llama 3.1 405BMeta AI88.6%

Human preference

Which answer people pick, blind

Ranked on LMArena Elo · 24 models reporting

  1. 1Gemini 2.5 ProGoogle DeepMind1439
  2. 2Kimi K2 InstructMoonshot AI1420
  3. 3Llama 4 MaverickMeta AI1417
  4. 4Grok 3xAI1402
  5. 5Gemini 2.5 FlashGoogle DeepMind1393

Method

How to read this

How is this leaderboard ranked?

Each field is decided by one benchmark, named in that field's header, and every model listed in it reported that same benchmark. Models are ranked within the field and given a percentile — the share of the field they are at least as good as. A model's overall position is the average of its percentiles across the fields it reports. Percentiles rather than raw ranks, because a field with 83 reporting models and one with 15 would otherwise reward the same achievement very differently.

Why does each field use only one benchmark?

Because averaging a whole group would rank on which test a lab chose rather than on capability. HumanEval is saturated near 92% while SWE-bench Verified sits around 80% for the same class of model, so a model that published only the easy one would float to the top of Coding. One deciding benchmark per field means every model in a list sat the same test. The other benchmarks in each group are still shown on model pages and in side-by-side comparisons.

Why do some well-known models not appear?

A model has to report at least 3 of the seven fields to be ranked. Below that, one strong benchmark would outrank a model measured across six, which would make the table misleading. Models under the threshold still have their own pages and still appear in the per-field lists below.

Is this a measure of which model is best?

No. It is a measure of which models published the best numbers on the tests they chose to publish. That is a real signal and a limited one — a lab can decline to report a benchmark it does badly on, and no one here re-ran anything. Read the per-field lists before the overall table; they are closer to the truth.

What are the seven fields?

Reasoning, Maths, Coding, Knowledge, Multimodal, Instruction following and Human preference. Each groups the benchmarks that measure the same thing, and a model's field score averages the ones it reports in that group.

Compare

A ranking is not a decision.

Pick the two or three models this page put in front of you and read them column by column — price, context, licence and every benchmark side by side.