Leaderboard
The top models, field by field.
One ranking across reasoning, maths, coding, knowledge, multimodal, instruction following and human preference — built from published figures, and honest about what that can and cannot tell you.
Across the 7 capability fields tracked here, Claude Opus 4.5 holds the highest average percentile among models reporting at least 3 fields. Rankings are computed per field and then averaged, never rolled into a single invented score. 58 of the 159 models in the index clear the reporting threshold.
Where these numbers come from
Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.
A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.
Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.
Overall
Best across every field
Ranked by average percentile across the fields each model reports. Fields shows how many of the 7 it reports at all — a high average over three fields is a narrower claim than the same average over six, and the column is there so you can see which you are looking at.
| # | Model | Fields | Avg. percentile | Top-3 finishes | Strongest field |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.5Anthropic | 4 / 7 | 98 | 4 | Maths #1 of 56 |
| 2 | GPT-5OpenAI | 4 / 7 | 96 | 2 | Multimodal #1 of 39 |
| 3 | Gemini 3 ProGoogle DeepMind | 4 / 7 | 95 | 3 | Reasoning #1 of 83 |
| 4 | Grok 4xAI | 3 / 7 | 90 | 1 | Reasoning #2 of 83 |
| 5 | gpt-oss-120bOpenAI | 3 / 7 | 88 | 0 | Knowledge #5 of 77 |
| 6 | Kimi K2 ThinkingMoonshot AI | 3 / 7 | 87 | 0 | Maths #4 of 56 |
| 7 | DeepSeek-V3.2DeepSeek | 3 / 7 | 84 | 1 | Knowledge #1 of 77 |
| 8 | Claude Sonnet 4.5Anthropic | 4 / 7 | 83 | 1 | Coding #2 of 33 |
| 9 | Qwen3-MaxAlibaba Qwen | 3 / 7 | 82 | 0 | Reasoning #6 of 83 |
| 10 | GPT-5 miniOpenAI | 3 / 7 | 79 | 0 | Coding #10 of 33 |
| 11 | o4-miniOpenAI | 4 / 7 | 79 | 0 | Multimodal #5 of 39 |
| 12 | Kimi K2 InstructMoonshot AI | 3 / 7 | 78 | 1 | Human preference #2 of 24 |
| 13 | o3OpenAI | 4 / 7 | 78 | 1 | Multimodal #2 of 39 |
| 14 | Llama 4 MaverickMeta AI | 4 / 7 | 77 | 1 | Human preference #3 of 24 |
| 15 | GLM-4.6Z.ai (Zhipu) | 3 / 7 | 76 | 0 | Maths #6 of 56 |
| 16 | Gemini 2.5 ProGoogle DeepMind | 5 / 7 | 76 | 1 | Human preference #1 of 24 |
| 17 | DeepSeek-V3.1DeepSeek | 4 / 7 | 74 | 1 | Knowledge #2 of 77 |
| 18 | gpt-oss-20bOpenAI | 3 / 7 | 72 | 0 | Maths #15 of 56 |
| 19 | Grok 3xAI | 3 / 7 | 70 | 0 | Human preference #4 of 24 |
| 20 | Gemini 2.5 FlashGoogle DeepMind | 4 / 7 | 69 | 0 | Human preference #5 of 24 |
| 21 | Claude Opus 4.1Anthropic | 3 / 7 | 69 | 0 | Coding #5 of 33 |
| 22 | Command ACohere | 3 / 7 | 68 | 1 | Instruction following #3 of 15 |
| 23 | Llama-3.1-Nemotron-Ultra-253BNVIDIA | 3 / 7 | 66 | 0 | Instruction following #4 of 15 |
| 24 | GLM-4.5Z.ai (Zhipu) | 3 / 7 | 66 | 0 | Maths #13 of 56 |
| 25 | Gemini 2.0 FlashGoogle DeepMind | 4 / 7 | 64 | 0 | Human preference #7 of 24 |
| 26 | Qwen3-235B-A22BAlibaba Qwen | 4 / 7 | 63 | 0 | Human preference #8 of 24 |
| 27 | Claude Sonnet 4Anthropic | 4 / 7 | 60 | 0 | Coding #7 of 33 |
| 28 | DeepSeek-R1DeepSeek | 5 / 7 | 60 | 1 | Knowledge #3 of 77 |
| 29 | Claude Haiku 4.5Anthropic | 3 / 7 | 59 | 0 | Coding #6 of 33 |
| 30 | Mistral Small 3.2 24BMistral AI | 3 / 7 | 58 | 1 | Instruction following #1 of 15 |
| 31 | DeepSeek-V3DeepSeek | 3 / 7 | 58 | 0 | Human preference #10 of 24 |
| 32 | Amazon Nova PremierAmazon | 3 / 7 | 58 | 0 | Knowledge #8 of 77 |
| 33 | Mistral Medium 3Mistral AI | 3 / 7 | 56 | 0 | Knowledge #11 of 77 |
| 34 | GPT-4.1OpenAI | 5 / 7 | 56 | 0 | Knowledge #6 of 77 |
| 35 | GLM-4.5-AirZ.ai (Zhipu) | 3 / 7 | 56 | 0 | Maths #16 of 56 |
| 36 | MiniMax-M2MiniMax | 3 / 7 | 54 | 0 | Coding #14 of 33 |
| 37 | o3-miniOpenAI | 3 / 7 | 53 | 0 | Reasoning #18 of 83 |
| 38 | Llama 3.3 70BMeta AI | 4 / 7 | 53 | 1 | Instruction following #2 of 15 |
| 39 | Llama 3.1 405BMeta AI | 4 / 7 | 52 | 0 | Instruction following #5 of 15 |
| 40 | Llama 4 ScoutMeta AI | 3 / 7 | 52 | 0 | Knowledge #17 of 77 |
Showing the top 40 of 58 ranked models. A model needs at least 3 of 7 fields to be ranked at all.
Field leaders
Who leads what
The top five in each field, with the score that put them there. These lists are the more reliable half of this page — no averaging across fields, no threshold, just the published numbers in order.
Reasoning
Multi-step logic on problems that cannot be looked up
Ranked on GPQA Diamond · 83 models reporting
- 1Gemini 3 ProGoogle DeepMind91.9%
- 2Grok 4xAI87.5%
- 3Claude Opus 4.5Anthropic87.0%
- 4GPT-5OpenAI85.7%
- 5Grok 4 FastxAI85.7%
Also in this field, on a different test: Claude Opus 5 reports 43.3% on Frontier-Bench v0.1. Its maker published no GPQA Diamond figure, so there is no like-for-like way to place it in the list above — these are not ranked against it or against each other.
- 1Claude Opus 4.5Anthropic96.0%
- 2Gemini 3 ProGoogle DeepMind95.0%
- 3GPT-5OpenAI94.6%
- 4Kimi K2 ThinkingMoonshot AI94.5%
- 5Grok 4xAI94.0%
- 1Claude Opus 4.5Anthropic80.9%
- 2Claude Sonnet 4.5Anthropic77.2%
- 3Gemini 3 ProGoogle DeepMind76.2%
- 4GPT-5OpenAI74.9%
- 5Claude Opus 4.1Anthropic74.5%
Also in this field, on a different test: Claude Fable 5.1 reports 81.2% on SWE-bench Pro; Claude Mythos 5.1 reports 60.9% on Terminal-Bench 4.0; Claude Opus 5 reports 89.1% on Terminal-Bench 2.1. Their makers published no SWE-bench Verified figure, so there is no like-for-like way to place them in the list above — these are not ranked against it or against each other.
- 1DeepSeek-V3.2DeepSeek85.0%
- 2DeepSeek-V3.1DeepSeek84.8%
- 3DeepSeek-R1DeepSeek84.0%
- 4Kimi K2 InstructMoonshot AI81.1%
- 5gpt-oss-120bOpenAI80.9%
- 1GPT-5OpenAI84.2%
- 2o3OpenAI82.9%
- 3Claude Opus 4.5Anthropic82.0%
- 4Gemini 2.5 ProGoogle DeepMind81.7%
- 5o4-miniOpenAI81.6%
- 1Mistral Small 3.2 24BMistral AI92.9%
- 2Llama 3.3 70BMeta AI92.1%
- 3Command ACohere90.9%
- 4Llama-3.1-Nemotron-Ultra-253BNVIDIA89.5%
- 5Llama 3.1 405BMeta AI88.6%
- 1Gemini 2.5 ProGoogle DeepMind1439
- 2Kimi K2 InstructMoonshot AI1420
- 3Llama 4 MaverickMeta AI1417
- 4Grok 3xAI1402
- 5Gemini 2.5 FlashGoogle DeepMind1393
Method
How to read this
How is this leaderboard ranked?
Each field is decided by one benchmark, named in that field's header, and every model listed in it reported that same benchmark. Models are ranked within the field and given a percentile — the share of the field they are at least as good as. A model's overall position is the average of its percentiles across the fields it reports. Percentiles rather than raw ranks, because a field with 83 reporting models and one with 15 would otherwise reward the same achievement very differently.
Why does each field use only one benchmark?
Because averaging a whole group would rank on which test a lab chose rather than on capability. HumanEval is saturated near 92% while SWE-bench Verified sits around 80% for the same class of model, so a model that published only the easy one would float to the top of Coding. One deciding benchmark per field means every model in a list sat the same test. The other benchmarks in each group are still shown on model pages and in side-by-side comparisons.
Why do some well-known models not appear?
A model has to report at least 3 of the seven fields to be ranked. Below that, one strong benchmark would outrank a model measured across six, which would make the table misleading. Models under the threshold still have their own pages and still appear in the per-field lists below.
Is this a measure of which model is best?
No. It is a measure of which models published the best numbers on the tests they chose to publish. That is a real signal and a limited one — a lab can decline to report a benchmark it does badly on, and no one here re-ran anything. Read the per-field lists before the overall table; they are closer to the truth.
What are the seven fields?
Reasoning, Maths, Coding, Knowledge, Multimodal, Instruction following and Human preference. Each groups the benchmarks that measure the same thing, and a model's field score averages the ones it reports in that group.
Compare
A ranking is not a decision.
Pick the two or three models this page put in front of you and read them column by column — price, context, licence and every benchmark side by side.