CorX Labs

Knowledge

MMLU-Pro

12,000 reasoning-heavy multiple-choice questions across 14 academic subjects, with ten options instead of four. The harder successor to MMLU.

MMLU-Pro is scored as a percentage of questions answered correctly. 77 of the 159 models in this index report a score for it. The highest published figure here is 85%, from DeepSeek-V3.2.


Reported scores

Top 20 on MMLU-Pro

Ordered by the figure each maker published. Models that have not reported this benchmark are not listed — an absent score is not a low score.

Models ranked by published MMLU-Pro score
#Model Published score
1DeepSeek-V3.2DeepSeek85%Best
2DeepSeek-V3.1DeepSeek84.8%
3DeepSeek-R1DeepSeek84%
4Kimi K2 InstructMoonshot AI81.1%
5gpt-oss-120bOpenAI80.9%
6GPT-4.1OpenAI80.5%
7Llama 4 MaverickMeta AI80.5%
8Amazon Nova PremierAmazon80%
9Claude 3.5 SonnetAnthropic78%
10Gemini 2.0 FlashGoogle DeepMind77.6%
11Mistral Medium 3Mistral AI76%
12DeepSeek-V3DeepSeek75.9%
13Gemini 1.5 ProGoogle DeepMind75.8%
14Grok 2xAI75.5%
15Amazon Nova ProAmazon75.5%
16GPT-4oOpenAI74.7%
17Llama 4 ScoutMeta AI74.3%
18ERNIE 4.5 300B-A47BBaidu74%
19Llama 3.1 405BMeta AI73.3%
20gpt-oss-20bOpenAI73.2%

Where these numbers come from

Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.

A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.

Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.

How to read this score

A high MMLU-Pro score means the model carries a lot of factual knowledge and can reason over it under multiple-choice conditions. The ten-option format makes lucky guessing worth 10% rather than 25%, so the spread between models is wider and more meaningful than on the original MMLU.

Where it is weak

It is still multiple choice, which is nothing like the open-ended work most people give a model. A model can pick the right option and still be unable to explain it, and the format rewards recall over judgement.