CorX Labs

Benchmarks

Compare 159 AI models, side by side.

Context windows, prices and published benchmark scores for every major model from 42 labs — searchable, sortable, and honest about what has not been measured.

The CorX Labs benchmark index tracks 159 language models from 42 companies, including OpenAI, Anthropic, Google DeepMind, Meta, Mistral, DeepSeek, Alibaba and xAI. For each model it records the context window, the price per million input and output tokens, the licence, and published scores on 14 standard evaluations — MMLU-Pro, GPQA Diamond, AIME, SWE-bench Verified and others. Pick any two to four models to see them column by column.

Where these numbers come from

Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.

A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.

Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.

The index

Every model in one table

Search by name or maker, filter by capability, and sort by any column. Tick two or more rows to compare them properly.

Showing 159 of 159 models

AI models with context window, price per million tokens and published benchmark scores. Sortable by any column.
Compare
Gemini 3 ProGoogle DeepMindGoogle DeepMind1M$2.00$12.0091.9%95%76.2%Proprietary
Grok 4xAIxAI256K$3.00$15.0087.5%94%72%Proprietary
Claude Opus 4.5AnthropicAnthropic200K$5.00$25.0087%96%80.9%Proprietary
GPT-5OpenAIOpenAI400K$1.25$10.0085.7%94.6%74.9%Proprietary
Grok 4 FastxAIxAI2M$0.20$0.5085.7%92%Proprietary
Qwen3-MaxAlibaba QwenAlibaba Qwen262K$1.20$6.0085.4%92.3%69.6%Proprietary
Kimi K2 ThinkingMoonshot AIMoonshot AI262K$0.60$2.5084.5%94.5%71.3%Modified MIT
Gemini 2.5 ProGoogle DeepMindGoogle DeepMind1M$1.25$10.0084%86.7%63.8%Proprietary
Claude Sonnet 4.5AnthropicAnthropic1M$3.00$15.0083.4%87%77.2%Proprietary
o3OpenAIOpenAI200K$2.00$8.0083.3%88.9%69.1%Proprietary
GLM-4.6Z.ai (Zhipu)Z.ai (Zhipu)205K$0.60$2.2082.9%93.9%68%MIT
GPT-5 miniOpenAIOpenAI400K$0.25$2.0082.3%91.1%71%Proprietary
o4-miniOpenAIOpenAI200K$1.10$4.4081.4%92.7%68.1%Proprietary
Claude Opus 4.1AnthropicAnthropic200K$15.00$75.0080.9%78%74.5%Proprietary
DeepSeek-V3.1DeepSeekDeepSeek131K$0.280$1.1484.8%80.1%88.4%66%MIT
gpt-oss-120bOpenAIOpenAI131K$0.10$0.5080.9%80.1%92.5%Apache 2.0
DeepSeek-V3.2DeepSeekDeepSeek164K$0.280$0.4285%79.9%89.3%MIT
o3-miniOpenAIOpenAI200K$1.10$4.4079.7%87.3%49.3%Proprietary
GLM-4.5Z.ai (Zhipu)Z.ai (Zhipu)131K$0.60$2.2079.1%91%64.2%MIT
Gemini 2.5 FlashGoogle DeepMindGoogle DeepMind1M$0.30$2.5078.3%78%Proprietary
MiniMax-M2MiniMaxMiniMax205K$0.30$1.2078%78%69.4%MIT
o1OpenAIOpenAI200K$15.00$60.0078%79.2%48.9%Proprietary
Llama-3.1-Nemotron-Ultra-253BNVIDIANVIDIA131K$0.60$1.8076%80.1%NVIDIA Open Model
Claude Sonnet 4AnthropicAnthropic200K$3.00$15.0075.4%70.5%72.7%Proprietary
Grok 3xAIxAI131K$3.00$15.0075.4%83.9%Proprietary
GLM-4.5-AirZ.ai (Zhipu)Z.ai (Zhipu)131K$0.20$1.1075%89.4%57.6%MIT
Claude Haiku 4.5AnthropicAnthropic200K$1.00$5.0073%77%73.3%Proprietary
DeepSeek-R1DeepSeekDeepSeek131K$0.550$2.1984%71.5%79.8%49.2%MIT
gpt-oss-20bOpenAIOpenAI131K$0.05$0.2073.2%71.5%90%Apache 2.0
Seed-OSS-36BByteDance SeedByteDance Seed524K$0.15$0.6071.4%91.7%Apache 2.0
GPT-5 nanoOpenAIOpenAI400K$0.05$0.4071.2%85.2%Proprietary
Hunyuan-A13BTencentTencent262K$0.30$1.0071.2%87.3%Tencent Hunyuan Community
Qwen3-235B-A22BAlibaba QwenAlibaba Qwen131K$0.20$0.6068.2%71.1%85.7%Apache 2.0
Magistral MediumMistral AIMistral AI41K$2.00$5.0070.8%73.6%Proprietary
Hermes 4 405BNous ResearchNous Research131K$1.00$3.0070.5%78.1%Llama 3.1 Community
Llama 4 MaverickMeta AIMeta AI1M$0.22$0.8580.5%69.8%Llama 4 Community
Phi-4-reasoning-plusMicrosoftMicrosoft33K$0.070$0.3569.3%78%MIT
Magistral SmallMistral AIMistral AI41K$0.50$1.5068.2%70.7%Apache 2.0
Claude 3.7 SonnetAnthropicAnthropic200K$3.00$15.0068%61.3%62.3%Proprietary
Sonar Reasoning ProPerplexityPerplexity127K$2.00$8.0068%Proprietary
Qwen3-32BAlibaba QwenAlibaba Qwen131K$0.10$0.3066.8%81.4%Apache 2.0
EXAONE 4.0 32BLG AI ResearchLG AI Research131K$0.15$0.4566.7%85.3%EXAONE AI Model License
Llama-3.3-Nemotron-Super-49BNVIDIANVIDIA131K$0.13$0.4066.7%67.5%NVIDIA Open Model
GPT-4.1OpenAIOpenAI1M$2.00$8.0080.5%66.3%54.6%Proprietary
Grok 3 MinixAIxAI131K$0.30$0.5066.2%90.8%Proprietary
Qwen3-30B-A3BAlibaba QwenAlibaba Qwen131K$0.08$0.29065.8%80.4%Apache 2.0
QwQ-32BAlibaba QwenAlibaba Qwen131K$0.15$0.2065.2%79.5%Apache 2.0
Claude 3.5 SonnetAnthropicAnthropic200K$3.00$15.0078%65%49%Proprietary
GPT-4.1 miniOpenAIOpenAI1M$0.40$1.6065%23.6%Proprietary
Gemini 2.5 Flash-LiteGoogle DeepMindGoogle DeepMind1M$0.10$0.4064.6%49.8%Proprietary
Nemotron Nano 9B v2NVIDIANVIDIA131K$0.04$0.1664%72.1%NVIDIA Open Model
Qwen3-14BAlibaba QwenAlibaba Qwen131K$0.06$0.2464%79.3%Apache 2.0
Mistral Medium 3Mistral AIMistral AI131K$0.40$2.0076%62.4%Proprietary
Gemini 2.0 FlashGoogle DeepMindGoogle DeepMind1M$0.10$0.4077.6%62.1%Proprietary
DeepSeek-R1-Distill-Qwen-32BDeepSeekDeepSeek131K$0.12$0.1862.1%72.6%MIT
Qwen3-8BAlibaba QwenAlibaba Qwen131K$0.035$0.13862%76%Apache 2.0
o1-miniOpenAIOpenAI128K$1.10$4.4060%63.6%Proprietary
Amazon Nova PremierAmazonAmazon1M$2.50$12.5080%59.9%Proprietary
DeepSeek-V3DeepSeekDeepSeek131K$0.27$1.1075.9%59.1%MIT
Gemini 1.5 ProGoogle DeepMindGoogle DeepMind2.1M$1.25$5.0075.8%59.1%Proprietary
Llama 4 ScoutMeta AIMeta AI10M$0.11$0.3474.3%57.2%Llama 4 Community
Phi-4MicrosoftMicrosoft16K$0.070$0.14070.4%56.1%MIT
Grok 2xAIxAI131K$2.00$10.0075.5%56%Grok 2 Community
Qwen3-4BAlibaba QwenAlibaba Qwen131K$0.02$0.0655.9%73.8%Apache 2.0
GPT-4oOpenAIOpenAI128K$2.50$10.0074.7%53.6%Proprietary
Gemini 2.0 Flash-LiteGoogle DeepMindGoogle DeepMind1M$0.075$0.3071.6%51.5%Proprietary
Reka Flash 3Reka AIReka AI33K$0.10$0.3051.2%65%Apache 2.0
Llama 3.1 405BMeta AIMeta AI131K$3.50$3.5073.3%51.1%Llama 3.1 Community
Gemini 1.5 FlashGoogle DeepMindGoogle DeepMind1M$0.075$0.3067.3%51%Proprietary
Llama 3.3 70BMeta AIMeta AI131K$0.23$0.4068.9%50.5%Llama 3.3 Community
Claude 3 OpusAnthropicAnthropic200K$15.00$75.0068.5%50.4%Proprietary
GPT-4.1 nanoOpenAIOpenAI1M$0.10$0.4050.3%Proprietary
Qwen2.5-72BAlibaba QwenAlibaba Qwen131K$0.35$0.4071.1%49%Qwen License
Mistral Large 2Mistral AIMistral AI131K$2.00$6.0069.9%48%Mistral Research
GPT-4 TurboOpenAIOpenAI128K$10.00$30.0063.7%48%Proprietary
Amazon Nova ProAmazonAmazon300K$0.80$3.2075.5%46.9%Proprietary
Llama 3.1 70BMeta AIMeta AI131K$0.12$0.3066.4%46.7%Llama 3.1 Community
Mistral Small 3.2 24BMistral AIMistral AI131K$0.10$0.3069.1%46.1%Apache 2.0
Gemma 3 27BGoogle DeepMindGoogle DeepMind131K$0.10$0.2067.5%42.4%Gemma Terms of Use
Claude 3.5 HaikuAnthropicAnthropic200K$0.80$4.0065%41.6%40.6%Proprietary
GPT-4o miniOpenAIOpenAI128K$0.15$0.6063.1%40.2%Proprietary
Gemma 3 12BGoogle DeepMindGoogle DeepMind131K$0.05$0.1060.6%34.9%Gemma Terms of Use
Llama 3.1 8BMeta AIMeta AI131K$0.03$0.0548.3%32.8%Llama 3.1 Community
Kimi K2 InstructMoonshot AIMoonshot AI131K$0.60$2.5081.1%65.8%Modified MIT
ERNIE 4.5 300B-A47BBaiduBaidu131K$0.280$1.1074%Apache 2.0
Qwen2.5-32BAlibaba QwenAlibaba Qwen131K$0.08$0.2069%Apache 2.0
Command ACohereCohere256K$2.50$10.0068%CC-BY-NC
Llama 3.2 90B VisionMeta AIMeta AI131K$0.35$0.4068%Llama 3.2 Community
Solar Pro 2UpstageUpstage66K$0.50$0.5066%Proprietary
Hunyuan-LargeTencentTencent262K$0.50$1.5060.2%Tencent Hunyuan Community
DeepSeek-Coder-V2DeepSeekDeepSeek131K$0.140$0.28060%DeepSeek License
Jamba 1.6 LargeAI21 LabsAI21 Labs256K$2.00$8.0060%Jamba Open Model
Yi-Large01.AI01.AI33K$3.00$3.0060%Proprietary
dots.llm1RedNote (Xiaohongshu)RedNote (Xiaohongshu)33K$0.20$0.6060%MIT
Amazon Nova LiteAmazonAmazon300K$0.06$0.2459%Proprietary
Falcon-H1 34BTII FalconTII Falcon262K$0.15$0.4558%Falcon LLM License
Qwen2.5-7BAlibaba QwenAlibaba Qwen131K$0.025$0.0556.3%Apache 2.0
Command R+CohereCohere128K$2.50$10.0056%CC-BY-NC
Gemma 2 27BGoogle DeepMindGoogle DeepMind8.2K$0.27$0.2756%Gemma Terms of Use
Mixtral 8x22BMistral AIMistral AI66K$0.90$0.9056%Apache 2.0
Reka CoreReka AIReka AI128K$2.00$2.0055%Proprietary
Phi-3.5-MoEMicrosoftMicrosoft131K$0.08$0.1654%MIT
Amazon Nova MicroAmazonAmazon128K$0.035$0.14051.6%Proprietary
Mistral NeMo 12BMistral AIMistral AI131K$0.03$0.07050%Apache 2.0
Yi-1.5-34B01.AI01.AI33K$0.15$0.1548%Apache 2.0
Llama 3.2 11B VisionMeta AIMeta AI131K$0.055$0.05547%Llama 3.2 Community
Ministral 8BMistral AIMistral AI131K$0.10$0.1047%Mistral Research
OLMo 2 32BAllen InstituteAllen Institute4.1K$0.20$0.4047%Apache 2.0
DBRX InstructDatabricksDatabricks33K$0.75$2.2545%Databricks Open Model
GLM-4-9BZ.ai (Zhipu)Z.ai (Zhipu)131K$0.03$0.0645%GLM License
Gemma 2 9BGoogle DeepMindGoogle DeepMind8.2K$0.06$0.0645%Gemma Terms of Use
Granite 3.3 8BIBMIBM131K$0.03$0.0645%Apache 2.0
MiniCPM4 8BOpenBMBOpenBMB33K$0.02$0.0545%Apache 2.0
Falcon 3 10BTII FalconTII Falcon33K$0.05$0.1044%Falcon LLM License
Gemma 3 4BGoogle DeepMindGoogle DeepMind131K$0.02$0.0443.6%Gemma Terms of Use
Jamba 1.6 MiniAI21 LabsAI21 Labs256K$0.20$0.4043%Jamba Open Model
Command R7BCohereCohere128K$0.037$0.1542%CC-BY-NC
LFM2-8B-A1BLiquid AILiquid AI33K$0.02$0.0540%LFM Open
Snowflake ArcticSnowflakeSnowflake4.1K$0.60$1.8040%Apache 2.0
GPT-3.5 TurboOpenAIOpenAI16K$0.50$1.5038%Proprietary
OLMo 2 13BAllen InstituteAllen Institute4.1K$0.10$0.2035%Apache 2.0
Llama 3.2 3BMeta AIMeta AI131K$0.015$0.02533%Llama 3.2 Community
Falcon 180BTII FalconTII Falcon2K$1.80$1.8030%Falcon 180B TII License
Mistral 7BMistral AIMistral AI33K$0.025$0.02530%Apache 2.0
Llama 3.2 1BMeta AIMeta AI131K$0.01$0.0222%Llama 3.2 Community
OpenELM 3BAppleApple2K$0.01$0.0220%Apple Sample Code License
Claude Fable 5AnthropicAnthropic1M$10.00$50.00Proprietary
Claude Fable 5.1AnthropicAnthropic1M$10.00$50.00Proprietary
Claude Mythos 5AnthropicAnthropic1M$10.00$50.00Proprietary
Claude Mythos 5.1AnthropicAnthropic1M$10.00$50.00Proprietary
Claude Opus 5AnthropicAnthropic1M$5.00$25.00Proprietary
Claude Sonnet 5AnthropicAnthropic1M$2.00$10.00Proprietary
Codestral 25.08Mistral AIMistral AI262K$0.30$0.90Mistral AI Non-Production
CorX1.5CorX LabsCorX Labs1KApache 2.0
CorX3.8-27BCorX LabsCorX Labs33KApache 2.0
DeepSeek-R1-Distill-Llama-8BDeepSeekDeepSeek131K$0.04$0.0450.4%MIT
Devstral MediumMistral AIMistral AI131K$0.40$2.0061.6%Proprietary
ERNIE X1BaiduBaidu131K$0.280$1.10Proprietary
GLM-5.3Z.ai (Zhipu)Z.ai (Zhipu)1M$1.40$4.40GLM-5.3 License
Grok Code Fast 1xAIxAI256K$0.20$1.5070.8%Proprietary
HyperCLOVA X SEED 14BNaverNaver33K$0.08$0.16HyperCLOVA X SEED License
Kimi K3Moonshot AIMoonshot AI1M$3.00$15.00Kimi K3 License
Kimi-Dev-72BMoonshot AIMoonshot AI131K$0.290$1.1560.4%Modified MIT
Ling-1TInclusionAI (Ant)InclusionAI (Ant)131K$0.50$2.0070.4%MIT
Mercury CoderInception LabsInception Labs33K$0.25$1.00Proprietary
MiniMax-M1MiniMaxMiniMax1M$0.40$2.1086%56%Apache 2.0
Molmo 72BAllen InstituteAllen Institute4.1K$0.35$0.40Apache 2.0
Palmyra X5WriterWriter1M$0.60$6.00Proprietary
Phi-4-multimodalMicrosoftMicrosoft131K$0.05$0.10MIT
Pixtral LargeMistral AIMistral AI131K$2.00$6.00Mistral Research
Qwen2.5-Coder-32BAlibaba QwenAlibaba Qwen131K$0.070$0.16Apache 2.0
Qwen2.5-VL-72BAlibaba QwenAlibaba Qwen131K$0.40$0.40Qwen License
Qwen3-Coder-480B-A35BAlibaba QwenAlibaba Qwen262K$0.30$1.2069.6%Apache 2.0
R1-1776PerplexityPerplexity131K$2.00$8.0078%MIT
Sarvam-MSarvam AISarvam AI33K$0.10$0.30Apache 2.0
Sonar ProPerplexityPerplexity200K$3.00$15.00Proprietary
Step-3StepFunStepFun66K$0.30$1.20Apache 2.0
TriStream-SVSCorX LabsSinging voice synthesisCorX LabsApache 2.0
xLAM-2-70BSalesforce AISalesforce AI131K$0.30$0.60CC-BY-NC

The tests

What each benchmark actually measures

A score is only useful if you know what it was measuring. These are the 14 evaluations tracked here, and what each one does and does not tell you.

MMLU-Pro

12,000 reasoning-heavy multiple-choice questions across 14 academic subjects, with ten options instead of four. The harder successor to MMLU.

Read the test

GPQA Diamond

198 graduate-level physics, chemistry and biology questions written to be Google-proof. PhD holders in the matching field score about 65%.

Read the test

AIME 2025

The American Invitational Mathematics Examination — 15 problems, integer answers, no partial credit. A standard test of multi-step maths reasoning.

Read the test

MATH-500

500 competition maths problems sampled from the MATH benchmark, graded on the final answer.

Read the test

SWE-bench Verified

500 human-validated GitHub issues from real Python repositories. The model must produce a patch that makes the project's own tests pass.

Read the test

SWE-bench Pro

A harder, contamination-resistant successor to SWE-bench Verified, drawn from commercial and copyleft repositories that were never public training data.

Read the test

Terminal-Bench 2.1

The 2.1 revision of the terminal agent benchmark. Scores on it are not comparable with Terminal-Bench 4.0 — the task set changed.

Read the test

Frontier-Bench v0.1

Novel problems built to resist memorisation, scored on whether the model gets anywhere at all. Absolute numbers are low by design.

Read the test

Terminal-Bench 4.0

End-to-end tasks in a real terminal — install, build, debug, run — scored on whether the machine ends up in the required state.

Read the test

LiveCodeBench

Competitive-programming problems collected after each model's training cutoff, so contamination cannot inflate the score.

Read the test

HumanEval

164 short Python functions written from a docstring. Saturated at the frontier — kept here for continuity with older models.

Read the test

MMMU

College-level questions that require reading charts, diagrams, tables and photographs alongside the text.

Read the test

IFEval

Verifiable instructions — word counts, formats, forbidden words — checked by a program rather than a judge model.

Read the test

LMArena Elo

Elo rating from blind pairwise votes by the public on LMArena. Measures what people prefer, not what is correct.

Read the test

Method, and what to distrust

This index is a collection of published figures, not an independent evaluation. That distinction matters more than it sounds, so here is exactly what was and was not done.

What is collected

For every model: the context window and maximum output length, the input and output modalities, the licence and whether weights are downloadable, the first-party API price per million tokens, the architecture where the lab has disclosed it, and the release month. For scores: whatever the lab published in its model card, system card, technical report or launch post, plus LMArena Elo where the model has been rated.

What is not done

No evaluation was re-run. No score was estimated, interpolated from a sibling model, or carried over from a previous version. Where a lab has not published a figure the cell says Not reported and stays empty, even when that leaves a gap in an otherwise full row.

Why self-reported scores are slippery

  • The scaffold moves the number. SWE-bench Verified in particular is a measure of a whole agent — retrieval, retries, test execution — not of a model alone. Two labs reporting the same benchmark may be running very different harnesses.
  • Attempts vary. A score taken at pass@1 and one taken with majority voting over many samples are not comparable, and the difference is often larger than the gap between two models.
  • Contamination. Older benchmarks leak into training data over time. HumanEval is effectively saturated; LiveCodeBench exists precisely because it collects problems published after a model's cutoff.
  • Reasoning budgets. A model with adjustable thinking can post a much higher score at a much higher cost per answer. The price column does not capture that, because tokens spent thinking are billed as output.

Prices

Prices are the standard first-party rate per million tokens, excluding batch discounts and cached-input rates unless noted on the model's own page. For open-weight models there is no first-party price, so the figure shown is a representative third-party hosting rate and is marked as such — the weights themselves are free to download.

Corrections

Every model records the month its row was last verified. If a figure is wrong or has been superseded, send a correction with a link to the source and it will be updated.

Trademarks

Company marks are shown to identify each lab's own models. All trademarks belong to their respective owners; CorX Labs is not affiliated with, endorsed by, or sponsored by any of the companies listed. Logo files are from the lobe-icons (MIT) and simple-icons (CC0) sets.

Questions

Common questions

What is the best AI model right now?

There is no single answer, which is why this page is a table rather than a ranking. The frontier models — GPT-5, Claude Opus 4.5, Gemini 3 Pro and Grok 4 — trade places depending on the test: reasoning benchmarks like GPQA Diamond, agentic coding benchmarks like SWE-bench Verified, and human preference on LMArena all pick different winners. Sort the table by the column that matches the work you are actually doing.

Which AI model is cheapest?

Among capable models, the open-weight ones hosted by third parties are usually cheapest — DeepSeek, Qwen, GLM and gpt-oss all sit far below the frontier proprietary models. Sort by In / M or Out / M to see the current order. Note that reasoning models generate many more output tokens than their price per token suggests, so a cheap reasoning model can cost more per answer than an expensive non-reasoning one.

What does open weights mean?

The lab has published the trained parameters, so you can download the model and run it on your own hardware. It does not necessarily mean the training data or code is public, and it does not always mean unrestricted commercial use — the Licence column records the actual terms, which range from Apache 2.0 and MIT through to non-commercial and custom community licences.

Did CorX Labs run these benchmarks?

No. Every score here is a published figure from the lab that built the model or from a public leaderboard, collected and put in one table. Most benchmark numbers in this industry are self-reported, and the scaffolding around a model can move a score by more than the difference between two models. Use them to build a shortlist, then test the shortlist on your own task.

How often is this updated?

Each model row records the month it was last checked against its sources. New models are added as they are released. If you spot a figure that is out of date or wrong, tell us and it will be corrected.

Can I compare more than two models?

Yes. Tick the boxes in the leaderboard and press Compare, or open the comparison tool and add up to four models side by side.

CorX Labs

We build models too.

CorX3.8-27B is Jamaica's first large open-weight LLM — a 27B Jamaican Patois assistant with open weights under Apache 2.0. It is in this index like everything else.