CorX Labs

Llama 3.3 70B

by Meta AI · United States
Open weightsTool calling131K context

405B-class quality distilled into 70B — for a year, the open-weight default.


Specification

The numbers

Maker
Meta AI
Released
2024-12
Parameters
70B
Architecture
Dense transformer
Context window
131,072 tokens
Max output
8,192 tokens
Input
Text
Output
Text
Reasoning
No
Tool calling
Yes
Knowledge cutoff
Not reported
Licence
Llama 3.3 Community
Availability
Not reported
Weights
Downloadable

Cost

Price per million tokens

Input
$0.23 / M tokens
Output
$0.40 / M tokens
Blended 3:1
$0.273

This model has open weights, so there is no first-party price. The figures above are a representative third-party hosting rate — you can also run it yourself for the cost of the hardware.

Published scores

Benchmarks

Figures published by Meta AI or taken from a public leaderboard. Row last checked 2026-08.

Llama 3.3 70B by capability category, with its rank among models reporting the same tests.
CategoryScore Rank
ReasoningGPQA Diamond50.5%70 of 83 reporting the same tests
MathsCompetition mathematics, graded on the final answerNot reported
CodingHumanEval88.4%12 of 46 reporting the same tests
KnowledgeMMLU-Pro68.9%27 of 77 reporting the same tests
MultimodalReading charts, diagrams and photographsNot reported
Instruction followingIFEval92.1%2 of 15 reporting the same tests
Human preferenceLMArena Elo125715 of 24 reporting the same tests

A category averages every benchmark in it that Llama 3.3 70B reports. The rank counts only models that report the same tests, so it never compares an average over three benchmarks against an average over one.

Every reported test

MMLU-ProKnowledge68.9%Rank 27 of 77 models reporting
GPQA DiamondReasoning50.5%Rank 70 of 83 models reporting
AIME 2025MathsNot reportedNo figure published
MATH-500MathsNot reportedNo figure published
SWE-bench VerifiedCodingNot reportedNo figure published
SWE-bench ProCodingNot reportedNo figure published
Terminal-Bench 2.1CodingNot reportedNo figure published
Frontier-Bench v0.1ReasoningNot reportedNo figure published
Terminal-Bench 4.0CodingNot reportedNo figure published
LiveCodeBenchCodingNot reportedNo figure published
HumanEvalCoding88.4%Rank 12 of 46 models reporting
MMMUMultimodalNot reportedNo figure published
IFEvalInstruction following92.1%Rank 2 of 15 models reporting
LMArena EloHuman preference1257Rank 15 of 24 models reporting

Where these numbers come from

Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.

A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.

Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.

Same maker

Other models from Meta AI

AI models with context window, price per million tokens and published benchmark scores. Sortable by any column.
Llama 4 MaverickMeta AIMeta AI1M$0.22$0.8580.5%69.8%Llama 4 Community
Llama 4 ScoutMeta AIMeta AI10M$0.11$0.3474.3%57.2%Llama 4 Community
Llama 3.2 90B VisionMeta AIMeta AI131K$0.35$0.4068%Llama 3.2 Community
Llama 3.2 3BMeta AIMeta AI131K$0.015$0.02533%Llama 3.2 Community
Llama 3.2 11B VisionMeta AIMeta AI131K$0.055$0.05547%Llama 3.2 Community
Llama 3.2 1BMeta AIMeta AI131K$0.01$0.0222%Llama 3.2 Community
Llama 3.1 405BMeta AIMeta AI131K$3.50$3.5073.3%51.1%Llama 3.1 Community
Llama 3.1 8BMeta AIMeta AI131K$0.03$0.0548.3%32.8%Llama 3.1 Community
Llama 3.1 70BMeta AIMeta AI131K$0.12$0.3066.4%46.7%Llama 3.1 Community