← Back to all datasets

Tech & AI · Live telemetry

LLM Knowledge & Coding Benchmarks

Published MMLU and HumanEval scores for leading large language models, as reported by each vendor. Not independently re-measured.

Unit
score (0–100)
  • GPT-4o88.7
  • Claude 3.5 Sonnet88.7
  • Llama 3.1 405B88.6
  • DeepSeek-V388.5
  • Gemini 1.5 Pro85.9
  • Mistral Large 284
Embed

Embed this chart page anywhere with an iframe.

<iframe src="https://axiostats.com/datasets/llm-leaderboard/" title="AxioStats — LLM Knowledge & Coding Benchmarks" width="100%" height="480" loading="lazy"></iframe>
API endpoint

The same data as this page, served as static JSON — no key required.

curl https://axiostats.com/api/datasets/llm-leaderboard.json
https://axiostats.com/api/datasets/llm-leaderboard.json

Executive summary

LLM Knowledge & Coding Benchmarks: GPT-4o tops the table at 88.7 MMLU, ahead of Claude 3.5 Sonnet (88.7). The spread from first to last (Mistral Large 2, 84) is 4.7 points.

Scores are as published by each vendor in their reports (not re-measured). MMLU measures knowledge accuracy; HumanEval measures code generation. See the 'source' column for the exact reference.

Source & metadata
Updated
2024-12-01
Methodology
Scores as published by each vendor in technical reports and launch announcements (GPT-4o, Claude 3.5 Sonnet, Llama 3.1, DeepSeek-V3, Gemini 1.5, Mistral Large 2). MMLU measures knowledge accuracy; HumanEval measures code generation (pass@1). AxioStats does not re-run benchmarks.

Data

ModelMMLUHumanEvalSource
GPT-4o88.790.2OpenAI GPT-4o report
Claude 3.5 Sonnet88.792Anthropic announcement
Llama 3.1 405B88.689Meta Llama 3.1 blog
DeepSeek-V388.582.6DeepSeek-V3 report
Gemini 1.5 Pro85.984.1Gemini 1.5 report
Mistral Large 28481.1Mistral announcement