Understanding Benchmarks
iAs of: September 2026
Benchmark scores cited here reflect the state as of September 2026. Frontier models are released roughly every month — find current leaderboards at llm-stats.com, artificialanalysis.ai, and swebench.com.
Knowledge
How do you know if one LLM is "better" than another? The AI industry uses benchmarks -- standardized tests that evaluate models across different categories. But not every benchmark measures the same thing, and no single benchmark gives the complete picture.
The most important benchmarks you should know:
MMLU (Massive Multitask Language Understanding)
- What it measures: Domain knowledge across 57 areas (mathematics, law, medicine, history, etc.)
- Format: Multiple-choice questions
- Typical scores: Top models score in the 88-92 % range (GPT-5.6, Claude Opus 5, Gemini 3.1 Pro all fall in this range)
- Limitation: Multiple choice tests recognition, not application. A model can pass MMLU and still explain things poorly. MMLU is now often described as "saturated" — top models barely differentiate themselves anymore.
HumanEval
- What it measures: Ability to solve programming tasks
- Format: Writing Python functions based on docstrings
- Typical scores: Top models solve 90-95% of tasks
- Limitation: Python only, relatively simple tasks. Real-world software development is far more complex.
SWE-bench
- What it measures: Ability to solve real GitHub issues (fix bugs, implement features)
- Format: The model receives an issue and must produce a working pull request
- Typical scores (as of September 2026): Top agents solve 90-97 % of issues on SWE-bench Verified — leading: Claude Opus 5 (96–97 %), GPT-5.6 Sol (96.2 %), Claude Mythos 5 (95.5 %, limited availability), Claude Fable 5 (95 %)
- Strength: Measures real software engineering skills, not just isolated coding tasks
!Caution: SWE-bench Verified is contaminated
In spring 2026, OpenAI's audit revealed that all tested frontier models (GPT-5.2, Claude Opus 4.5, Gemini 3 Flash) could verbatim reproduce parts of the "gold patches" or problem statements — meaning the tasks likely leaked into training data. OpenAI no longer reports Verified scores and recommends SWE-bench Pro, where top scores currently sit around 46 %. When comparing models, always check which SWE-bench variant is being cited.
Arena ELO (Chatbot Arena)
- What it measures: User preference in blind comparisons
- Format: Two models answer the same question anonymously, the user picks the better answer
- Strength: Measures real user satisfaction, not just technical metrics
- Limitation: Subjective, influenced by response length and style
Understand
Model Comparison: Open-Source LLMs
Hover over an axis for detail values. Click the legend to show/hide models.
Models (click to show/hide)
Why individual benchmarks can be misleading
A model can lead on MMLU and still perform worse in everyday use than a model with a lower score. Why?
- Contamination: Benchmark questions may have appeared in the training data
- Overfitting: Models can be optimized for specific benchmarks
- Real tasks are different: MMLU asks multiple choice, but you need detailed explanations. HumanEval tests simple functions, but you need complex systems.
That's why it's important to consider multiple benchmarks together and run your own tests for your specific use case.
Useful Comparison Platforms
- artificialanalysis.ai: Compares models by speed, price, and quality. Especially useful for cost optimization.
- lmcouncil.ai: Automated evaluation of LLMs on real-world tasks with multiple evaluators.
Apply
When you need to choose the right LLM for a project, follow this approach:
- Define the use case: What exactly should the model do? (Write code? Summarize text? Advise customers?)
- Identify relevant benchmarks: Code -> HumanEval + SWE-bench. General knowledge -> MMLU. User satisfaction -> Arena ELO.
- Check price and speed: artificialanalysis.ai shows cost per token and latency.
- Run your own evaluation: Send 20-50 typical queries to 2-3 models and compare the results.
*The Best Strategy
Don't trust any single benchmark. Instead, create an eval set with 30-50 questions that reflect your real use case, and test 2-3 candidate models with it. It takes half a day and saves months of frustration.
Beyond Individual Benchmarks: The Path to AGI
Individual benchmarks measure isolated capabilities. But how close are LLMs to general intelligence? AI researcher Dan Hendrycks (UC Berkeley, known for MMLU) has proposed a framework based on the Cattell-Horn-Carroll Theory (CHC) — an established intelligence model from psychology.
The idea: Human intelligence consists of multiple cognitive dimensions, and AGI (Artificial General Intelligence) would be achieved when an AI system reaches at least the level of an educated adult in all of these dimensions:
| Dimension | What it measures | LLMs today |
|---|---|---|
| Knowledge | Factual and domain knowledge | Strong (MMLU 88-92 %) |
| Reasoning | Logical inference | Strong (especially with Chain-of-Thought) |
| Working Memory | Actively holding information | Good (limited by Context Window) |
| Processing Speed | Fast information processing | Very strong |
| Creativity | Generating new ideas | Surprisingly good |
| Long-term Memory | Retrieving permanently learned info | Weak — LLMs forget between sessions |
| Perception | Processing sensory information | Growing (multimodal) |
| Social Cognition | Understanding others' intentions and emotions | Superficial |
iThe 'Jagged Profile'
LLMs have an uneven profile: superhuman in knowledge and speed, but with critical gaps in long-term memory and true world understanding. AGI requires all dimensions to reach a minimum level — not just some.
This makes benchmarks like MMLU or HumanEval one piece of the puzzle, but not the whole picture. A model can score 90% on MMLU and still fail at tasks requiring long-term memory over several days.
Reflect
A new LLM advertises having the highest MMLU score of all time. What is the correct conclusion?