Comprehensive performance comparison across major AI labs · Data sourced from official model cards and benchmark publications
Last updated: February 19, 2026 · Click column headers to sort
Lab
Category
Model | ReasoningGPQA Diamond | KnowledgeMMLU | KnowledgeMMMLU | CodingHumanEval | CodingSWE-bench Ve… | MathAIME 2024 | MathMATH-500 | ReasoningHLE | ReasoningARC-AGI-2 | CodingLiveCodeBench | MultimodalMMMU | Instruction FollowingIFEval | ReasoningARC-AGI-1 | EQEQ-Bench 3 (… |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Claude 3.7 Sonnet Anthropic · 2025-02 | 68 | 88 | — | 93.7 | 62.3 | — | 80 | — | — | — | — | 86.5 | 32 | 1083.7 |
Claude Opus 4.1 Anthropic · 2025-04 | 79.2 | — | 89.5 | 92 | 72.8 | — | 94.8 | — | — | — | — | 91.4 | 60 | 1407.9 |
Claude Opus 4.5 Anthropic · 2025-10 | 87 | — | 90.8 | — | 80.9 | 100 | 97.5 | 32 | 37.6 | — | 78.5 | 93.1 | 76 | 1683.1 |
Claude Opus 4.6 Anthropic · 2026-02 | 91.3 | — | — | — | 80.8 | — | 97.6 | 40 | 75.2 | — | — | — | 85 | 1961 |
Claude Sonnet 4.5 Anthropic · 2025-10 | 83.4 | — | 89.1 | — | 82 | 87 | 96.2 | 19.8 | — | 64 | 75.4 | 92.8 | 64 | 1501.4 |
Claude Sonnet 4.6 Anthropic · 2026-02 | 74.1 | 79.1 | — | — | 79.6 | — | 97.8 | 19.1 | 58.3 | — | — | — | — | — |
DeepSeek V3 DeepSeek · 2024-12 | 59.1 | 88.5 | — | 82.6 | 42 | — | 90.2 | — | — | — | — | 87.5 | — | 1048.7 |
DeepSeek-R1 DeepSeek · 2025-01 | 71.5 | 90.8 | — | — | 49.2 | 79.8 | 97.3 | — | — | 65.9 | — | 83.3 | 45 | 1185.6 |
Gemini 2.5 Flash Google · 2025-04 | 70.5 | — | 85.1 | — | 49.6 | 83 | 91.2 | — | — | — | 68 | 87.5 | 42 | 1074.2 |
Gemini 2.5 Pro Google · 2025-03 | 84 | — | 89.2 | — | 63.8 | 92 | 95.2 | 21.6 | — | 70.4 | 75.8 | 89.5 | 63 | 1347.8 |
Gemini 3 Flash Google · 2025-12 | 82.1 | — | 87.5 | — | 62 | 90 | 93.5 | — | — | — | 74.2 | 89 | 55 | — |
Gemini 3 Pro Google · 2025-11 | 91.9 | — | 91.8 | — | 76.2 | 100 | 97 | 45.8 | 31.1 | 79.5 | 81 | 91.5 | 80 | 1629.5 |
Gemini 3.1 Pro Google · 2026-02 | 94.3 | — | 92.6 | — | 80.6 | — | — | 44.4 | 77.1 | — | — | — | — | — |
Gemini Deep Think Google · 2026-02 | 93.8 | — | — | — | — | — | — | 48.4 | 84.6 | — | — | — | — | — |
GPT-4.1 OpenAI · 2025-04 | 66.3 | 90.2 | — | 91.5 | 54.6 | — | 90.2 | — | — | — | — | 88.4 | 52 | 1137.7 |
GPT-4o OpenAI · 2024-05 | 53.6 | 88.7 | — | 90.2 | 38.4 | — | 76.6 | — | — | — | 69.1 | 86.1 | 21 | 1321.9 |
GPT-5 OpenAI · 2025-08 | 88.4 | — | — | — | 74.9 | 94.6 | 97.3 | 35.2 | 18 | — | 84.2 | 92.5 | 72 | 1456.9 |
GPT-5.1 OpenAI · 2025-08 | 88.1 | — | 91 | — | 76.3 | 100 | 97.8 | — | 17.6 | — | 85.4 | — | 75 | 1727.6 |
GPT-5.2 OpenAI · 2025-12 | 92.4 | — | 91 | — | 80 | 100 | 98.5 | — | 52.9 | — | 86.5 | — | 82 | 1637 |
Grok 3 xAI · 2025-02 | 68.2 | 88.5 | — | 89.3 | 48.5 | 86.7 | 93 | — | — | — | — | — | 34 | 1180.6 |
Grok 4 xAI · 2025-07 | 87.5 | — | 86.6 | — | 75 | 95 | 97 | 25.4 | — | — | 76.5 | — | 71 | 1131.7 |
Grok 4.20 xAI · 2026-02 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Kimi K2 Thinking Moonshot · 2025-07 | 84.5 | — | — | — | 71.3 | 99.1 | 97 | 44.9 | — | 83.1 | — | — | 68 | 1622.5 |
Llama 4 Maverick Meta · 2025-04 | 69.8 | 88.4 | — | 85.5 | — | — | 86 | — | — | — | 73.4 | 90 | — | 833.2 |
Llama 4 Scout Meta · 2025-04 | 57.2 | 85.8 | — | 81.2 | — | — | 79.6 | — | — | — | 69.4 | 87.6 | — | 626.8 |
o3 OpenAI · 2025-04 | 83.3 | — | — | — | 69.1 | 98.4 | 96.7 | — | — | — | 82.9 | 91.8 | 87.5 | 1500 |
o4-mini OpenAI · 2025-04 | 81.4 | — | — | — | 68.1 | 99.5 | 96.3 | — | — | — | 79.6 | 90.2 | 72 | 1210.6 |
Qwen 3 Alibaba · 2025-04 | 71.1 | 89.5 | — | 88.4 | — | 87.5 | 95 | — | — | 62.5 | — | 88.2 | — | 1167.9 |
Data sourced from official model cards, blog posts, arXiv papers, and independent evaluations. Curated by Natural20 and BuseyBench.
Founded by Matthew A. Mishak, Esq. — Harvard Business School Executive Education Graduate, MIT Sloan Artificial Intelligence Graduate.
LegalTek.ai proves you don't have to choose between speed and care, scale and quality, efficiency and ethics. Dedicated to closing the justice gap through ethical AI adoption.
Mapped to and operationalizing ABA Formal Opinion 512 (the ABA does not endorse vendor frameworks). COUNSEL stands for: Confidentiality, Oversight, Understanding, Notification, Scrutiny, Equity, and Lifetime Learning.
A managed AI service for legal professionals with human oversight, ethical guardrails, and COUNSEL Framework compliance.
Website: https://legaltek.ai | Twitter: @legaltek_ai