News · High impactBack to News

Claude Opus 5 Benchmark Results Revealed: Top Rankings Across Frontier Model Comparison Define New Baseline

Anthropic's newly released Claude Opus 5 achieves standout scores across key AI benchmarks, leading Humanity's Last Exam (HLE) at 64.7%, scoring 91.59% on MMLU-Pro, and reaching 97.00% on SWE-bench Verified, outperforming models including GPT-5.6 Sol, Claude Fable 5, Kimi K3, and Opus 4.8.

Academic and Advanced Reasoning Performance

Anthropic has officially released full benchmark evaluation results for its next-generation flagship model, Claude Opus 5. The results demonstrate that the model sets new records across multiple benchmark suites evaluating complex academic reasoning, domain expertise, and multi-step logic.

Academic and Reasoning Metrics

  • 64.7%

    Humanity's Last Exam (HLE) score, ranking first among all public 2026 frontier models.

  • 91.59%

    MMLU-Pro accuracy, demonstrating broad interdisciplinary expert knowledge.

  • 95.0%

    GPQA Diamond frontier science evaluation score, proving advanced logical reasoning.

On Humanity's Last Exam (HLE), known as a rigorous evaluation designed to test frontier reasoning against unsearchable problems, Claude Opus 5 achieved a leading score of 64.7%. This performance confirms the model's capacity for deep deduction when tackling novel complex academic challenges. Additionally, on MMLU-Pro, Opus 5 reached an accuracy of 91.59%, demonstrating comprehensive mastery across diverse academic disciplines.

Frontier Multi-Model Comparison and Competitive Landscape

To provide a comprehensive assessment of Claude Opus 5's positioning, independent evaluation organizations compared it against leading frontier models, including GPT-5.6 Sol, Claude Fable 5, Kimi K3, Qwen3.7 Max, DeepSeek-V4 Pro, and predecessor Claude Opus 4.8.

SWE-bench Verified Cross-Model Resolution Rate (%)

Claude Opus 597%
GPT-5.6 Sol96.2%
Claude Fable 595%
Kimi K393.4%
Qwen3.7 Max90.4%
DeepSeek-V4 Pro88.6%
Claude Opus 4.884.2%

In head-to-head metrics, Claude Opus 5 establishes strong competitive performance across academic reasoning and software engineering:

Evaluated ModelRelease DateHumanity's Last Exam (HLE)MMLU-ProGPQA DiamondSWE-bench Verified
Claude Opus 52026-0764.7%91.6%95.0%97.0%
GPT-5.6 Sol2026-0762.8%91.2%94.6%96.2%
Claude Fable 52026-0663.2%90.8%94.5%95.0%
Kimi K32026-0761.5%89.8%92.4%93.4%
Qwen3.7 Max2026-0558.2%89.6%91.5%90.4%
DeepSeek-V4 Pro2026-0456.4%88.5%90.2%88.6%
Claude Opus 4.82026-0554.1%86.4%88.5%84.2%

Comparative data shows that compared to predecessor Opus 4.8, Opus 5 improves by 10.6 percentage points on HLE and 12.8 percentage points on SWE-bench. Against flagship Claude Fable 5 (95.0%) and OpenAI GPT-5.6 Sol (96.2%), Opus 5 achieves superior performance across agentic programming and academic reasoning.

Agentic Coding and Software Engineering Benchmark

In real-world software engineering and agentic coding benchmarks, Claude Opus 5 continues the Claude family's performance leadership. By pairing natively with Thinking by Default and granular Effort controls, the model significantly improves autonomous debugging and codebase refactoring capabilities.

On the SWE-bench Verified evaluation measuring resolution of real GitHub issues, Claude Opus 5 achieved a 97.00% resolution rate. This proves that Opus 5 excels not only at localized algorithmic prompts, but also at navigating multi-file dependencies and executing complex repository repairs.

Independent Verification and Overall Impact

Beyond Anthropic's official technical report, independent benchmarking organizations including Artificial Analysis, Vals AI, and Vellum completed evaluations of Claude Opus 5. Independent testing confirmed that under High / Max Effort parameters, the model sustains top-tier success rates in long-context consistency and tool execution.

View Evaluation Methodology & Environment Notes

Benchmark metrics are sourced from Anthropic official technical announcements and third-party evaluations published in July 2026. SWE-bench Verified used standard single-agent evaluation harnesses; GPQA and HLE were measured with Thinking enabled by default.

Overall, Claude Opus 5's top performance across HLE, MMLU-Pro, and SWE-bench re-establishes a benchmark standard in frontier AI reasoning and agentic deployment. Delivering major benchmark advances while maintaining $5/$25 per million tokens pricing provides solid foundation for enterprise AI agent deployments.

Next step

Keep tracking Anthropic

Continue along the same topic.

Open entity record