8 models
Ranked by SWE-Bench Pro
Reasoning Tools 69.2

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents

anthropic/claude-opus-4-8 2026-05-28 1M context $5/M input $25/M output
44 providers
Reasoning Tools 64.3

Stronger Opus tier for advanced software work and high-stakes reasoning

anthropic/claude-opus-4-7 2026-04-16 1M context $5/M input $25/M output
40 providers
Reasoning Tools JSON Open weights 62.1

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools JSON 59.1

Agent-ready GPT for coding and computer-use workflows at a lower cost

openai/gpt-5.4 2026-03-05 1.05M context $2.5/M input $15/M output
42 providers
Reasoning Tools JSON 58.6

Default frontier GPT for coding, computer use, research, and knowledge work

openai/gpt-5.5 2026-04-23 1.05M context $5/M input $30/M output
46 providers
Reasoning Tools 51.9

High-end Claude for difficult coding, planning, and slower expert reasoning

anthropic/claude-opus-4-6 2026-02-05 1M context $5/M input $25/M output
37 providers
Reasoning Tools JSON Open weights 18.0

Open MoE flagship with million-token context for coding and long agent runs

deepseek/deepseek-v4-pro 2026-04-24 1M context $0.435/M input $0.87/M output
62 providers
Reasoning Tools 14.9

Claude workhorse for coding agents, careful analysis, and production cost control

anthropic/claude-sonnet-4-6 2026-02-17 1M context $3/M input $15/M output
45 providers