4 models
Ranked by GPQA Diamond
Reasoning Tools 94.2

Stronger Opus tier for advanced software work and high-stakes reasoning

anthropic/claude-opus-4-7 2026-04-16 1M context $5/M input $25/M output
40 providers
Reasoning Tools JSON 93.6

Default frontier GPT for coding, computer use, research, and knowledge work

openai/gpt-5.5 2026-04-23 1.05M context $5/M input $30/M output
46 providers
Reasoning Tools JSON 92.8

Agent-ready GPT for coding and computer-use workflows at a lower cost

openai/gpt-5.4 2026-03-05 1.05M context $2.5/M input $15/M output
42 providers
Reasoning Tools JSON Open weights 91.2

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers