4 models
Ranked by SWE-Atlas Refactoring
Reasoning Tools 48.6

Stronger Opus tier for advanced software work and high-stakes reasoning

anthropic/claude-opus-4-7 2026-04-16 1M context $5/M input $25/M output
40 providers
Reasoning Tools JSON 44.8

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

openai/gpt-5.5 2026-04-23 1.05M context $5/M input $30/M output
46 providers
Reasoning Tools JSON 33.8

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

google/gemini-3.1-pro-preview 2026-02-19 1.04858M context $2/M input $12/M output
31 providers
Reasoning Tools 32.2

Claude workhorse for coding agents, careful analysis, and production cost control

anthropic/claude-sonnet-4-6 2026-02-17 1M context $3/M input $15/M output
45 providers