6 models
Ranked by SWE-Bench Multilingual
Reasoning Tools 89.5

Strongest Claude Opus model for coding, agents, and professional work

anthropic/claude-opus-5 2026-07-24 1M context $5/M input $25/M output
32 providers
Open weights 78.9

Large coding-reasoning model for agentic software tasks and RL search

deepreinforce/ornith-1.0-397b 2026-06-25 262.144K context Input not listed Output not listed
Tools JSON 78.3

Everyday Claude agent model for coding, planning, browsing, and general work

anthropic/claude-sonnet-5 2026-06-30 1M context $2/M input $10/M output
35 providers
Reasoning Tools JSON 78.0

xAI's Grok model for chat, coding, agentic tools, and lower hallucination risk

xai/grok-4.5 2026-07-08 500K context $2/M input $6/M output
26 providers
Open weights 69.3

Large coding-reasoning model for agentic software tasks and RL search

deepreinforce/ornith-1.0-35b 2026-06-25 262.144K context Input not listed Output not listed
Open weights 52.0

Open coding-reasoning model for repository tasks and self-improving agents

deepreinforce/ornith-1.0-9b 2026-06-25 262.144K context Input not listed Output not listed