6 models
Ranked by Terminal-Bench Hard
Reasoning Tools JSON 26.5

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Tools Open weights 18.9

Mistral's coding-agent model for repository work, terminal tasks, and software fixes

mistral/devstral-2512 2025-12-09 262.144K context $0.4/M input $2/M output
13 providers
Open weights 15.9

Mistral's largest general model for enterprise agents, coding, and multilingual reasoning

mistral/mistral-large-2512 2024-11-01 262.144K context $0.5/M input $1.5/M output
12 providers
Reasoning Tools JSON 13.6

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Reasoning Tools Open weights 6.1

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

deepseek/deepseek-r1 2025-01-20 64K context $0.7/M input $2.5/M output
14 providers
Reasoning Tools JSON 4.5

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

google/gemini-2.5-flash-lite 2025-06-17 1.04858M context $0.1/M input $0.4/M output
18 providers