7 models
Ranked by Terminal-Bench Hard
Reasoning Tools JSON 26.5

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Open weights 15.9

Mistral's largest general model for enterprise agents, coding, and multilingual reasoning

mistral/mistral-large-2512 2024-11-01 262.144K context $0.5/M input $1.5/M output
12 providers
Reasoning Tools JSON 13.6

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Tools 8.3

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-08-06 2024-08-06 128K context $2.5/M input $10/M output
7 providers
Tools 8.3

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-11-20 2024-11-20 128K context $2.5/M input $10/M output
8 providers
Reasoning Tools JSON 4.5

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

google/gemini-2.5-flash-lite 2025-06-17 1.04858M context $0.1/M input $0.4/M output
18 providers
3.8

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

mistral/mistral-medium-2505 2025-05-07 131.072K context $0.4/M input $2/M output
10 providers