4 models
Ranked by Terminal-Bench Hard
Reasoning Tools JSON 26.5

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Reasoning Tools Open weights 17.4

Fast Mistral production model for chat, extraction, and cost-sensitive agents

mistral/mistral-small-2603 2026-03-16 256K context $0.15/M input $0.6/M output
12 providers
Reasoning Tools JSON 13.6

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Tools Open weights 3.0

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

meta/llama-3.3-70b-instruct 2024-12-06 128K context $0.1/M input $0.32/M output
25 providers