5 models
Ranked by Terminal-Bench Hard
Reasoning Tools JSON 37.9

Coding-optimized GPT model for repository edits, reviews, and agentic software work

openai/gpt-5-codex 2025-09-15 400K context $1.1/M input $9/M output
13 providers
Reasoning Tools JSON Open weights 35.6

Newer StepFun flash model for faster agents, coding, and multimodal prompts

stepfun/step-3.7-flash 2026-05-29 256K context $0.185/M input $1.11/M output
17 providers
Reasoning Tools JSON 26.5

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Reasoning Tools Open weights 17.4

Fast Mistral production model for chat, extraction, and cost-sensitive agents

mistral/mistral-small-2603 2026-03-16 256K context $0.15/M input $0.6/M output
12 providers
Reasoning Tools JSON 13.6

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers