9 models
Ranked by Terminal-Bench Hard
Reasoning Tools JSON 37.9

Coding-optimized GPT model for repository edits, reviews, and agentic software work

openai/gpt-5-codex 2025-09-15 400K context $1.1/M input $9/M output
13 providers
Reasoning Open weights 32.6

StepFun flash model for efficient multimodal reasoning, coding, and tool use

stepfun/step-3.5-flash-2603 2026-04-02 256K context $0.1/M input $0.3/M output
6 providers
Open weights 27.3

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token....

stepfun/step-3.5-flash 2026-01-29 262.144K context $0.1/M input $0.3/M output
14 providers
Reasoning Tools JSON 26.5

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
20.5

Flagship Qwen3 model for coding agents, complex reasoning, and tool use

alibaba/qwen3-max 2025-09-23 262.144K context $1.2/M input $6/M output
20 providers
Reasoning Tools JSON 13.6

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Tools 8.3

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded...

openai/gpt-4o-2024-11-20 2024-11-20 128K context $2.5/M input $10/M output
8 providers
Tools 8.3

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is...

openai/gpt-4o-2024-08-06 2024-08-06 128K context $2.5/M input $10/M output
7 providers
3.8

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

mistral/mistral-medium-2505 2025-05-07 131.072K context $0.4/M input $2/M output
10 providers