13 models
Ranked by Terminal-Bench Hard
Reasoning Tools JSON Open weights 35.6

Newer StepFun flash model for faster agents, coding, and multimodal prompts

stepfun/step-3.7-flash 2026-05-29 256K context $0.185/M input $1.11/M output
17 providers
Open weights 27.3

StepFun flash lane for quick multimodal reasoning and coding assistance

stepfun/step-3.5-flash 2026-01-29 256K context $0.1/M input $0.3/M output
14 providers
Reasoning Tools Open weights 25.0

Late GLM-4 workhorse for coding agents, reasoning, and structured tasks

zhipuai/glm-4.6 2025-09-30 204.8K context $0.6/M input $2.2/M output
18 providers
20.5

Flagship Qwen3 model for coding agents, complex reasoning, and tool use

alibaba/qwen3-max 2025-09-23 262.144K context $1.2/M input $6/M output
20 providers
Tools Open weights 18.9

Mistral's coding-agent model for repository work, terminal tasks, and software fixes

mistral/devstral-2512 2025-12-09 262.144K context $0.4/M input $2/M output
13 providers
Reasoning Tools Open weights 17.4

Fast Mistral production model for chat, extraction, and cost-sensitive agents

mistral/mistral-small-2603 2026-03-16 256K context $0.15/M input $0.6/M output
12 providers
Open weights 15.9

Mistral's largest general model for enterprise agents, coding, and multilingual reasoning

mistral/mistral-large-2512 2024-11-01 262.144K context $0.5/M input $1.5/M output
12 providers
Tools Open weights 15.2

Smaller Qwen coder for efficient local agents and repo-level fixes

alibaba/qwen3-coder-30b-a3b-instruct 2025-04 262.144K context $0.45/M input $2.25/M output
13 providers
Tools 8.3

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-08-06 2024-08-06 128K context $2.5/M input $10/M output
7 providers
Tools 8.3

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-11-20 2024-11-20 128K context $2.5/M input $10/M output
8 providers
Reasoning Tools Open weights 6.1

Classic open reasoning model for transparent math, coding, and deliberate problem solving

deepseek/deepseek-r1 2025-01-20 128K context $0.7/M input $2.5/M output
14 providers
3.8

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

mistral/mistral-medium-2505 2025-05-07 131.072K context $0.4/M input $2/M output
10 providers
Tools Open weights 3.0

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

meta/llama-3.3-70b-instruct 2024-12-06 128K context $0.1/M input $0.32/M output
25 providers