8 models
Ranked by Terminal-Bench 2.1
Reasoning Tools JSON Open weights 88.3

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

moonshotai/kimi-k3 2026-07-16 1.04858M context $3/M input $15/M output
63 providers
Reasoning Tools JSON 84.7

Cost-efficient GPT-5.6 model for fast, high-volume workloads

openai/gpt-5.6-luna 2026-07-09 1.05M context $0.2/M input $1.2/M output
36 providers
Reasoning Tools JSON 83.3

xAI's Grok model for chat, coding, agentic tools, and lower hallucination risk

xai/grok-4.5 2026-07-08 500K context $2/M input $6/M output
26 providers
Reasoning Tools JSON Open weights 82.7

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools 70.8

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

meituan/longcat-2.0 2026-06-30 1M context $0.3/M input $1.2/M output
5 providers
Reasoning Tools JSON Open weights 65.1

Strong GLM coding model for agentic engineering, terminals, and repository generation

zhipuai/glm-5.1 2026-04-07 200K context $1.4/M input $4.4/M output
45 providers
Reasoning Tools JSON Open weights 64.7

Open MoE flagship with million-token context for coding and long agent runs

deepseek/deepseek-v4-pro 2026-04-24 1M context $0.435/M input $0.87/M output
62 providers
Reasoning Tools JSON Open weights 64.3

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

moonshotai/kimi-k2.6 2026-04-21 262.144K context $0.95/M input $4/M output
64 providers