5 models
Ranked by Terminal-Bench 2.1
Reasoning Tools JSON Open weights 88.3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3 2026-07-16 1.04858M context $2.303/M input $11.55/M output
63 providers
Reasoning Tools JSON Open weights 82.7

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools Open weights 66.0

MiniMax multimodal model for long-context coding, perception, and agent planning

minimax/MiniMax-M3 2026-06-01 1.04858M context $0.3/M input $1.2/M output
37 providers
Reasoning Tools JSON Open weights 64.7

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

deepseek/deepseek-v4-pro 2026-04-24 1.024M context $0.779/M input $1.558/M output
62 providers
Reasoning Tools JSON Open weights 64.3

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

moonshotai/kimi-k2.6 2026-04-21 262.144K context $0.95/M input $4/M output
64 providers