3 models
Ranked by AutomationBench
Reasoning Tools JSON Open weights 30.8

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

moonshotai/kimi-k3 2026-07-16 1.04858M context $3/M input $15/M output
63 providers
Reasoning Tools 27.3

Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows

alibaba/qwen3.8-max-preview 2026-07-19 1M context $2/M input $6/M output
6 providers
Reasoning Tools 17.4

Claude model for creative writing, analysis, and controlled agent workflows

anthropic/claude-fable-5 2026-06-09 1M context $10/M input $50/M output
37 providers