5 models
Ranked by AutomationBench
Reasoning Tools JSON Open weights 30.8

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

moonshotai/kimi-k3 2026-07-16 1.04858M context $3/M input $15/M output
63 providers
Reasoning Tools JSON 30.4

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

google/gemini-3.7-flash 2026-08-13 1.04858M context $0.75/M input $3.75/M output
23 providers
Reasoning Tools 26.0

Strongest Claude Opus model for coding, agents, and professional work

anthropic/claude-opus-5 2026-07-24 1M context $5/M input $25/M output
32 providers
Reasoning Tools JSON Open weights 25.1

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

deepseek/deepseek-v4-flash-0731 2026-07-31 1M context $0.035/M input $0.07/M output
43 providers
Reasoning Tools 17.4

Claude model for creative writing, analysis, and controlled agent workflows

anthropic/claude-fable-5 2026-06-09 1M context $10/M input $50/M output
37 providers