2 models
Ranked by GPQA Diamond
Reasoning Tools JSON 95.5

Quality-first multi-agent model for hard research, analysis, and competitions

sakana/fugu-ultra 2026-06-15 1M context $5/M input $30/M output
11 providers
Reasoning Tools Open weights 86.6

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

xiaomi/mimo-v2.5-pro 2026-04-22 1.04858M context $0.435/M input $0.87/M output
24 providers