Large coding-reasoning model for agentic software tasks and RL search
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Open flagship GLM for long-horizon coding agents and million-token context work
MiniMax multimodal model for long-context coding, perception, and agent planning
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Open MiniMax flagship for coding agents, office automation, and complex environments
Large coding-reasoning model for agentic software tasks and RL search
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Open coding-reasoning model for repository tasks and self-improving agents
Cohere coding model for practical software engineering and agentic edits
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Earlier MiniMax agent model for practical coding and productivity tasks
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Strong GLM coding model for agentic engineering, terminals, and repository generation
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Open multimodal Llama for strong reasoning with efficient everyday serving
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Ornith 1.0 397Bdeepreinforce/ornith-1.0-397b | 62.2 | 262.144K | — | — | 2026-06-25 | |||
| GLM-5.2zhipuai/glm-5.2 | 62.1 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| MiniMax-M3minimax/MiniMax-M3 | 59.0 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 57.2 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| StepFun: Step 3.7 Flashstepfun/step-3.7-flash | 56.3 | 256K | $0.2 | $1.15 | 2026-05-29 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 56.2 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| Ornith 1.0 35Bdeepreinforce/ornith-1.0-35b | 50.4 | 262.144K | — | — | 2026-06-25 | |||
| Poolside: Laguna XS 2.1poolside/laguna-xs-2.1 | 47.6 | 262.144K | $0.06 | $0.12 | 2026-07-02 | |||
| Ornith 1.0 9Bdeepreinforce/ornith-1.0-9b | 42.9 | 262.144K | — | — | 2026-06-25 | |||
| North Mini Codecohere/north-mini-code-1-0 | 40.2 | 256K | — | — | 2026-06-09 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 38.7 | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 36.8 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| MoonshotAI: Kimi K2.6moonshotai/kimi-k2.6 | 27.3 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| GLM-5.1zhipuai/glm-5.1 | 19.8 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | 18.0 | 1.024M | $0.948 | $1.896 | 2026-04-24 | |||
| GLM-4.6zhipuai/glm-4.6 | 9.7 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 5.2 | 1M | $0.14 | $0.59 | 2025-04-05 |