DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Open flagship GLM for long-horizon coding agents and million-token context work
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
MiniMax multimodal model for long-context coding, perception, and agent planning
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Strong GLM coding model for agentic engineering, terminals, and repository generation
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
Faster GLM-5 lane for coding agents that need lower latency
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Prior MiniMax coding model for agent workflows, office edits, and automation
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...
Mature GLM model for dependable coding, reasoning, and structured agent tasks
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Smaller Qwen coder for efficient local agents and repo-level fixes
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1.04858M | $0.04 | $0.08 | 2026-07-31 | |||
| MoonshotAI: Kimi K3moonshotai/kimi-k3 | 1.04858M | $2.303 | $11.55 | 2026-07-16 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.71 | $3.5 | 2026-06-12 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 1.024M | $0.067 | $0.134 | 2026-04-24 | |||
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | 1.024M | $0.799 | $1.598 | 2026-04-24 | |||
| MoonshotAI: Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| GLM-5V-Turbozhipuai/glm-5v-turbo | 200K | $5 | $22 | 2026-04-01 | |||
| GLM-5-Turbozhipuai/glm-5-turbo | 200K | $0.9 | $3.7 | 2026-03-16 | |||
| NVIDIA: Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | 262.144K | $0.085 | $0.4 | 2026-03-11 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MoonshotAI: Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.45 | $2.25 | 2026-01 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| DeepSeek: DeepSeek V3.2deepseek/deepseek-v3.2 | 163.84K | $0.269 | $0.4 | 2025-12-01 | |||
| OpenAI: gpt-oss-120bopenai/gpt-oss-120b | 131.072K | $0.037 | $0.17 | 2025-08-05 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 |