Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen vision-language model for visual reasoning, documents, and agent tasks
Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space
Open MiMo model for multimodal coding agents and long-context automation
Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports
Image model for prompt-driven generation, editing, and visual design workflows
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Agentic model for autonomous multi-step research, synthesis, and cited reports
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Open multimodal Qwen MoE for local agents that need vision, audio, and code
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
Safety model for policy screening, moderation, and risk-aware routing workflows
Stronger Opus tier for advanced software work and high-stakes reasoning
Fast Grok coding model tuned for agentic engineering and iterative edits
Low-latency speech generation with steerable prompts and expressive audio tags
Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Strong GLM coding model for agentic engineering, terminals, and repository generation
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Earlier Qwen multimodal workhorse for million-token agent and document tasks
Open Gemma instruction model for efficient chat and self-hosted deployments
Open Gemma instruction model for efficient chat and self-hosted deployments
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
Video model for prompt-guided generation, editing, and motion workflows
Reranking model for improving retrieval quality in search and recommendation systems
High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications
Music generation model for short 30-second clips, loops, and previews from text or image prompts
Music generation model for full-length songs from text or images with vocals and structure
Nemotron model for efficient reasoning, coding, and specialized AI agents
MiMo omni model for text, image, video, audio, and agents
Low-latency M2.7 variant for interactive coding plans and agent loops
Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks
Open MiniMax flagship for coding agents, office automation, and complex environments
Strong small GPT for coding subagents, quick tool use, and high-volume work
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron multimodal model for visual reasoning and agentic AI workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Faster GLM-5 lane for coding agents that need lower latency
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Grok model for agentic tool use, reasoning, coding, and live assistance
Reasoning Grok for document-heavy analysis and long-horizon tool use
Agent-ready GPT for coding and computer-use workflows at a lower cost
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
Low-latency Gemini model for high-volume multimodal and agent workloads
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| Gemini Embedding 2google/gemini-embedding-2 | 8.192K | $0.2 | — | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Deep Research Max Previewgoogle/deep-research-max-preview-04-2026 | 1.04858M | — | — | 2026-04-21 | |||
| GPT-Image-2openai/gpt-image-2 | Not documented | $5 | $30 | 2026-04-21 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Gemini Deep Research Previewgoogle/deep-research-preview-04-2026 | 1.04858M | — | — | 2026-04-21 | |||
| Hy3 previewtencent/hy3-preview | 256K | $0.066 | $0.26 | 2026-04-20 | |||
| Qwen3.6 Max Previewalibaba/qwen3.6-max-preview | 262.144K | $1.3 | $7.8 | 2026-04-20 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| Grok 4.3xai/grok-4.3 | 1M | $1.25 | $2.5 | 2026-04-17 | |||
| Nemotron 3 Content Safetynvidia/nemotron-3-content-safety | 128K | — | — | 2026-04-16 | |||
| Claude Opus 4.7anthropic/claude-opus-4-7 | 1M | $5 | $25 | 2026-04-16 | |||
| Grok Build 0.1xai/grok-build-0.1 | 256K | $1 | $2 | 2026-04-16 | |||
| Gemini 3.1 Flash TTS Previewgoogle/gemini-3.1-flash-tts-preview | 8.192K | $1 | $20 | 2026-04-15 | |||
| Gemini Robotics-ER 1.6 Previewgoogle/gemini-robotics-er-1.6-preview | 131.072K | $1 | $5 | 2026-04-14 | |||
| Muse Spark 1.1meta/muse-spark-1.1 | 1.04858M | $1.25 | $4.25 | 2026-04-08 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 256K | $0.1 | $0.3 | 2026-04-02 | |||
| Qwen3.6 Plusalibaba/qwen3.6-plus | 1M | $0.5 | $3 | 2026-04-02 | |||
| Gemma 4 E4B ITgoogle/gemma-4-E4B-it | 131.072K | $0.02 | $0.1 | 2026-04-02 | |||
| Gemma 4 E2B ITgoogle/gemma-4-E2B-it | 131.072K | $0.04 | $0.08 | 2026-04-02 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| GLM-5V-Turbozhipuai/glm-5v-turbo | 200K | $5 | $22 | 2026-04-01 | |||
| Veo 3.1 Lite Previewgoogle/veo-3.1-lite-generate-preview | 1.024K | — | — | 2026-03-31 | |||
| Llama Nemotron Rerank VL 1B v2nvidia/llama-nemotron-rerank-vl-1b-v2 | 128K | — | — | 2026-03-31 | |||
| Gemini 3.1 Flash Live Previewgoogle/gemini-3.1-flash-live-preview | 131.072K | $0.75 | $4.5 | 2026-03-26 | |||
| Lyria 3 Clip Previewgoogle/lyria-3-clip-preview | 131.072K | — | — | 2026-03-25 | |||
| Lyria 3 Pro Previewgoogle/lyria-3-pro-preview | 131.072K | — | — | 2026-03-25 | |||
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | — | — | 2026-03-24 | |||
| MiMo-V2-Omnixiaomi/mimo-v2-omni | 262.144K | $0.14 | $0.28 | 2026-03-18 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| MiMo-V2-Proxiaomi/mimo-v2-pro | 1.04858M | $0.435 | $0.87 | 2026-03-18 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron VoiceChatnvidia/nemotron-voicechat | 128K | — | — | 2026-03-16 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| GLM-5-Turbozhipuai/glm-5-turbo | 200K | $0.9 | $3.7 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Grok 4.20 (Non-Reasoning)xai/grok-4.20-0309-non-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| Grok 4.20 (Reasoning)xai/grok-4.20-0309-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| GPT-5.4openai/gpt-5.4 | 1.05M | $2.5 | $15 | 2026-03-05 | |||
| GPT-5.4 Proopenai/gpt-5.4-pro | 1.05M | $30 | $180 | 2026-03-05 | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| GPT-5.3 Chat (latest)openai/gpt-5.3-chat-latest | 128K | $1.75 | $14 | 2026-03-03 |