Claude model for creative writing, analysis, and controlled agent workflows
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Safety model for policy screening, moderation, and risk-aware routing workflows
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
MiniMax multimodal model for long-context coding, perception, and agent planning
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Image model for prompt-driven generation, editing, and visual design workflows
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
Cohere's stronger command model for multilingual agents and enterprise workflows
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Low-latency Gemini model for high-volume multimodal and agent workloads
Compact GPT model for low-latency assistance and high-volume workloads
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Open Nemotron omni model combining reasoning with text, vision, and audio
Qwen vision-language model for visual reasoning, documents, and agent tasks
Default frontier GPT for coding, computer use, research, and knowledge work
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open MiMo model for multimodal coding agents and long-context automation
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports
Agentic model for autonomous multi-step research, synthesis, and cited reports
Open multimodal Qwen MoE for local agents that need vision, audio, and code
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
Fast Grok coding model tuned for agentic engineering and iterative edits
Stronger Opus tier for advanced software work and high-stakes reasoning
Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Open Gemma instruction model for efficient chat and self-hosted deployments
Open Gemma instruction model for efficient chat and self-hosted deployments
Earlier Qwen multimodal workhorse for million-token agent and document tasks
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
Reranking model for improving retrieval quality in search and recommendation systems
High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications
Music generation model for full-length songs from text or images with vocals and structure
Music generation model for short 30-second clips, loops, and previews from text or image prompts
MiMo omni model for text, image, video, audio, and agents
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
Strong small GPT for coding subagents, quick tool use, and high-volume work
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Reasoning Grok for document-heavy analysis and long-horizon tool use
Grok model for agentic tool use, reasoning, coding, and live assistance
Agent-ready GPT for coding and computer-use workflows at a lower cost
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Low-latency Gemini model for high-volume multimodal and agent workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Claude Fable 5anthropic/claude-fable-5 | 1M | $10 | $50 | 2026-06-09 | |||
| Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safety | 128K | $0.2 | $0.2 | 2026-06-04 | |||
| Qwen3.7 Plusalibaba/qwen3.7-plus | 1M | $0.5 | $3 | 2026-06-02 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| Nano Banana 2google/gemini-3.1-flash-image | 131.072K | $0.5 | $60 | 2026-05-28 | |||
| Claude Opus 4.8anthropic/claude-opus-4-8 | 1M | $5 | $25 | 2026-05-28 | |||
| Command A Pluscohere/command-a-plus-05-2026 | 128K | $2.5 | $10 | 2026-05-20 | |||
| Gemini 3.5 Flashgoogle/gemini-3.5-flash | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | 1.04858M | $0.25 | $1.5 | 2026-05-07 | |||
| GPT-5.5 Instantopenai/gpt-5.5-instant | 400K | $5 | $30 | 2026-05-05 | |||
| Mistral Medium (latest)mistral/mistral-medium-latest | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Mistral Medium 3.5mistral/mistral-medium-2604 | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Nemotron 3 Nano Omni 30B A3B Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 256K | $0.2 | $0.8 | 2026-04-28 | |||
| Qwen3.6 Flashalibaba/qwen3.6-flash | 1M | $0.188 | $1.125 | 2026-04-27 | |||
| GPT-5.5openai/gpt-5.5 | 1.05M | $5 | $30 | 2026-04-23 | |||
| GPT-5.5 Proopenai/gpt-5.5-pro | 1.05M | $30 | $180 | 2026-04-23 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Deep Research Max Previewgoogle/deep-research-max-preview-04-2026 | 1.04858M | — | — | 2026-04-21 | |||
| Gemini Deep Research Previewgoogle/deep-research-preview-04-2026 | 1.04858M | — | — | 2026-04-21 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| Grok 4.3xai/grok-4.3 | 1M | $1.25 | $2.5 | 2026-04-17 | |||
| Grok Build 0.1xai/grok-build-0.1 | 256K | $1 | $2 | 2026-04-16 | |||
| Claude Opus 4.7anthropic/claude-opus-4-7 | 1M | $5 | $25 | 2026-04-16 | |||
| Gemini Robotics-ER 1.6 Previewgoogle/gemini-robotics-er-1.6-preview | 131.072K | $1 | $5 | 2026-04-14 | |||
| Muse Spark 1.1meta/muse-spark-1.1 | 1.04858M | $1.25 | $4.25 | 2026-04-08 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 E4B ITgoogle/gemma-4-E4B-it | 131.072K | $0.02 | $0.1 | 2026-04-02 | |||
| Gemma 4 E2B ITgoogle/gemma-4-E2B-it | 131.072K | $0.04 | $0.08 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Qwen3.6 Plusalibaba/qwen3.6-plus | 1M | $0.5 | $3 | 2026-04-02 | |||
| GLM-5V-Turbozhipuai/glm-5v-turbo | 200K | $5 | $22 | 2026-04-01 | |||
| Llama Nemotron Rerank VL 1B v2nvidia/llama-nemotron-rerank-vl-1b-v2 | 128K | — | — | 2026-03-31 | |||
| Gemini 3.1 Flash Live Previewgoogle/gemini-3.1-flash-live-preview | 131.072K | $0.75 | $4.5 | 2026-03-26 | |||
| Lyria 3 Pro Previewgoogle/lyria-3-pro-preview | 131.072K | — | — | 2026-03-25 | |||
| Lyria 3 Clip Previewgoogle/lyria-3-clip-preview | 131.072K | — | — | 2026-03-25 | |||
| MiMo-V2-Omnixiaomi/mimo-v2-omni | 262.144K | $0.14 | $0.28 | 2026-03-18 | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Grok 4.20 (Reasoning)xai/grok-4.20-0309-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| Grok 4.20 (Non-Reasoning)xai/grok-4.20-0309-non-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| GPT-5.4openai/gpt-5.4 | 1.05M | $2.5 | $15 | 2026-03-05 | |||
| GPT-5.4 Proopenai/gpt-5.4-pro | 1.05M | $30 | $180 | 2026-03-05 | |||
| GPT-5.3 Chat (latest)openai/gpt-5.3-chat-latest | 128K | $1.75 | $14 | 2026-03-03 | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| Qwen3.5 Flashalibaba/qwen3.5-flash | 1M | $0.029 | $0.287 | 2026-02-23 |