Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
Qwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to the Qwen3.8 Max Preview. It is...
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December...
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | 200K | $15 | $75 | — | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 262.144K | Free | Free | — | |||
| Z.ai: GLM 5.1z-ai/glm-5.1 | 200K | $0.966 | $3.036 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 200K | $1.1 | $4.4 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 512.288K | $0.6 | $3.6 | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | 128K | $0.075 | $0.2 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 262.144K | $0.26 | $2.08 | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 131.072K | $0.227 | $0.91 | — | |||
| MoonshotAI: Kimi K2 0905moonshotai/kimi-k2-0905 | 262.144K | $0.6 | $2.5 | — | |||
| Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | 262.144K | $1.5 | $7.5 | — | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 200K | $3 | $15 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 262.144K | $0.25 | $0.75 | — | |||
| Anthropic: Claude Opus 4.7anthropic/claude-opus-4.7 | 1M | $5 | $25 | — | |||
| OpenAI: GPT-5 Imageopenai/gpt-5-image | 400K | $10 | $10 | — | |||
| Qwen: Qwen3.8 Max (0803)qwen/qwen3.8-max | 1M | $2 | $6 | — | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 202.752K | $1.2 | $4 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 524.288K | $0.3 | $1.2 | — | |||
| OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch | 1.05M | $0.1 | $0.6 | — | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 262.144K | Free | Free | — | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 262.144K | $0.55 | $3.5 | — | |||
| Mistral: Devstral 2 2512mistralai/devstral-2512 | 262.144K | $0.4 | $2 | — | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 131.072K | $0.4 | $2 | — | |||
| Inception: Mercury 2inception/mercury-2 | 128K | $0.25 | $0.75 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 1M | $1.25 | $2.5 | — | |||
| MoonshotAI: Kimi K3 (batch)moonshotai/kimi-k3:batch | 1.04858M | $3 | $15 | — | |||
| SpaceXAI: Grok 4.20x-ai/grok-4.20 | 2M | $1.25 | $2.5 | — | |||
| OpenAI: GPT-6 Astra (batch)openai/gpt-6-astra:batch | 1.05M | $5 | $25 | — | |||
| Qwen: Qwen3.8 Max (0902)qwen/qwen3.8-max-0902 | 1M | $2 | $6 | — | |||
| SpaceXAI: Grok 4.5x-ai/grok-4.5 | 500K | $2 | $6 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 327.68K | $0.1 | $0.3 | — | |||
| Amazon: Nova Pro 1.0amazon/nova-pro-v1 | 300K | $0.8 | $3.2 | — | |||
| Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | 131.072K | $0.061 | $0.4 | — | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 262.144K | Free | Free | — | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 256K | Free | Free | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 1.04858M | Free | Free | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 262.144K | $0.075 | $0.625 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 131.072K | $0.15 | $0.6 | — | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 524.288K | $1 | $4.05 | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 262.144K | $0.2 | $0.2 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 1M | $3 | $15 | — | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 1.04858M | Free | Free | — | |||
| Qwen: Qwen3 Coder 480B A35B (free)qwen/qwen3-coder:free | 262K | Free | Free | — | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 262.144K | $0.3 | $1.2 | — | |||
| IBM: Granite 4.2 8Bibm-granite/granite-4.2-8b | 131.072K | $0.06 | $0.25 | — | |||
| DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek/deepseek-v4-flash-0731:batch | 1.04858M | $0.11 | $0.33 | — | |||
| OpenAI: GPT-5.6 Terra (batch)openai/gpt-5.6-terra:batch | 1.05M | $1 | $6 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 262.144K | $0.01 | $0.03 | — |