Ministral 8B is an 8B parameter model featuring a unique interleaved sliding-window attention pattern for faster, memory-efficient inference. Designed for edge use cases, it supports up to 128k context length...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
No provider description is available for this model yet.
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
No provider description is available for this model yet.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
No provider description is available for this model yet.
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
No provider description is available for this model yet.
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
No provider description is available for this model yet.
No provider description is available for this model yet.
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Mistral: Ministral 8Bmistralai/ministral-8b | 128K | $0.11 | $0.11 | — | |||
| Sakana: Fugu Maxsakana/fugu-max | 1M | $2 | $6 | — | |||
| openai/gpt-5.6-solopenrouter/openai/gpt-5.6-sol | 1.05M | $2 | $10 | — | |||
| Anthropic: Claude Opus 4.6 (batch)anthropic/claude-opus-4.6:batch | 1M | $2.5 | $12.5 | — | |||
| Z.ai: GLM 4.5Vz-ai/glm-4.5v | 65.536K | $0.6 | $1.8 | — | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 128K | $0.075 | $0.3 | — | |||
| IBM: Granite 4.2 8Bibm-granite/granite-4.2-8b | 131.072K | $0.06 | $0.25 | — | |||
| glm-4.5vzai/glm-4.5v | 128K | $0.6 | $1.8 | — | |||
| Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| OpenAI: GPT-5.4 Pro (batch)openai/gpt-5.4-pro:batch | 1.05M | $15 | $90 | — | |||
| SpaceXAI: Grok 4.6x-ai/grok-4.6 | 500K | $2 | $6 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 262.144K | $0.021 | $0.063 | — | |||
| Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | 200K | $15 | $75 | — | |||
| OpenAI: GPT-5.6 Terra (batch)openai/gpt-5.6-terra:batch | 1.05M | $1 | $6 | — | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1M | Free | Free | — | |||
| gpt-chat-latestazure_ai/gpt-chat-latest | 272K | $5 | $30 | — | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 400K | $0.625 | $5 | — | |||
| model-routerazure_ai/model-router | 200K | $0.14 | — | — | |||
| cohere-command-aazure_ai/cohere-command-a | 131.072K | $2.5 | $10 | — | |||
| xai/grok-4.3vertex_ai/xai/grok-4.3 | 200K | $1.25 | $2.5 | — | |||
| grok-4-20-reasoningazure_ai/grok-4-20-reasoning | 262K | $1.25 | $2.5 | — | |||
| xai/grok-4.6vertex_ai/xai/grok-4.6 | 524.288K | $2 | $6 | — | |||
| grok-4-20-non-reasoningazure_ai/grok-4-20-non-reasoning | 262K | $1.25 | $2.5 | — | |||
| Qwen: Qwen3 Coder Plusqwen/qwen3-coder-plus | 1M | $0.65 | $3.25 | — | |||
| us.openai.gpt-6-astrabedrock_converse/us.openai.gpt-6-astra | 1.05M | $11 | $55 | — | |||
| global.openai.gpt-6-astrabedrock_converse/global.openai.gpt-6-astra | 1.05M | $10 | $50 | — | |||
| lyria-3.5gemini/lyria-3.5 | 1.04858M | — | — | — | |||
| Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| gpt-6-astraazure/gpt-6-astra | 922K | $10 | $50 | — | |||
| Mistral: Codestral 2508mistralai/codestral-2508 | 256K | $0.3 | $0.9 | — | |||
| us/gpt-6-astraazure/us/gpt-6-astra | 922K | $11 | $55 | — | |||
| Codestral-2501azure_ai/codestral-2501 | 256K | $0.3 | $0.9 | — | |||
| MiniMax: MiniMax M2minimax/minimax-m2 | 204.8K | $0.255 | $1.02 | — | |||
| FW-Nemotron-Lightning-3.5-30B-A3Bazure_ai/fw-nemotron-lightning-3.5-30b-a3b | 262.144K | $0.06 | $0.22 | — | |||
| MAI-Thinking-1azure_ai/mai-thinking-1 | 256K | $2 | $8 | — | |||
| grok-4.6azure_ai/grok-4.6 | 200K | $2 | $6 | — | |||
| databricks-claude-fable-5-1databricks/databricks-claude-fable-5-1 | 1M | $10 | $50 | — | |||
| databricks-gemini-3-1-flash-imagedatabricks/databricks-gemini-3-1-flash-image | 131.072K | — | — | — | |||
| databricks-gemini-3-pro-imagedatabricks/databricks-gemini-3-pro-image | 65.536K | — | — | — | |||
| databricks-gemini-3-8-flashdatabricks/databricks-gemini-3-8-flash | 1.04858M | — | — | — | |||
| databricks-gemini-3-7-flashdatabricks/databricks-gemini-3-7-flash | 1.04858M | — | — | — | |||
| databricks-gemini-3-6-flashdatabricks/databricks-gemini-3-6-flash | 1.04858M | $1.875 | $9.375 | — | |||
| databricks-gemini-3-5-flashdatabricks/databricks-gemini-3-5-flash | 1.04858M | $1.875 | $11.25 | — | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 1M | $0.32 | $1.28 | — | |||
| databricks-gemini-3-5-flash-litedatabricks/databricks-gemini-3-5-flash-lite | 1.04858M | $0.375 | $3.125 | — | |||
| databricks-glm-5-3databricks/databricks-glm-5-3 | 1.04858M | $1.4 | $4.4 | — | |||
| databricks-gpt-5-6-soldatabricks/databricks-gpt-5-6-sol | 922K | $4 | $20 | — | |||
| databricks-gpt-5-6-terradatabricks/databricks-gpt-5-6-terra | 922K | $2.5 | $15 | — | |||
| databricks-gpt-5-6-lunadatabricks/databricks-gpt-5-6-luna | 922K | $1 | $6 | — | |||
| databricks-grok-4-6databricks/databricks-grok-4-6 | 500K | $2.5 | $7.5 | — |