Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
No provider description is available for this model yet.
No provider description is available for this model yet.
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...
Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
No provider description is available for this model yet.
No provider description is available for this model yet.
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
No provider description is available for this model yet.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
No provider description is available for this model yet.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 200K | $3 | $15 | — | |||
| qwen3-vl-32b-instructqwen_ai_platform/qwen3-vl-32b-instruct | 131.072K | $0.16 | $0.64 | — | |||
| qwen3-vl-32b-thinkingqwen_ai_platform/qwen3-vl-32b-thinking | 131.072K | $0.16 | $2.87 | — | |||
| qwen3-vl-plusqwen_ai_platform/qwen3-vl-plus | 260.096K | — | — | — | |||
| zai-org/glm-5v-turbonovita/zai-org/glm-5v-turbo | 204.8K | $1.2 | $4 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 400K | $0.1 | $0.625 | — | |||
| qwen3.5-plusqwen_ai_platform/qwen3.5-plus | 991.808K | — | — | — | |||
| qwen3.7-maxqwen_ai_platform/qwen3.7-max | 991.808K | $2.5 | $7.5 | — | |||
| Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | 128K | $0.075 | $0.2 | — | |||
| Meta: Muse Glimmer 30B (batch)meta/muse-glimmer-30b:batch | 131.072K | $0.175 | $0.75 | — | |||
| TheDrummer: Cydonia 24B V4.1thedrummer/cydonia-24b-v4.1 | 131.072K | $0.3 | $0.5 | — | |||
| qwen3.8-maxqwen_ai_platform/qwen3.8-max | 991.808K | $2 | $6 | — | |||
| qwq-plusqwen_ai_platform/qwq-plus | 98.304K | $0.8 | $2.4 | — | |||
| databricks-deepseek-v4-flash-0731databricks/databricks-deepseek-v4-flash-0731 | 1M | $0.14 | $0.28 | — | |||
| databricks-deepseek-v4-pro-0813databricks/databricks-deepseek-v4-pro-0813 | 1M | $1.32 | $3.96 | — | |||
| zai-org/GLM-5.3-Flashfriendliai/zai-org/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1.04758M | $0.2 | $0.8 | — | |||
| GigaChat-2gigachat/gigachat-2 | 128K | — | — | — | |||
| google/gemma-4-31B-itdeepinfra/google/gemma-4-31b-it | 262.144K | $0.13 | $0.38 | — | |||
| zai-org/GLM-4.7deepinfra/zai-org/glm-4.7 | 202.752K | $0.4 | $1.75 | — | |||
| Mistral: Mistral Medium 3mistralai/mistral-medium-3 | 131.072K | $0.4 | $2 | — | |||
| claude-fable-5-1@defaultvertex_ai-anthropic_models/claude-fable-5-1@default | 1M | $10 | $50 | — | |||
| glm-5.2zai/glm-5.2 | 1M | $1.4 | $4.4 | — | |||
| MiniMaxAI/MiniMax-M2.7-Turbodeepinfra/minimaxai/minimax-m2.7-turbo | 196.608K | $0.38 | $1.7 | — | |||
| Qwen/Qwen3.5-9Bdeepinfra/qwen/qwen3.5-9b | 262.144K | $0.1 | $0.15 | — | |||
| gemma-4-31bcerebras/gemma-4-31b | 131.072K | $0.99 | $1.49 | — | |||
| Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | 262.144K | $0.3 | $1 | — | |||
| OpenAI: GPT-5.6 Terra (batch)openai/gpt-5.6-terra:batch | 1.05M | $1 | $6 | — | |||
| OpenAI: GPT-5.4 (batch)openai/gpt-5.4:batch | 1.05M | $1.25 | $7.5 | — | |||
| google.gemma-4-e2bbedrock_mantle/google.gemma-4-e2b | 128K | $0.04 | $0.08 | — | |||
| qwen3.5-122b-a10blibertai/qwen3.5-122b-a10b | 262.144K | $0.25 | $1.75 | — | |||
| openai/gpt-oss-120b-Ultradeepinfra/openai/gpt-oss-120b-ultra | 131.072K | $0.2 | $0.95 | — | |||
| deepseek-ai/DeepSeek-V4-Flashdeepinfra/deepseek-ai/deepseek-v4-flash | 1.04858M | $0.09 | $0.18 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 1M | $1.5 | $7.5 | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 204.8K | $0.3 | $1.2 | — | |||
| anthropic/claude-haiku-4-5deepinfra/anthropic/claude-haiku-4-5 | 200K | $1 | $5 | — | |||
| XiaomiMiMo/MiMo-V2.5-Prodeepinfra/xiaomimimo/mimo-v2.5-pro | 1.04858M | $1 | $3 | — | |||
| Anthropic: Claude Sonnet 5 (batch)anthropic/claude-sonnet-5:batch | 1M | $1 | $5 | — | |||
| minimax/minimax-m2.7-highspeednovita/minimax/minimax-m2.7-highspeed | 204.8K | $0.6 | $2.4 | — | |||
| Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 1.04858M | $1.25 | $10 | — | |||
| OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | 1.04758M | $1 | $4 | — | |||
| fallback_generalizationsunknown/fallback_generalizations | Not documented | — | — | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 262.144K | $0.021 | $0.063 | — | |||
| google/gemini-3.1-prodeepinfra/google/gemini-3.1-pro | 1M | $2 | $12 | — | |||
| google/gemma-4-26B-A4B-itdeepinfra/google/gemma-4-26b-a4b-it | 262.144K | $0.07 | $0.34 | — | |||
| zai-org/glm-5.1novita/zai-org/glm-5.1 | 204.8K | $1.38 | $4.4 | — | |||
| Qwen/Qwen3.6-27Bdeepinfra/qwen/qwen3.6-27b | 262.144K | $0.32 | $3.2 | — | |||
| anthropic/claude-opus-4-7deepinfra/anthropic/claude-opus-4-7 | 1M | $5 | $25 | — | |||
| moonshotai/kimi-k2.6novita/moonshotai/kimi-k2.6 | 262.144K | $0.8 | $3.4 | — | |||
| gpt-oss-20bdarkbloom/gpt-oss-20b | 131.072K | $0.015 | $0.07 | — |