*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...
No provider description is available for this model yet.
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...
No provider description is available for this model yet.
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
No provider description is available for this model yet.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
No provider description is available for this model yet.
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
No provider description is available for this model yet.
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
No provider description is available for this model yet.
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 262.144K | $0.021 | $0.063 | — | |||
| Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | 128K | $0.075 | $0.2 | — | |||
| deepseek/deepseek-r1/communitynovita/deepseek/deepseek-r1/community | 64K | $4 | $4 | — | |||
| OpenAI: o4 Mini (batch)openai/o4-mini:batch | 200K | $0.55 | $2.2 | — | |||
| Inception: Mercury 2inception/mercury-2 | 128K | $0.25 | $0.75 | — | |||
| mistral-vibe-cli-fastmistral/mistral-vibe-cli-fast | 262.144K | $0.15 | $0.6 | — | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Qwen: Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | — | |||
| OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | 1.04758M | $1 | $4 | — | |||
| deepseek/deepseek-v3/communitynovita/deepseek/deepseek-v3/community | 64K | $0.89 | $0.89 | — | |||
| Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 1.04858M | $1.25 | $10 | — | |||
| OpenAI: gpt-oss-120b (batch)openai/gpt-oss-120b:batch | 131.072K | $0.15 | $0.6 | — | |||
| Anthropic: Claude Sonnet 5 (batch)anthropic/claude-sonnet-5:batch | 1M | $1 | $5 | — | |||
| anthropic/claude-sonnet-5deepinfra/anthropic/claude-sonnet-5 | 1M | $2 | $10 | — | |||
| gpt-chat-latestazure_ai/gpt-chat-latest | 272K | $5 | $30 | — | |||
| databricks-gpt-5-4-nanodatabricks/databricks-gpt-5-4-nano | 272K | $0.2 | $1.25 | — | |||
| model-routerazure_ai/model-router | 200K | $0.14 | — | — | |||
| minimax/minimax-m2.7-highspeednovita/minimax/minimax-m2.7-highspeed | 204.8K | $0.6 | $2.4 | — | |||
| cohere-command-aazure_ai/cohere-command-a | 131.072K | $2.5 | $10 | — | |||
| xai/grok-4.3vertex_ai/xai/grok-4.3 | 200K | $1.25 | $2.5 | — | |||
| grok-4-20-reasoningazure_ai/grok-4-20-reasoning | 262K | $1.25 | $2.5 | — | |||
| xai/grok-4.6vertex_ai/xai/grok-4.6 | 524.288K | $2 | $6 | — | |||
| grok-4-20-non-reasoningazure_ai/grok-4-20-non-reasoning | 262K | $1.25 | $2.5 | — | |||
| us.openai.gpt-6-astrabedrock_converse/us.openai.gpt-6-astra | 1.05M | $11 | $55 | — | |||
| moonshotai/kimi-k2.7-codenovita/moonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | — | |||
| global.openai.gpt-6-astrabedrock_converse/global.openai.gpt-6-astra | 1.05M | $10 | $50 | — | |||
| lyria-3.5gemini/lyria-3.5 | 1.04858M | — | — | — | |||
| OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch | 400K | $0.025 | $0.2 | — | |||
| moonshotai/Kimi-K2.7-Codedeepinfra/moonshotai/kimi-k2.7-code | 262.144K | $0.68 | $3.4 | — | |||
| MoonshotAI: Kimi K2 0711moonshotai/kimi-k2 | 131.072K | $0.57 | $2.3 | — | |||
| us/gpt-6-astraazure/us/gpt-6-astra | 922K | $11 | $55 | — | |||
| Codestral-2501azure_ai/codestral-2501 | 256K | $0.3 | $0.9 | — | |||
| FW-Nemotron-Lightning-3.5-30B-A3Bazure_ai/fw-nemotron-lightning-3.5-30b-a3b | 262.144K | $0.06 | $0.22 | — | |||
| MAI-Thinking-1azure_ai/mai-thinking-1 | 256K | $2 | $8 | — | |||
| grok-4.6azure_ai/grok-4.6 | 200K | $2 | $6 | — | |||
| databricks-claude-fable-5-1databricks/databricks-claude-fable-5-1 | 1M | $10 | $50 | — | |||
| nvidia/Nemotron-3-Nano-30B-A3Bdeepinfra/nvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | — | |||
| databricks-gpt-5-4-minidatabricks/databricks-gpt-5-4-mini | 272K | $0.75 | $4.5 | — | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 400K | $0.125 | $1 | — | |||
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 131.072K | $0.2 | $1 | — | |||
| zai-org/GLM-5deepinfra/zai-org/glm-5 | 202.752K | $0.6 | $2.08 | — | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 262.144K | $0.15 | $1.2 | — | |||
| Google: Gemini 3.8 Flash (batch)google/gemini-3.8-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| databricks-gemini-3-7-flashdatabricks/databricks-gemini-3-7-flash | 1.04858M | — | — | — | |||
| ByteDance/Seed-2.0-prodeepinfra/bytedance/seed-2.0-pro | 256K | $0.5 | $3 | — | |||
| databricks-gpt-5-4databricks/databricks-gpt-5-4 | 272K | $2.5 | $15 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 1M | $1.5 | $7.5 | — | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 198K | $0.43 | $1.75 | — | |||
| ByteDance/Seed-2.0-codedeepinfra/bytedance/seed-2.0-code | 256K | $0.5 | $3 | — | |||
| Anthropic: Claude Opus 4.5 (batch)anthropic/claude-opus-4.5:batch | 200K | $2.5 | $12.5 | — |