Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Open GPT reasoning model for self-hosted agents and controllable deployments
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Classic open reasoning model for transparent math, coding, and deliberate problem solving
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Open GPT reasoning model for self-hosted agents and controllable deployments
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
Open multimodal Gemma instruction model for multilingual text generation and image understanding
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 14.2 | 262.144K | $0.01 | $0.03 | — | |||
| Cohere: Command Acohere/command-a | 13.9 | 256K | $2.5 | $10 | — | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 13.6 | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | 13.6 | 1M | Free | Free | — | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 13.6 | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 13.6 | 262.144K | Free | Free | — | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 12.7 | 131.072K | $0.23 | $2.3 | — | |||
| OpenAI: gpt-oss-120b (batch)openai/gpt-oss-120b:batch | 12.3 | 131.072K | $0.15 | $0.6 | — | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 12.3 | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| IBM: Granite 4.2 8Bibm-granite/granite-4.2-8b | 11.8 | 131.072K | $0.06 | $0.25 | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 11.5 | 262.144K | $0.15 | $0.6 | — | |||
| Inception: Mercury 2inception/mercury-2 | 11.5 | 128K | $0.25 | $0.75 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 11.5 | 262.144K | $0.075 | $0.3 | — | |||
| DeepSeek-R1deepseek/deepseek-r1 | 11.4 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 11.0 | 200K | $1.1 | $4.4 | — | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 10.9 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| Qwen: Qwen3 Coder Nextqwen/qwen3-coder-next | 10.1 | 262.144K | $0.12 | $0.8 | — | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 9.8 | 81.92K | $0.2 | $2.4 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 9.7 | 262.144K | $0.5 | $1.5 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 9.7 | 163.84K | $0.25 | $1 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 9.7 | 262.144K | $0.25 | $0.75 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 9.6 | 1.04758M | $0.05 | $0.2 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 9.6 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 9.4 | 65.536K | Free | Free | — | |||
| Mistral: Devstral 2 2512mistralai/devstral-2512 | 9.4 | 262.144K | $0.4 | $2 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 9.3 | 131.072K | $0.1 | $0.32 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 9.3 | 128K | $0.2 | $0.696 | — | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 9.0 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 9.0 | 131.072K | $0.05 | $0.2 | — | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 8.9 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 7.8 | 131.072K | $0.05 | $0.08 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 7.8 | 131.072K | $0.15 | $0.6 | — | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 7.2 | 40.96K | $0.08 | $0.28 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 6.5 | 327.68K | $0.1 | $0.3 | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 6.4 | 131.072K | $0.227 | $0.91 | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 6.0 | 262.144K | $0.2 | $0.2 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 5.5 | 262.144K | $0.075 | $0.075 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 5.5 | 262.144K | $0.15 | $0.15 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 5.2 | 131.072K | $0.117 | $0.455 | — | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 4.9 | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 4.8 | 131.072K | $0.1 | $0.1 | — | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 3.8 | 131.072K | $0.05 | $0.1 | 2025-03-12 |