Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Affordable GPT-4.1 lane for fast coding help and structured extraction
Thinking Kimi model for slower research passes, planning, and hard technical questions
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Open GPT reasoning model for self-hosted agents and controllable deployments
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Classic open reasoning model for transparent math, coding, and deliberate problem solving
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Small omni GPT for cheap multimodal assistance and production-scale traffic
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Open multimodal Gemma instruction model for multilingual text generation and image understanding
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 3.1 | 131.072K | $0.4 | $2 | — | |||
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 3.1 | 131.072K | $0.2 | $1 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 3.1 | 131.072K | Free | Free | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 2.4 | 262.144K | $0.5 | $1.5 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 2.4 | 262.144K | $0.25 | $0.75 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 2.3 | 262.144K | $0.01 | $0.03 | — | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 2.1 | 262.144K | $0.15 | $1.2 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 2.0 | 256K | Free | Free | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1.8 | 1.04758M | $0.2 | $0.8 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 1.8 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 1.7 | 200K | $0.55 | $2.2 | — | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 1.4 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 1.4 | 262.144K | $0.075 | $0.3 | — | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 1.4 | 131.072K | $0.05 | $0.2 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 1.4 | 131.072K | $0.15 | $0.6 | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 1.4 | 262.144K | $0.15 | $0.6 | — | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 1.3 | 131.072K | $0.23 | $2.3 | — | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1.2 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.2 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.2 | 1.04758M | $0.05 | $0.2 | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 1.1 | 262.144K | $0.2 | $0.2 | — | |||
| DeepSeek-R1deepseek/deepseek-r1 | 1.1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 1.1 | 256K | Free | Free | — | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 1.0 | 128K | $0.075 | $0.3 | — | |||
| GPT-4o miniopenai/gpt-4o-mini | 1.0 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 1.0 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 0.9 | 81.92K | $0.2 | $2.4 | — | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 0.9 | 40.96K | $0.08 | $0.28 | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 0.9 | 131.072K | $0.227 | $0.91 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 0.9 | 200K | $1.1 | $4.4 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 0.8 | 163.84K | $0.25 | $1 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 0.8 | 131.072K | $0.117 | $0.455 | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 0.8 | 131.072K | $0.1 | $0.1 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 0.6 | 262.144K | $0.15 | $0.15 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 0.6 | 128K | $0.2 | $0.696 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 0.6 | 262.144K | $0.075 | $0.075 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 0.5 | 131.072K | $0.05 | $0.08 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 0.5 | 327.68K | $0.1 | $0.3 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 0.3 | 131.072K | $0.1 | $0.32 | — | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 0.3 | 65.536K | Free | Free | — | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 0.1 | 131.072K | $0.05 | $0.1 | 2025-03-12 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 0.1 | 131.072K | $0.08 | $0.16 | 2025-03-12 |