Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Thinking Kimi model for slower research passes, planning, and hard technical questions
Affordable GPT-4.1 lane for fast coding help and structured extraction
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
Open GPT reasoning model for self-hosted agents and controllable deployments
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
Classic open reasoning model for transparent math, coding, and deliberate problem solving
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Small omni GPT for cheap multimodal assistance and production-scale traffic
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Open multimodal Gemma instruction model for multilingual text generation and image understanding
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 3.1 | 131.072K | $0.2 | $1 | — | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 3.1 | 131.072K | $0.4 | $2 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 3.1 | 1M | $0.3 | $2.5 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 2.4 | 262.144K | $0.5 | $1.5 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 2.4 | 262.144K | $0.25 | $0.75 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 2.3 | 262.144K | $0.01 | $0.03 | — | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 2.1 | 262.144K | $0.15 | $1.2 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 2.0 | 256K | Free | Free | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1.8 | 1.04758M | $0.2 | $0.8 | — | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 1.8 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 1.7 | 200K | $0.55 | $2.2 | — | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 1.4 | 131.072K | $0.05 | $0.2 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 1.4 | 262.144K | $0.075 | $0.3 | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 1.4 | 262.144K | $0.15 | $0.6 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 1.4 | 131.072K | $0.15 | $0.6 | — | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 1.4 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 1.3 | 131.072K | $0.23 | $2.3 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.2 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1.2 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.2 | 1.04758M | $0.05 | $0.2 | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 1.1 | 262.144K | $0.2 | $0.2 | — | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 1.1 | 256K | Free | Free | — | |||
| DeepSeek-R1deepseek/deepseek-r1 | 1.1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 1.0 | 128K | $0.075 | $0.3 | — | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 1.0 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| GPT-4o miniopenai/gpt-4o-mini | 1.0 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 0.9 | 40.96K | $0.08 | $0.28 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 0.9 | 200K | $1.1 | $4.4 | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 0.9 | 131.072K | $0.227 | $0.91 | — | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 0.9 | 81.92K | $0.2 | $2.4 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 0.8 | 163.84K | $0.25 | $1 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 0.8 | 131.072K | $0.117 | $0.455 | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 0.8 | 131.072K | $0.1 | $0.1 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 0.6 | 262.144K | $0.15 | $0.15 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 0.6 | 128K | $0.2 | $0.696 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 0.6 | 262.144K | $0.075 | $0.075 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 0.5 | 131.072K | $0.05 | $0.08 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 0.5 | 327.68K | $0.1 | $0.3 | — | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 0.3 | 65.536K | Free | Free | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 0.3 | 131.072K | $0.1 | $0.32 | — | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 0.1 | 131.072K | $0.05 | $0.1 | 2025-03-12 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 0.1 | 131.072K | $0.08 | $0.16 | 2025-03-12 |