Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Thinking Kimi model for slower research passes, planning, and hard technical questions
Affordable GPT-4.1 lane for fast coding help and structured extraction
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Open GPT reasoning model for self-hosted agents and controllable deployments
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Classic open reasoning model for transparent math, coding, and deliberate problem solving
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
Small omni GPT for cheap multimodal assistance and production-scale traffic
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Open multimodal Gemma instruction model for multilingual text generation and image understanding
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 3.1 | 131.072K | $0.4 | $2 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 3.1 | 1M | $0.3 | $2.5 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 3.1 | 131.072K | Free | Free | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 2.4 | 262.144K | $0.25 | $0.75 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 2.4 | 262.144K | $0.5 | $1.5 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 2.3 | 262.144K | $0.01 | $0.03 | — | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 2.1 | 262.144K | $0.15 | $1.2 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 2.0 | 256K | Free | Free | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1.8 | 1.04758M | $0.2 | $0.8 | — | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 1.8 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 1.7 | 200K | $0.55 | $2.2 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 1.4 | 262.144K | $0.075 | $0.3 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 1.4 | 131.072K | $0.15 | $0.6 | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 1.4 | 262.144K | $0.15 | $0.6 | — | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 1.4 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 1.4 | 131.072K | $0.05 | $0.2 | — | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 1.3 | 131.072K | $0.23 | $2.3 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.2 | 1.04758M | $0.05 | $0.2 | — | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1.2 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.2 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 1.1 | 262.144K | $0.2 | $0.2 | — | |||
| DeepSeek-R1deepseek/deepseek-r1 | 1.1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 1.1 | 256K | Free | Free | — | |||
| GPT-4o miniopenai/gpt-4o-mini | 1.0 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 1.0 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 1.0 | 128K | $0.075 | $0.3 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 0.9 | 200K | $1.1 | $4.4 | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 0.9 | 131.072K | $0.227 | $0.91 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 0.8 | 131.072K | $0.117 | $0.455 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 0.8 | 163.84K | $0.25 | $1 | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 0.8 | 131.072K | $0.1 | $0.1 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 0.6 | 262.144K | $0.15 | $0.15 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 0.6 | 262.144K | $0.075 | $0.075 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 0.6 | 128K | $0.2 | $0.696 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 0.5 | 327.68K | $0.1 | $0.3 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 0.5 | 131.072K | $0.05 | $0.08 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 0.3 | 131.072K | $0.1 | $0.32 | — | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 0.1 | 131.072K | $0.05 | $0.1 | 2025-03-12 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 0.1 | 131.072K | $0.08 | $0.16 | 2025-03-12 |