gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
Affordable GPT-4.1 lane for fast coding help and structured extraction
Thinking Kimi model for slower research passes, planning, and hard technical questions
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Open GPT reasoning model for self-hosted agents and controllable deployments
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Classic open reasoning model for transparent math, coding, and deliberate problem solving
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Small omni GPT for cheap multimodal assistance and production-scale traffic
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
Open multimodal Gemma instruction model for multilingual text generation and image understanding
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 3.1 | 131.072K | Free | Free | — | |||
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 3.1 | 131.072K | $0.2 | $1 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 3.1 | 1M | $0.3 | $2.5 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 2.4 | 262.144K | $0.25 | $0.75 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 2.4 | 262.144K | $0.5 | $1.5 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 2.3 | 262.144K | $0.01 | $0.03 | — | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 2.1 | 262.144K | $0.15 | $1.2 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 2.0 | 256K | Free | Free | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 1.8 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1.8 | 1.04758M | $0.2 | $0.8 | — | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 1.7 | 200K | $0.55 | $2.2 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 1.4 | 262.144K | $0.075 | $0.3 | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 1.4 | 262.144K | $0.15 | $0.6 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 1.4 | 131.072K | $0.15 | $0.6 | — | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 1.4 | 131.072K | $0.05 | $0.2 | — | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 1.4 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 1.3 | 131.072K | $0.23 | $2.3 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.2 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.2 | 1.04758M | $0.05 | $0.2 | — | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1.2 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 1.1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 1.1 | 256K | Free | Free | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 1.1 | 262.144K | $0.2 | $0.2 | — | |||
| GPT-4o miniopenai/gpt-4o-mini | 1.0 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 1.0 | 128K | $0.075 | $0.3 | — | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 1.0 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 0.9 | 40.96K | $0.08 | $0.28 | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 0.9 | 131.072K | $0.227 | $0.91 | — | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 0.9 | 81.92K | $0.2 | $2.4 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 0.9 | 200K | $1.1 | $4.4 | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 0.8 | 131.072K | $0.1 | $0.1 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 0.8 | 131.072K | $0.117 | $0.455 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 0.8 | 163.84K | $0.25 | $1 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 0.6 | 262.144K | $0.075 | $0.075 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 0.6 | 128K | $0.2 | $0.696 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 0.6 | 262.144K | $0.15 | $0.15 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 0.5 | 327.68K | $0.1 | $0.3 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 0.5 | 131.072K | $0.05 | $0.08 | — | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 0.3 | 65.536K | Free | Free | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 0.3 | 131.072K | $0.1 | $0.32 | — | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 0.1 | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 0.1 | 131.072K | $0.05 | $0.1 | 2025-03-12 |