Small GPT-5 for responsive agents, coding help, and everyday automation
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Small omni GPT for cheap multimodal assistance and production-scale traffic
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
Compact GPT model for low-latency assistance and high-volume workloads
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| GPT-5 Miniopenai/gpt-5-mini | 15.6 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 15.6 | 400K | $0.125 | $1 | — | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 15.3 | 40.96K | $0.08 | $0.28 | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 14.4 | 262.144K | $0.2 | $0.2 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 14.4 | 256K | Free | Free | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 13.8 | 256K | Free | Free | — | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 13.8 | 131.072K | $0.227 | $0.91 | — | |||
| GPT-4openai/gpt-4 | 13.1 | 8.192K | $30 | $60 | 2023-11-06 | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 12.1 | 81.92K | $0.2 | $2.4 | — | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 11.9 | 65.536K | Free | Free | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 11.9 | 131.072K | $0.1 | $0.32 | — | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 11.4 | 128K | $0.075 | $0.3 | — | |||
| GPT-4o miniopenai/gpt-4o-mini | 11.4 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 11.1 | 1.04758M | $0.05 | $0.2 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 11.1 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| OpenAI: GPT-3.5 Turbo (batch)openai/gpt-3.5-turbo:batch | 10.7 | 16.385K | $0.25 | $0.75 | — | |||
| GPT-3.5-turboopenai/gpt-3.5-turbo | 10.7 | 16.385K | $0.5 | $1.5 | 2023-03-01 | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 9.7 | 262.144K | $0.075 | $0.075 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 9.7 | 262.144K | $0.15 | $0.15 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 9.0 | 131.072K | $0.117 | $0.455 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 8.2 | 327.68K | $0.1 | $0.3 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 5.4 | 131.072K | $0.05 | $0.08 | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 4.8 | 131.072K | $0.1 | $0.1 | — | |||
| Google: Gemma 3n 4Bgoogle/gemma-3n-e4b-it | 3.2 | 32.768K | $0.06 | $0.12 | — |