Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Small GPT-5 for responsive agents, coding help, and everyday automation
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 17.4 | 262.144K | $0.15 | $1.2 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 16.3 | 200K | $1.1 | $4.4 | — | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 16.3 | 200K | $0.55 | $2.2 | — | |||
| GPT-5 Miniopenai/gpt-5-mini | 15.6 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 15.6 | 400K | $0.125 | $1 | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 14.4 | 262.144K | $0.2 | $0.2 | — | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 14.4 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 14.4 | 256K | Free | Free | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 13.8 | 256K | Free | Free | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 11.1 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 11.1 | 1.04758M | $0.05 | $0.2 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 9.7 | 262.144K | $0.15 | $0.15 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 9.7 | 262.144K | $0.075 | $0.075 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 8.2 | 327.68K | $0.1 | $0.3 | — |