GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Strong small GPT for coding subagents, quick tool use, and high-volume work
Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Open MiMo model for multimodal coding agents and long-context automation
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Small GPT-5 for responsive agents, coding help, and everyday automation
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Thinking Kimi model for slower research passes, planning, and hard technical questions
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
Google's proven reasoning model for coding, math, and multimodal analysis
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Low-latency Gemini model for high-volume multimodal and agent workloads
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Affordable GPT-4.1 lane for fast coding help and structured extraction
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 24.6 | 400K | $0.375 | $2.25 | — | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 24.6 | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| inclusionAI: Ling 3.0 Tiny (free)inclusionai/ling-3.0-tiny:free | 24.5 | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 24.3 | 256K | $0.312 | $1.25 | — | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 23.4 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 23.4 | 1M | Free | Free | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 23.2 | 204.8K | $0.3 | $1.2 | — | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 22.7 | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 22.7 | 1.04858M | $0.15 | $1.25 | — | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 22.3 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 21.9 | 262.144K | $0.3 | $2 | — | |||
| Qwen: Qwen3.5-9B (batch)qwen/qwen3.5-9b:batch | 21.8 | 262.144K | $0.17 | $0.25 | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 21.8 | 262.144K | $0.1 | $0.15 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 21.2 | 1M | $1.5 | $7.5 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 21.2 | 400K | $0.1 | $0.625 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 21.2 | 1M | $3 | $15 | — | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 21.2 | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 20.2 | 256K | Free | Free | — | |||
| LongCat-2.0meituan/longcat-2.0 | 19.7 | 1M | $0.3 | $1.2 | 2026-06-30 | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 19.1 | 262.144K | $0.55 | $3.5 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 18.8 | 262.144K | $0.1 | $0.9 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 18.4 | 1M | $0.3 | $2.5 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 17.6 | 200K | $1 | $5 | — | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 17.6 | 200K | $0.5 | $2.5 | — | |||
| GPT-5 Miniopenai/gpt-5-mini | 17.4 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 17.4 | 400K | $0.125 | $1 | — | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 17.2 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 16.9 | 262.144K | $0.15 | $1.2 | — | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 16.7 | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 16.7 | 1.04858M | $0.625 | $5 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 16.2 | 262.144K | $0.26 | $2.08 | — | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 16.0 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 15.7 | 200K | $0.55 | $2.2 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 15.4 | 262.144K | $0.39 | $0.97 | — | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 15.4 | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 15.4 | 262.144K | Free | Free | — | |||
| Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | 14.9 | 262.144K | $1.5 | $7.5 | — | |||
| Mistral: Mistral Medium 3.5 (batch)mistralai/mistral-medium-3-5:batch | 14.9 | 262.144K | $0.75 | $3.75 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 14.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 14.8 | 1.04758M | $0.2 | $0.8 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 14.5 | 256K | Free | Free | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 14.2 | 262.144K | $0.01 | $0.03 | — | |||
| Cohere: Command Acohere/command-a | 13.9 | 256K | $2.5 | $10 | — | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 13.6 | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | 13.6 | 1M | Free | Free | — | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 13.6 | 262.144K | Free | Free | — | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 13.6 | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 11.5 | 262.144K | $0.15 | $0.6 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 11.5 | 262.144K | $0.075 | $0.3 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 11.0 | 200K | $1.1 | $4.4 | — |