Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Classic open reasoning model for transparent math, coding, and deliberate problem solving
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
Compact GPT model for low-latency assistance and high-volume workloads
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
Thinking Kimi model for slower research passes, planning, and hard technical questions
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Open GPT reasoning model for self-hosted agents and controllable deployments
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Affordable GPT-4.1 lane for fast coding help and structured extraction
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
Small GPT-5 for responsive agents, coding help, and everyday automation
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Small omni GPT for cheap multimodal assistance and production-scale traffic
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Compact GPT model for low-latency assistance and high-volume workloads
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window and is designed for enterprise tasks...
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| inclusionAI: Ling 3.0 Tiny (free)inclusionai/ling-3.0-tiny:free | 26.5 | 262.144K | Free | Free | — | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 25.8 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 25.3 | 262.144K | $0.01 | $0.03 | — | |||
| DeepSeek-R1deepseek/deepseek-r1 | 24.6 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 24.2 | 128K | $5 | $15 | 2024-05-13 | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 23.0 | 1M | $0.3 | $2.5 | — | |||
| IBM: Granite 4.2 8Bibm-granite/granite-4.2-8b | 22.4 | 131.072K | $0.06 | $0.25 | — | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 22.1 | 131.072K | $0.23 | $2.3 | — | |||
| OpenAI: GPT-4 Turbo (batch)openai/gpt-4-turbo:batch | 21.5 | 128K | $5 | $15 | — | |||
| GPT-4 Turboopenai/gpt-4-turbo | 21.5 | 128K | $10 | $30 | 2023-11-06 | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 21.2 | 163.84K | $0.25 | $1 | — | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 21.0 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 20.7 | 131.072K | $0.05 | $0.2 | — | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 20.7 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 20.7 | 131.072K | Free | Free | — | |||
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 20.5 | 131.072K | $0.2 | $1 | — | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 20.5 | 131.072K | $0.4 | $2 | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 20.2 | 1.04758M | $0.2 | $0.8 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 20.2 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 20.1 | 262.144K | $0.25 | $0.75 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 20.1 | 262.144K | $0.5 | $1.5 | — | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 17.4 | 262.144K | $0.15 | $1.2 | — | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 16.3 | 200K | $0.55 | $2.2 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 16.3 | 200K | $1.1 | $4.4 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 16.3 | 128K | $0.2 | $0.696 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 16.2 | 131.072K | $0.15 | $0.6 | — | |||
| GPT-5 Miniopenai/gpt-5-mini | 15.6 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 15.6 | 400K | $0.125 | $1 | — | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 15.3 | 40.96K | $0.08 | $0.28 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 14.4 | 256K | Free | Free | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 14.4 | 262.144K | $0.2 | $0.2 | — | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 14.4 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 13.8 | 131.072K | $0.227 | $0.91 | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 13.8 | 256K | Free | Free | — | |||
| GPT-4openai/gpt-4 | 13.1 | 8.192K | $30 | $60 | 2023-11-06 | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 12.1 | 81.92K | $0.2 | $2.4 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 11.9 | 131.072K | $0.1 | $0.32 | — | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 11.9 | 65.536K | Free | Free | — | |||
| GPT-4o miniopenai/gpt-4o-mini | 11.4 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 11.4 | 128K | $0.075 | $0.3 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 11.1 | 1.04758M | $0.05 | $0.2 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 11.1 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| GPT-3.5-turboopenai/gpt-3.5-turbo | 10.7 | 16.385K | $0.5 | $1.5 | 2023-03-01 | |||
| OpenAI: GPT-3.5 Turbo (batch)openai/gpt-3.5-turbo:batch | 10.7 | 16.385K | $0.25 | $0.75 | — | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 10.1 | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 9.7 | 262.144K | $0.075 | $0.075 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 9.7 | 262.144K | $0.15 | $0.15 | — | |||
| IBM: Granite 4.1 8Bibm-granite/granite-4.1-8b | 9.5 | 131.072K | $0.05 | $0.1 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 9.0 | 131.072K | $0.117 | $0.455 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 8.2 | 327.68K | $0.1 | $0.3 | — |