Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Strong small GPT for coding subagents, quick tool use, and high-volume work
Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Open MiMo model for multimodal coding agents and long-context automation
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Small GPT-5 for responsive agents, coding help, and everyday automation
Thinking Kimi model for slower research passes, planning, and hard technical questions
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
Google's proven reasoning model for coding, math, and multimodal analysis
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Low-latency Gemini model for high-volume multimodal and agent workloads
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Affordable GPT-4.1 lane for fast coding help and structured extraction
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 25.4 | 1M | $1.25 | $2.5 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 24.8 | 262.144K | Free | Free | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 24.8 | 131.072K | $0.06 | $0.18 | — | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 24.8 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 24.6 | 400K | $0.375 | $2.25 | — | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 24.6 | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| inclusionAI: Ling 3.0 Tiny (free)inclusionai/ling-3.0-tiny:free | 24.5 | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 24.3 | 256K | $0.312 | $1.25 | — | |||
| OpenAI: gpt-oss-120b (free)openai/gpt-oss-120b:free | 23.8 | 131.072K | Free | Free | — | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 23.4 | 1M | Free | Free | — | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 23.4 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 23.2 | 204.8K | $0.3 | $1.2 | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 22.7 | 1.04858M | $0.15 | $1.25 | — | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 22.7 | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 22.3 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 21.9 | 262.144K | $0.3 | $2 | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 21.8 | 262.144K | $0.1 | $0.15 | — | |||
| Qwen: Qwen3.5-9B (batch)qwen/qwen3.5-9b:batch | 21.8 | 262.144K | $0.17 | $0.25 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 21.2 | 1M | $1.5 | $7.5 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 21.2 | 400K | $0.1 | $0.625 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 21.2 | 1M | $3 | $15 | — | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 21.2 | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 20.2 | 256K | Free | Free | — | |||
| LongCat-2.0meituan/longcat-2.0 | 19.7 | 1M | $0.3 | $1.2 | 2026-06-30 | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 19.1 | 262.144K | $0.55 | $3.5 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 18.8 | 262.144K | $0.1 | $0.9 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 18.4 | 1M | $0.3 | $2.5 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 17.6 | 200K | $1 | $5 | — | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 17.6 | 200K | $0.5 | $2.5 | — | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 17.4 | 400K | $0.125 | $1 | — | |||
| GPT-5 Miniopenai/gpt-5-mini | 17.4 | 400K | $0.25 | $2 | 2025-08-07 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 17.2 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 16.9 | 262.144K | $0.15 | $1.2 | — | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 16.7 | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 16.7 | 1.04858M | $0.625 | $5 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 16.2 | 262.144K | $0.26 | $2.08 | — | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 16.0 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 15.7 | 200K | $0.55 | $2.2 | — | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 15.4 | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 15.4 | 262.144K | Free | Free | — | |||
| DeepSeek: DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | 15.4 | 131.072K | $0.27 | $1 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 15.4 | 262.144K | $0.39 | $0.97 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 15.2 | 131.072K | Free | Free | — | |||
| Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | 14.9 | 262.144K | $1.5 | $7.5 | — | |||
| Mistral: Mistral Medium 3.5 (batch)mistralai/mistral-medium-3-5:batch | 14.9 | 262.144K | $0.75 | $3.75 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 14.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 14.8 | 1.04758M | $0.2 | $0.8 | — | |||
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 14.7 | 131.072K | $0.2 | $1 | — | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 14.7 | 131.072K | $0.4 | $2 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 14.5 | 256K | Free | Free | — |