Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Open MiMo model for multimodal coding agents and long-context automation
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Strong small GPT for coding subagents, quick tool use, and high-volume work
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Fast Gemini model balancing multimodal reasoning, tool use, and cost
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
O-series reasoning model for hard analysis, math, coding, and planning
The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Open Gemma instruction model for efficient chat and self-hosted deployments
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
Low-latency Gemini model for high-volume multimodal and agent workloads
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 59.1 | 262.144K | $0.25 | $1 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 58.6 | 1.04858M | Free | Free | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 58.6 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 58.6 | 524.288K | $0.3 | $1.2 | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 57.0 | 131.072K | $0.06 | $0.18 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 57.0 | 262.144K | Free | Free | — | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 56.8 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 56.1 | 400K | $0.1 | $0.625 | — | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 56.1 | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 56.1 | 400K | $0.375 | $2.25 | — | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 56.1 | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 55.9 | 1M | $0.32 | $1.28 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 54.5 | 1M | $0.325 | $1.95 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 53.7 | 262.144K | $0.3 | $2 | — | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 52.9 | 524.288K | $0.5 | $1.2 | — | |||
| Inkling Smallthinkingmachines/inkling-small | 52.9 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 52.9 | 1.04858M | Free | Free | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 52.1 | 1M | $1.5 | $7.5 | — | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 52.1 | 1.04858M | Free | Free | — | |||
| Inklingthinkingmachines/inkling | 52.1 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 52.1 | 524.288K | $1 | $4.05 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 52.1 | 1M | $3 | $15 | — | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 51.5 | 256K | $1 | $2 | — | |||
| GPT-5.1openai/gpt-5.1 | 49.4 | 400K | $1.25 | $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 49.4 | 400K | $0.625 | $5 | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 49.3 | 1.04858M | $0.15 | $1.25 | — | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 49.3 | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 48.2 | 262.144K | $0.55 | $3.5 | — | |||
| Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | 46.9 | 262.144K | $1.5 | $7.5 | — | |||
| Mistral: Mistral Medium 3.5 (batch)mistralai/mistral-medium-3-5:batch | 46.9 | 262.144K | $0.75 | $3.75 | — | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 46.8 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 45.7 | 262.144K | $0.26 | $2.08 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 43.9 | 200K | $1 | $5 | — | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 43.9 | 200K | $0.5 | $2.5 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 43.4 | 262.144K | $0.39 | $0.97 | — | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 43.4 | 262.144K | Free | Free | — | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 43.4 | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 42.2 | 1M | $1 | $2 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 42.2 | 1M | $1.25 | $2.5 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 41.9 | 262.144K | $0.1 | $0.9 | — | |||
| o1openai/o1 | 39.7 | 200K | $15 | $60 | 2024-12-05 | |||
| OpenAI: o1 (batch)openai/o1:batch | 39.7 | 200K | $7.5 | $30 | — | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 39.6 | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 39.3 | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 39.3 | 262.144K | Free | Free | — | |||
| GPT-5openai/gpt-5 | 37.8 | 400K | $1.25 | $10 | 2025-08-07 | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 37.8 | 400K | $0.625 | $5 | — | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 37.6 | 200K | $3 | $15 | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 37.0 | 256K | $0.312 | $1.25 | — | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 34.7 | 1.04858M | $0.25 | $1.5 | 2026-03-03 |