MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Strong small GPT for coding subagents, quick tool use, and high-volume work
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 31.0 | 1.04858M | Free | Free | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 30.8 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 30.8 | 524.288K | $0.3 | $1.2 | — | |||
| MoonshotAI: Kimi K2.7 Code (batch)moonshotai/kimi-k2.7-code:batch | 30.3 | 262.144K | $0.95 | $4 | — | |||
| Gemini 3.6 Flashgoogle/gemini-3.6-flash | 30.2 | 1.04858M | $0.75 | $3.75 | 2026-07-21 | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 30.2 | 1.04858M | $0.375 | $1.875 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 30.0 | 262.144K | Free | Free | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 30.0 | 131.072K | $0.06 | $0.18 | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 29.6 | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 29.0 | 1M | $0.325 | $1.95 | — | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 28.9 | 256K | $1 | $2 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 27.5 | 512.288K | $0.6 | $3.6 | — | |||
| Muse Spark 1.1meta/muse-spark-1.1 | 27.5 | 1.04858M | $1.25 | $4.25 | 2026-04-08 | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 27.3 | 1.04858M | $0.75 | $4.5 | — | |||
| Gemini 3.5 Flashgoogle/gemini-3.5-flash | 27.3 | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| GPT-5openai/gpt-5 | 26.5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 26.5 | 400K | $0.625 | $5 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 26.2 | 202.752K | $0.4 | $1.75 | — | |||
| MiniMax: MiniMax M2.7 (free)minimax/minimax-m2.7:free | 25.9 | 196.608K | Free | Free | — | |||
| Z.ai: GLM 5.1z-ai/glm-5.1 | 25.2 | 200K | $0.966 | $3.036 | — | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 25.0 | 524.288K | $0.5 | $1.2 | — | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 25.0 | 1.04858M | Free | Free | — | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 24.3 | 524.288K | $1 | $4.05 | — | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 24.3 | 1.04858M | Free | Free | — | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 23.9 | 1M | $1.475 | $4.425 | — | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 21.7 | 1M | Free | Free | — | |||
| GPT-5.1openai/gpt-5.1 | 21.6 | 400K | $1.25 | $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 21.6 | 400K | $0.625 | $5 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 21.0 | 262.144K | $0.021 | $0.063 | — | |||
| inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 20.1 | 262.144K | $0.075 | $0.625 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 20.1 | 262.144K | $0.3 | $2 | — | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 19.7 | 1M | $0.32 | $1.28 | — | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 19.7 | 400K | $0.375 | $2.25 | — | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 19.7 | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 18.6 | 198K | $0.43 | $1.75 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 17.7 | 400K | $0.1 | $0.625 | — | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 17.7 | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 17.6 | 200K | $3 | $15 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 17.5 | 1M | $1.5 | $7.5 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 17.5 | 1M | $3 | $15 | — | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 17.2 | 1M | $1 | $2 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 17.2 | 1M | $1.25 | $2.5 | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 16.8 | 204.8K | $0.3 | $1.2 | — | |||
| inclusionAI: Ling 3.0 Tiny (free)inclusionai/ling-3.0-tiny:free | 16.0 | 262.144K | Free | Free | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 15.9 | 1.04858M | $0.15 | $1.25 | — | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 15.9 | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| LongCat-2.0meituan/longcat-2.0 | 15.9 | 1M | $0.3 | $1.2 | 2026-06-30 | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 15.1 | 262.144K | $0.3 | $1.2 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 15.0 | 262.144K | $0.1 | $0.9 | — | |||
| OpenAI: gpt-oss-120b (free)openai/gpt-oss-120b:free | 13.2 | 131.072K | Free | Free | — |