MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results...
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 1186.0 | 1.04858M | Free | Free | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 1185.0 | 202.752K | $0.4 | $1.75 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 1184.0 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 1184.0 | 524.288K | $0.3 | $1.2 | — | |||
| OpenAI: GPT-5.2 (batch)openai/gpt-5.2:batch | 1172.0 | 400K | $0.875 | $7 | — | |||
| OpenAI: GPT-5.2openai/gpt-5.2 | 1172.0 | 400K | $1.75 | $14 | 2025-12-11 | |||
| OpenAI: GPT-5.3-Codexopenai/gpt-5.3-codex | 1171.0 | 400K | $1.75 | $14 | 2026-02-05 | |||
| Z.ai: GLM 5 Turboz-ai/glm-5-turbo | 1171.0 | 202.752K | $1.2 | $4 | — | |||
| StepFun: Step 3.7 Flashstepfun/step-3.7-flash | 1170.0 | 256K | $0.2 | $1.15 | 2026-05-29 | |||
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | 1170.0 | 1.024M | $0.85 | $1.7 | 2026-04-24 | |||
| Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1170.0 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Xiaomi: MiMo-V2.5xiaomi/mimo-v2.5 | 1164.0 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| OpenAI: GPT-5openai/gpt-5 | 1161.0 | 400K | $1.25 | $10 | 2025-08-07 | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 1161.0 | 400K | $0.625 | $5 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 1160.0 | 200K | $1 | $5 | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 1160.0 | 204.8K | $0.3 | $1.2 | — | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 1160.0 | 200K | $0.5 | $2.5 | — | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 1157.0 | 1M | $1 | $2 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 1157.0 | 1M | $1.25 | $2.5 | — | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 1151.0 | 1M | $0.32 | $1.28 | — | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 1148.0 | 262.144K | $0.78 | $3.9 | — | |||
| Z.ai: GLM 5.1z-ai/glm-5.1 | 1146.0 | 200K | $0.966 | $3.036 | — | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 1142.0 | 400K | $0.125 | $1 | — | |||
| OpenAI: GPT-5 Miniopenai/gpt-5-mini | 1142.0 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5.1openai/gpt-5.1 | 1138.0 | 400K | $1.25 | $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 1138.0 | 400K | $0.625 | $5 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1134.0 | 1M | $0.325 | $1.95 | — | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 1131.0 | 202.752K | $1.2 | $4 | — | |||
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 1131.0 | 1.024M | $0.068 | $0.135 | 2026-04-24 | |||
| Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 1126.0 | 262.144K | $0.25 | $1 | — | |||
| OpenAI: GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | 1125.0 | 400K | $0.25 | $2 | 2025-11-13 | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 1119.0 | 1.04858M | Free | Free | — | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 1119.0 | 524.288K | $1 | $4.05 | — | |||
| Thinking Machines: Inklingthinkingmachines/inkling | 1119.0 | 1.04858M | $1 | $4.05 | 2026-07-15 | |||
| DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek/deepseek-v4-flash-0731:batch | 1116.0 | 1.04858M | $0.11 | $0.33 | — | |||
| DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1116.0 | 1.04858M | $0.065 | $0.18 | 2026-07-31 | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1114.0 | 1M | $0.26 | $1.56 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 1103.0 | 512.288K | $0.6 | $3.6 | — | |||
| NVIDIA: Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | 1101.0 | 256K | $0.625 | $3.125 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1101.0 | 1M | Free | Free | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 1089.0 | 262.144K | $0.25 | $0.75 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 1089.0 | 262.144K | $0.5 | $1.5 | — | |||
| Arcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1061.0 | 262.144K | $0.25 | $0.8 | 2026-04-01 |