Fast GLM vision model for screenshots, documents, and multimodal agent tasks
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
Nemotron model for efficient reasoning, coding, and specialized AI agents
Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks
Open MiniMax flagship for coding agents, office automation, and complex environments
MiMo omni model for text, image, video, audio, and agents
Low-latency M2.7 variant for interactive coding plans and agent loops
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Faster GLM-5 lane for coding agents that need lower latency
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Efficient Mistral model for fast chat, extraction, and production assistants
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Reasoning Grok for document-heavy analysis and long-horizon tool use
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across...
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party...
Claude workhorse for coding agents, careful analysis, and production cost control
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen multimodal MoE for visual agents and long technical tasks
Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,...
Flagship ByteDance Seed 2.0 model for complex multimodal reasoning and long-horizon agent workflows
Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across...
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...
High-speed MiniMax model for low-latency coding and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results...
High-end Claude for difficult coding, planning, and slower expert reasoning
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token....
Budget GLM lane for fast coding help, routing, and everyday automation
Efficient GLM model for fast reasoning, coding, and agent workflows
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...
ByteDance Seed model for multimodal reasoning, long-context analysis, and agent workflows
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
MiMo flash model for fast multimodal assistance and agent workflows
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,...
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...
GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GLM-5V-Turbozhipuai/glm-5v-turbo | 200K | $5 | $22 | 2026-04-01 | |||
| Arcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinking | 262.144K | $0.25 | $0.8 | 2026-04-01 | |||
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | — | — | 2026-03-24 | |||
| MiMo-V2-Proxiaomi/mimo-v2-pro | 1.04858M | $0.435 | $0.87 | 2026-03-18 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| MiMo-V2-Omnixiaomi/mimo-v2-omni | 262.144K | $0.14 | $0.28 | 2026-03-18 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nano | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| OpenAI: GPT-5.4 Miniopenai/gpt-5.4-mini | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| GLM-5-Turbozhipuai/glm-5-turbo | 200K | $0.9 | $3.7 | 2026-03-16 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| NVIDIA: Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | 262.144K | $0.085 | $0.4 | 2026-03-11 | |||
| Grok 4.20 (Reasoning)xai/grok-4.20-0309-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| OpenAI: GPT-5.4 Proopenai/gpt-5.4-pro | 1.05M | $30 | $180 | 2026-03-05 | |||
| OpenAI: GPT-5.4openai/gpt-5.4 | 1.05M | $2.5 | $15 | 2026-03-05 | |||
| Google: Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| Qwen3.5 Flashalibaba/qwen3.5-flash | 1M | $0.029 | $0.287 | 2026-02-23 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 35B-A3Balibaba/qwen3.5-35b-a3b | 262.144K | $0.25 | $2 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Google: Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Google: Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Claude Sonnet 4.6anthropic/claude-sonnet-4-6 | 1M | $3 | $15 | 2026-02-17 | |||
| Qwen3.5 Plusalibaba/qwen3.5-plus | 1M | $0.4 | $2.4 | 2026-02-16 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| ByteDance Seed: Seed-2.0-Minibytedance-seed/seed-2.0-mini | 262.144K | $0.1 | $0.4 | 2026-02-14 | |||
| Seed 2.0 Probytedance-seed/seed-2.0-pro | 256K | $0.475 | $2.375 | 2026-02-14 | |||
| ByteDance Seed: Seed-2.0-Litebytedance-seed/seed-2.0-lite | 262.144K | $0.25 | $2 | 2026-02-14 | |||
| ByteDance Seed: Seed-2.0-Codebytedance-seed/seed-2.0-code | 262.144K | $0.5 | $3 | 2026-02-14 | |||
| MiniMax-M2.5-highspeedminimax/MiniMax-M2.5-highspeed | 204.8K | $0.6 | $2.4 | 2026-02-13 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| OpenAI: GPT-5.3-Codexopenai/gpt-5.3-codex | 400K | $1.75 | $14 | 2026-02-05 | |||
| Claude Opus 4.6anthropic/claude-opus-4-6 | 1M | $5 | $25 | 2026-02-05 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| StepFun: Step 3.5 Flashstepfun/step-3.5-flash | 262.144K | $0.1 | $0.3 | 2026-01-29 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| GLM-4.7-FlashXzhipuai/glm-4.7-flashx | 200K | $0.07 | $0.4 | 2026-01-19 | |||
| MoonshotAI: Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.45 | $2.25 | 2026-01 | |||
| Seed 1.8bytedance-seed/seed-1-8 | 256K | $0.119 | $1.187 | 2025-12-28 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1.04858M | $0.5 | $3 | 2025-12-17 | |||
| MiMo-V2-Flashxiaomi/mimo-v2-flash | 262.144K | $0.14 | $0.28 | 2025-12-16 | |||
| NVIDIA: Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| OpenAI: GPT-5.2 Proopenai/gpt-5.2-pro | 400K | $21 | $168 | 2025-12-11 | |||
| OpenAI: GPT-5.2openai/gpt-5.2 | 400K | $1.75 | $14 | 2025-12-11 | |||
| OpenAI: GPT-5.2-Codexopenai/gpt-5.2-codex | 400K | $1.75 | $14 | 2025-12-11 |