Strong GLM coding model for agentic engineering, terminals, and repository generation
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Open Gemma instruction model for efficient chat and self-hosted deployments
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Open Gemma instruction model for efficient chat and self-hosted deployments
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Reranking model for improving retrieval quality in search and recommendation systems
Nemotron model for efficient reasoning, coding, and specialized AI agents
Open MiniMax flagship for coding agents, office automation, and complex environments
Low-latency M2.7 variant for interactive coding plans and agent loops
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron multimodal model for visual reasoning and agentic AI workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Turkish-language multimodal instruct model built on Gemma 3 12B for e-commerce text, chat, and image-text tasks
Efficient Indian-language reasoning model for chat, coding, and multilingual work
Large open Qwen multimodal MoE for visual agents and long technical tasks
High-speed MiniMax model for low-latency coding and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
StepFun flash lane for quick multimodal reasoning and coding assistance
Lightly post-trained 398B MoE chat model for creative work, long-context prompts, and tool-using agents
High-accuracy OCR model for extracting text from documents, screenshots, receipts, and natural scenes
Safety model for policy screening, moderation, and risk-aware routing workflows
Budget GLM lane for fast coding help, routing, and everyday automation
Efficient GLM model for fast reasoning, coding, and agent workflows
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
MiMo flash model for fast multimodal assistance and agent workflows
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Compact multimodal coding model for repository exploration, file editing, and software agents
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
GLM vision model for visual reasoning, documents, and multimodal agents
Lightweight GLM vision model for visual reasoning, documents, and multimodal agents
Mistral coding agent model for repository tasks and software engineering workflows
Compact multimodal Mistral model for local assistants, edge agents, and efficient tool use
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
Reasoning-tuned 26B MoE model with 3B active parameters for agents, tools, and multi-step workloads
DeepSeek chat model for instruction following, coding, and analysis
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Experimental chat-tuned 6B MoE model with 1B active parameters for low-resource chat and instruction following
Thinking Kimi model for slower research passes, planning, and hard technical questions
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 256K | $0.1 | $0.3 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 E4B ITgoogle/gemma-4-E4B-it | 131.072K | $0.02 | $0.1 | 2026-04-02 | |||
| Gemma 4 E2B ITgoogle/gemma-4-E2B-it | 131.072K | $0.04 | $0.08 | 2026-04-02 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| Llama Nemotron Rerank VL 1B v2nvidia/llama-nemotron-rerank-vl-1b-v2 | 128K | — | — | 2026-03-31 | |||
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | — | — | 2026-03-24 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron VoiceChatnvidia/nemotron-voicechat | 128K | — | — | 2026-03-16 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 35B-A3Balibaba/qwen3.5-35b-a3b | 262.144K | $0.25 | $2 | 2026-02-23 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Trendyol Asure 12Btrendyol/asure-12b | 131.072K | $0.1 | $0.5 | 2026-02-19 | |||
| Sarvam 30Bsarvam/sarvam-30b | 128K | $0.02 | $0.1 | 2026-02-18 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| MiniMax-M2.5-highspeedminimax/MiniMax-M2.5-highspeed | 204.8K | $0.6 | $2.4 | 2026-02-13 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| Llama Nemotron Embed VL 1B v2nvidia/llama-nemotron-embed-vl-1b-v2 | 32.768K | — | — | 2026-02-10 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| Step 3.5 Flashstepfun/step-3.5-flash | 256K | $0.1 | $0.3 | 2026-01-29 | |||
| Trinity Large Previewarcee-ai/trinity-large-preview | 524.288K | — | — | 2026-01-27 | |||
| DeepSeek OCR 2deepseek/deepseek-ocr-2 | 8.192K | $0.03 | $0.03 | 2026-01-27 | |||
| Nemotron Content Safety Reasoning 4Bnvidia/nemotron-content-safety-reasoning-4b | 128K | — | — | 2026-01-22 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| GLM-4.7-FlashXzhipuai/glm-4.7-flashx | 200K | $0.07 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| MiMo-V2-Flashxiaomi/mimo-v2-flash | 262.144K | $0.14 | $0.28 | 2025-12-16 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Devstral Small 2mistral/devstral-small-2 | 262.144K | $0.1 | $0.3 | 2025-12-09 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| GLM-4.6Vzhipuai/glm-4.6v | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| GLM-4.6V-Flashzhipuai/glm-4.6v-flash | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| Devstral 2 (latest)mistral/devstral-medium-latest | 262.144K | $0.4 | $2 | 2025-12-02 | |||
| Ministral 14Bmistral/ministral-14b | 262.144K | $0.2 | $0.2 | 2025-12-02 | |||
| DeepSeek Reasonerdeepseek/deepseek-reasoner | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| Trinity Miniarcee-ai/trinity-mini | 131.072K | $0.045 | $0.15 | 2025-12-01 | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| Trinity Nano Previewarcee-ai/trinity-nano-preview | 131.072K | — | — | 2025-12-01 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 |