Strong GLM coding model for agentic engineering, terminals, and repository generation
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Open Gemma instruction model for efficient chat and self-hosted deployments
Open Gemma instruction model for efficient chat and self-hosted deployments
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Open Gemma instruction model for efficient chat and self-hosted deployments
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Reranking model for improving retrieval quality in search and recommendation systems
Nemotron model for efficient reasoning, coding, and specialized AI agents
Open MiniMax flagship for coding agents, office automation, and complex environments
Low-latency M2.7 variant for interactive coding plans and agent loops
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Efficient Mistral model for fast chat, extraction, and production assistants
Nemotron multimodal model for visual reasoning and agentic AI workflows
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Turkish-language multimodal instruct model built on Gemma 3 12B for e-commerce text, chat, and image-text tasks
Efficient Indian-language reasoning model for chat, coding, and multilingual work
Large open Qwen multimodal MoE for visual agents and long technical tasks
High-speed MiniMax model for low-latency coding and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
StepFun flash lane for quick multimodal reasoning and coding assistance
High-accuracy OCR model for extracting text from documents, screenshots, receipts, and natural scenes
Lightly post-trained 398B MoE chat model for creative work, long-context prompts, and tool-using agents
Safety model for policy screening, moderation, and risk-aware routing workflows
Efficient GLM model for fast reasoning, coding, and agent workflows
Budget GLM lane for fast coding help, routing, and everyday automation
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
MiMo flash model for fast multimodal assistance and agent workflows
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
GLM vision model for visual reasoning, documents, and multimodal agents
Lightweight GLM vision model for visual reasoning, documents, and multimodal agents
Mistral coding agent model for repository tasks and software engineering workflows
DeepSeek chat model for instruction following, coding, and analysis
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
Experimental chat-tuned 6B MoE model with 1B active parameters for low-resource chat and instruction following
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Reasoning-tuned 26B MoE model with 3B active parameters for agents, tools, and multi-step workloads
Thinking Kimi model for slower research passes, planning, and hard technical questions
Kimi reasoning model for long-horizon research, planning, and tool use
Safety model for policy screening, moderation, and risk-aware routing workflows
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Gemma 4 E2B ITgoogle/gemma-4-E2B-it | 131.072K | $0.04 | $0.08 | 2026-04-02 | |||
| Gemma 4 E4B ITgoogle/gemma-4-E4B-it | 131.072K | $0.02 | $0.1 | 2026-04-02 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 256K | $0.1 | $0.3 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| Llama Nemotron Rerank VL 1B v2nvidia/llama-nemotron-rerank-vl-1b-v2 | 128K | — | — | 2026-03-31 | |||
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | — | — | 2026-03-24 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron VoiceChatnvidia/nemotron-voicechat | 128K | — | — | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 35B-A3Balibaba/qwen3.5-35b-a3b | 262.144K | $0.25 | $2 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Trendyol Asure 12Btrendyol/asure-12b | 131.072K | $0.1 | $0.5 | 2026-02-19 | |||
| Sarvam 30Bsarvam/sarvam-30b | 128K | $0.02 | $0.1 | 2026-02-18 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| MiniMax-M2.5-highspeedminimax/MiniMax-M2.5-highspeed | 204.8K | $0.6 | $2.4 | 2026-02-13 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| Llama Nemotron Embed VL 1B v2nvidia/llama-nemotron-embed-vl-1b-v2 | 32.768K | — | — | 2026-02-10 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| Step 3.5 Flashstepfun/step-3.5-flash | 256K | $0.1 | $0.3 | 2026-01-29 | |||
| DeepSeek OCR 2deepseek/deepseek-ocr-2 | 8.192K | $0.03 | $0.03 | 2026-01-27 | |||
| Trinity Large Previewarcee-ai/trinity-large-preview | 524.288K | — | — | 2026-01-27 | |||
| Nemotron Content Safety Reasoning 4Bnvidia/nemotron-content-safety-reasoning-4b | 128K | — | — | 2026-01-22 | |||
| GLM-4.7-FlashXzhipuai/glm-4.7-flashx | 200K | $0.07 | $0.4 | 2026-01-19 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| MiMo-V2-Flashxiaomi/mimo-v2-flash | 262.144K | $0.14 | $0.28 | 2025-12-16 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| GLM-4.6Vzhipuai/glm-4.6v | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| GLM-4.6V-Flashzhipuai/glm-4.6v-flash | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| Devstral 2 (latest)mistral/devstral-medium-latest | 262.144K | $0.4 | $2 | 2025-12-02 | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| DeepSeek Reasonerdeepseek/deepseek-reasoner | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| Trinity Nano Previewarcee-ai/trinity-nano-preview | 131.072K | — | — | 2025-12-01 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| Trinity Miniarcee-ai/trinity-mini | 131.072K | $0.045 | $0.15 | 2025-12-01 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | 262.144K | $1.15 | $8 | 2025-11-06 | |||
| GPT OSS Safeguard 20Bopenai/gpt-oss-safeguard-20b | 131.072K | $0.07 | $0.2 | 2025-10-29 |