DeepSeek chat model for instruction following, coding, and analysis
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Kimi reasoning model for long-horizon research, planning, and tool use
Thinking Kimi model for slower research passes, planning, and hard technical questions
Safety model for policy screening, moderation, and risk-aware routing workflows
Safety model for policy screening, moderation, and risk-aware routing workflows
Safety model for policy screening, moderation, and risk-aware routing workflows
Nemotron multimodal model for visual reasoning and agentic AI workflows
Efficient open MiniMax model built for coding agents and tool-heavy workflows
Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads
Compact open-weight hybrid Granite model for lightweight enterprise chat and tool calling
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Qwen vision-language thinking model for visual reasoning, documents, and agent tasks
Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
Gemma 3 27B tuned by AI Singapore for Southeast Asian languages and instruction following
Open multimodal reasoning model for transparent analysis of text and images
Flagship Indian-language reasoning model for enterprise multilingual applications
Efficient Qwen thinking model for local reasoning, math, and coding agents
Qwen instruction model for multilingual chat, reasoning, and tool use
Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes
Compact Nemotron model for efficient reasoning and deployable AI agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Cohere vision model for multilingual document analysis, OCR, and image understanding
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Open Mistral reasoning model for transparent step-by-step problem solving
Sparse MoE Qwen model with 3B active parameters for efficient chat and reasoning
Nemotron model for efficient reasoning, coding, and specialized AI agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Flagship Nemotron model for high-throughput reasoning and complex agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Llama with long-context vision for efficient multimodal agents
Dense open Qwen model for self-hosted chat, reasoning, and coding
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Smaller Qwen coder for efficient local agents and repo-level fixes
Large open Qwen MoE for multilingual reasoning, coding, and tool use
March 2025 checkpoint of DeepSeek-V3 with improved reasoning and coding
Efficient multimodal model for instruction following, coding, reasoning, and function calling
Cohere command model for multilingual enterprise agents, tools, and chat
Open multimodal Gemma instruction model for efficient text generation and image understanding
Open multimodal Gemma instruction model for multilingual text generation and image understanding
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
Open reasoning model from the Qwen team for math, coding, and step-by-step problem solving
Open Command R model optimized for Arabic enterprise chat, RAG, and cultural knowledge
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek Chatdeepseek/deepseek-chat | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | 262.144K | $1.15 | $8 | 2025-11-06 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| GPT OSS Safeguard 20Bopenai/gpt-oss-safeguard-20b | 131.072K | $0.07 | $0.2 | 2025-10-29 | |||
| GPT OSS Safeguard 120Bopenai/gpt-oss-safeguard-120b | 131.072K | $0.15 | $0.6 | 2025-10-29 | |||
| Llama 3.1 Nemotron Safety Guard 8B v3nvidia/llama-3.1-nemotron-safety-guard-8b-v3 | 128K | — | — | 2025-10-28 | |||
| Nemotron Nano 12B v2 VLnvidia/nemotron-nano-12b-v2-vl | 128K | $0.2 | $0.6 | 2025-10-28 | |||
| MiniMax-M2minimax/MiniMax-M2 | 204.8K | $0.3 | $1.2 | 2025-10-27 | |||
| Granite-4.0-H-Smallibm/granite-4-h-small | 131.072K | $0.064 | $0.265 | 2025-10-02 | |||
| Granite-4.0-H-Microibm/granite-4-h-micro | 131.072K | — | — | 2025-10-02 | |||
| GLM-4.6zhipuai/glm-4.6 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | 2025-09-23 | |||
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 | |||
| Gemma-SEA-LION-v4-27B-ITaisingapore/gemma-sea-lion-v4-27b-it | 128K | — | — | 2025-09-23 | |||
| Magistral Small 1.2mistral/magistral-small-2509 | 131.072K | $0.5 | $1.5 | 2025-09-18 | |||
| Sarvam 105Bsarvam/sarvam-105b | 131.072K | $0.04 | $0.16 | 2025-09-01 | |||
| Qwen3-Next 80B-A3B (Thinking)alibaba/qwen3-next-80b-a3b-thinking | 131.072K | $0.5 | $6 | 2025-09 | |||
| Qwen3-Next 80B-A3B Instructalibaba/qwen3-next-80b-a3b-instruct | 131.072K | $0.5 | $2 | 2025-09 | |||
| Command A Reasoningcohere/command-a-reasoning-08-2025 | 256K | $2.5 | $10 | 2025-08-21 | |||
| DeepSeek-V3.1deepseek/deepseek-v3.1 | 131.072K | $0.19 | $0.71 | 2025-08-21 | |||
| Nemotron Nano 9B v2nvidia/nemotron-nano-9b-v2 | 131.072K | $0.06 | $0.23 | 2025-08-18 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Command A Visioncohere/command-a-vision-07-2025 | 128K | $2.5 | $10 | 2025-07-31 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.4 | $0.4 | 2025-07-25 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 | |||
| Mistral Nemotronnvidia/mistral-nemotron | 128K | — | — | 2025-06-11 | |||
| Magistral Smallmistral/magistral-small-2506 | 131.072K | $0.5 | $1.5 | 2025-06-10 | |||
| Qwen3 30B A3Balibaba/qwen3-30b-a3b | 131.072K | $0.08 | $0.29 | 2025-04-28 | |||
| Llama 3.1 Nemotron 70B Instructnvidia/llama-3.1-nemotron-70b-instruct | 128K | — | — | 2025-04-15 | |||
| Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | 131.072K | — | — | 2025-04-07 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 10M | $0.17 | $0.66 | 2025-04-05 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| Qwen3 235B-A22Balibaba/qwen3-235b-a22b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| DeepSeek V3 0324deepseek/deepseek-v3-0324 | 163.84K | $0.2 | $0.8 | 2025-03-24 | |||
| Mistral Small 3.1 24Bmistral/mistral-small-3-1-24b-instruct-2503 | 128K | $0.106 | $0.318 | 2025-03-17 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 | $10 | 2025-03-13 | |||
| Gemma 3 4B ITgoogle/gemma-3-4b-it | 131.072K | $0.04 | $0.08 | 2025-03-12 | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 131.072K | $0.05 | $0.1 | 2025-03-12 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| QwQ 32Balibaba/qwq-32b | 131.072K | $0.287 | $0.861 | 2025-03-05 | |||
| Command R7B Arabiccohere/command-r7b-arabic-02-2025 | 128K | $0.037 | $0.15 | 2025-02-27 |