Safety model for policy screening, moderation, and risk-aware routing workflows
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Nemotron multimodal model for visual reasoning and agentic AI workflows
Efficient open MiniMax model built for coding agents and tool-heavy workflows
ByteDance Seed model for long-context reasoning, instruction following, and tool-assisted tasks
Fast Claude model for responsive assistance, classification, and lightweight agents
Fast Claude lane for lightweight agents, office tasks, and responsive chat
Specialized Gemini 2.5 model for browser-control agents that automate UI tasks
Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning
Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
Qwen vision-language thinking model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open multimodal reasoning model for transparent analysis of text and images
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Fully open 70B multilingual LLM supporting 1800+ languages with 65K context. Trained on 15T tokens of compliant open data. Apache 2.0, EU AI Act compliant.
Flagship Indian-language reasoning model for enterprise multilingual applications
Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen thinking model for local reasoning, math, and coding agents
Low-latency ByteDance Seed model for high-throughput chat, extraction, and lightweight tool use
Translation model for multilingual conversion, localization, and cross-language workflows
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes
Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows
Compact Nemotron model for efficient reasoning and deployable AI agents
ByteDance Seed multimodal model for image understanding, visual reasoning, and tool-assisted tasks
GLM vision model for visual reasoning, documents, and multimodal agents
Small GPT-5 for responsive agents, coding help, and everyday automation
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Qwen coding model for software agents, repository edits, and code reasoning
Efficient GLM model for fast reasoning, coding, and agent workflows
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Hosted Qwen coder for software agents, repo edits, and long-context code
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Instruct model with native audio input for speech understanding and tool use
Mistral coding agent model for repository tasks and software engineering workflows
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT OSS Safeguard 120Bopenai/gpt-oss-safeguard-120b | 131.072K | $0.15 | $0.6 | 2025-10-29 | |||
| Nemotron Nano 12B v2 VLnvidia/nemotron-nano-12b-v2-vl | 128K | $0.2 | $0.6 | 2025-10-28 | |||
| MiniMax-M2minimax/MiniMax-M2 | 204.8K | $0.3 | $1.2 | 2025-10-27 | |||
| Seed 1.6bytedance-seed/seed-1-6 | 256K | $0.119 | $1.187 | 2025-10-15 | |||
| Claude Haiku 4.5anthropic/claude-haiku-4-5-20251001 | 200K | $1 | $5 | 2025-10-15 | |||
| Claude Haiku 4.5 (latest)anthropic/claude-haiku-4-5 | 200K | $1 | $5 | 2025-10-15 | |||
| Gemini 2.5 Computer Use Previewgoogle/gemini-2.5-computer-use-preview-10-2025 | 128K | $1.25 | $10 | 2025-10-07 | |||
| GPT-5 Proopenai/gpt-5-pro | 400K | $15 | $120 | 2025-10-06 | |||
| Granite-4.0-H-Smallibm/granite-4-h-small | 131.072K | $0.064 | $0.265 | 2025-10-02 | |||
| GLM-4.6zhipuai/glm-4.6 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Claude Sonnet 4.5anthropic/claude-sonnet-4-5-20250929 | 200K | $3 | $15 | 2025-09-29 | |||
| Claude Sonnet 4.5 (latest)anthropic/claude-sonnet-4-5 | 200K | $3 | $15 | 2025-09-29 | |||
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 | |||
| Qwen3 Maxalibaba/qwen3-max | 262.144K | $1.2 | $6 | 2025-09-23 | |||
| Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | 2025-09-23 | |||
| Qwen3-VL Plusalibaba/qwen3-vl-plus | 262.144K | $0.2 | $1.6 | 2025-09-23 | |||
| Magistral Small 1.2mistral/magistral-small-2509 | 131.072K | $0.5 | $1.5 | 2025-09-18 | |||
| GPT-5-Codexopenai/gpt-5-codex | 400K | $1.1 | $9 | 2025-09-15 | |||
| Apertus 70Bswiss-ai/apertus-70b | 65.536K | $0.46 | $2.42 | 2025-09-02 | |||
| Sarvam 105Bsarvam/sarvam-105b | 131.072K | $0.04 | $0.16 | 2025-09-01 | |||
| Qwen3-Next 80B-A3B Instructalibaba/qwen3-next-80b-a3b-instruct | 131.072K | $0.5 | $2 | 2025-09 | |||
| Qwen3-Next 80B-A3B (Thinking)alibaba/qwen3-next-80b-a3b-thinking | 131.072K | $0.5 | $6 | 2025-09 | |||
| Seed 1.6 Flashbytedance-seed/seed-1-6-flash | 256K | $0.022 | $0.223 | 2025-08-28 | |||
| Command A Translatecohere/command-a-translate-08-2025 | 8K | $2.5 | $10 | 2025-08-28 | |||
| DeepSeek-V3.1deepseek/deepseek-v3.1 | 131.072K | $0.19 | $0.71 | 2025-08-21 | |||
| Command A Reasoningcohere/command-a-reasoning-08-2025 | 256K | $2.5 | $10 | 2025-08-21 | |||
| Nemotron Nano 9B v2nvidia/nemotron-nano-9b-v2 | 131.072K | $0.06 | $0.23 | 2025-08-18 | |||
| Seed 1.6 Visionbytedance-seed/seed-1-6-vision | 256K | $0.119 | $1.187 | 2025-08-15 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Claude Opus 4.1 (latest)anthropic/claude-opus-4-1 | 200K | $15 | $75 | 2025-08-05 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| GLM-4.5-Flashzhipuai/glm-4.5-flash | 131.072K | — | — | 2025-07-28 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.4 | $0.4 | 2025-07-25 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Voxtral Small 24B 2507mistral/voxtral-small-24b-2507 | 32.768K | $0.1 | $0.3 | 2025-07-15 | |||
| Voxtral Mini 3B 2507mistral/voxtral-mini-3b-2507 | 32.768K | $0.04 | $0.04 | 2025-07-15 | |||
| Voxtral Small (latest)mistral/voxtral-small-latest | 32K | $0.1 | $0.3 | 2025-07-15 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Devstral Mediummistral/devstral-medium-2507 | 128K | $0.4 | $2 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 |