Translation model for multilingual conversion, localization, and cross-language workflows
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Low-latency ByteDance Seed model for high-throughput chat, extraction, and lightweight tool use
Nano Banana image model for fast generation, edits, and character-consistent assets
Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes
Compact Nemotron model for efficient reasoning and deployable AI agents
ByteDance Seed multimodal model for image understanding, visual reasoning, and tool-assisted tasks
GLM vision model for visual reasoning, documents, and multimodal agents
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Small GPT-5 for responsive agents, coding help, and everyday automation
Open GPT reasoning model for self-hosted agents and controllable deployments
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Cohere vision model for multilingual document analysis, OCR, and image understanding
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Efficient GLM model for fast reasoning, coding, and agent workflows
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen coding model for software agents, repository edits, and code reasoning
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Nemotron model for efficient reasoning, coding, and specialized AI agents
Hosted Qwen coder for software agents, repo edits, and long-context code
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Instruct model with native audio input for speech understanding and tool use
Mistral coding agent model for repository tasks and software engineering workflows
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Gemini workhorse for multimodal apps where latency and price matter
Google's proven reasoning model for coding, math, and multimodal analysis
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
High-effort o3 tier for difficult technical reasoning and careful answers
Open Mistral reasoning model for transparent step-by-step problem solving
Balanced Claude model for coding, analysis, agent workflows, and cost control
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Flagship model for demanding analysis, coding, and production agent workflows
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Multimodal model for complex analysis, long-context understanding, tool use, and model distillation
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
Sparse MoE Qwen model with 3B active parameters for efficient chat and reasoning
OpenAI image model for production generation, edits, and brand-safe visual workflows
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Fast o-series model for compact reasoning, coding, and tool use
Nemotron model for efficient reasoning, coding, and specialized AI agents
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Command A Translatecohere/command-a-translate-08-2025 | 8K | $2.5 | $10 | 2025-08-28 | |||
| Seed 1.6 Flashbytedance-seed/seed-1-6-flash | 256K | $0.022 | $0.223 | 2025-08-28 | |||
| Nano Bananagoogle/gemini-2.5-flash-image | 32.768K | $0.3 | $30 | 2025-08-26 | |||
| Command A Reasoningcohere/command-a-reasoning-08-2025 | 256K | $2.5 | $10 | 2025-08-21 | |||
| DeepSeek-V3.1deepseek/deepseek-v3.1 | 131.072K | $0.19 | $0.71 | 2025-08-21 | |||
| Nemotron Nano 9B v2nvidia/nemotron-nano-9b-v2 | 131.072K | $0.06 | $0.23 | 2025-08-18 | |||
| Seed 1.6 Visionbytedance-seed/seed-1-6-vision | 256K | $0.119 | $1.187 | 2025-08-15 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| GPT-5 Chat (latest)openai/gpt-5-chat-latest | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| Claude Opus 4.1 (latest)anthropic/claude-opus-4-1 | 200K | $15 | $75 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| Command A Visioncohere/command-a-vision-07-2025 | 128K | $2.5 | $10 | 2025-07-31 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| GLM-4.5-Flashzhipuai/glm-4.5-flash | 131.072K | — | — | 2025-07-28 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.4 | $0.4 | 2025-07-25 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Voxtral Small (latest)mistral/voxtral-small-latest | 32K | $0.1 | $0.3 | 2025-07-15 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Devstral Mediummistral/devstral-medium-2507 | 128K | $0.4 | $2 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| Mistral Nemotronnvidia/mistral-nemotron | 128K | — | — | 2025-06-11 | |||
| o3-proopenai/o3-pro | 200K | $20 | $80 | 2025-06-10 | |||
| Magistral Smallmistral/magistral-small-2506 | 131.072K | $0.5 | $1.5 | 2025-06-10 | |||
| Claude Sonnet 4anthropic/claude-sonnet-4-20250514 | 200K | $3 | $15 | 2025-05-22 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 | $75 | 2025-05-22 | |||
| Claude Opus 4 (latest)anthropic/claude-opus-4-0 | 200K | — | — | 2025-05-22 | |||
| Claude Sonnet 4 (latest)anthropic/claude-sonnet-4-0 | 200K | $2.898 | $14.493 | 2025-05-22 | |||
| Solar Pro 2upstage/solar-pro2 | 65.536K | $0.15 | $0.6 | 2025-05-20 | |||
| Gemini Embedding 001google/gemini-embedding-001 | 2.048K | $0.15 | — | 2025-05-20 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 131.072K | $0.4 | $2 | 2025-05-07 | |||
| Nova Premieramazon/nova-premier | 1M | — | — | 2025-04-30 | |||
| Palmyra X5writer/palmyra-x5 | 1.04M | $0.6 | $6 | 2025-04-28 | |||
| Palmyra X4writer/palmyra-x4 | 122.88K | $2.5 | $10 | 2025-04-28 | |||
| Qwen3 30B A3Balibaba/qwen3-30b-a3b | 131.072K | $0.08 | $0.29 | 2025-04-28 | |||
| GPT-Image-1openai/gpt-image-1 | Not documented | $5 | $40 | 2025-04-24 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| Llama 3.1 Nemotron 70B Instructnvidia/llama-3.1-nemotron-70b-instruct | 128K | — | — | 2025-04-15 |