Open GPT reasoning model for self-hosted agents and controllable deployments
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Google's proven reasoning model for coding, math, and multimodal analysis
Fast Gemini workhorse for multimodal apps where latency and price matter
High-effort o3 tier for difficult technical reasoning and careful answers
Balanced Claude model for coding, analysis, agent workflows, and cost control
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
Sparse MoE Qwen model with 3B active parameters for efficient chat and reasoning
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Fast o-series model for compact reasoning, coding, and tool use
Dense open Qwen model for self-hosted chat, reasoning, and coding
Large open Qwen MoE for multilingual reasoning, coding, and tool use
O-series reasoning model for hard analysis, math, coding, and planning
Cohere command model for multilingual enterprise agents, tools, and chat
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
Sonar search model for autonomous research and citation-backed long-form reports
Classic open reasoning model for transparent math, coding, and deliberate problem solving
Smaller o-series reasoner for economical coding, math, and planning tasks
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
O-series reasoning model for hard analysis, math, coding, and planning
Compact open Llama model for lightweight chat, drafting, and self-hosting
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
Qwen instruction model for multilingual chat, reasoning, and tool use
Web-grounded Sonar for multi-step research questions that need cited reasoning
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.4 | $0.4 | 2025-07-25 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| o3-proopenai/o3-pro | 200K | $20 | $80 | 2025-06-10 | |||
| Claude Sonnet 4anthropic/claude-sonnet-4-20250514 | 200K | $3 | $15 | 2025-05-22 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 | $75 | 2025-05-22 | |||
| Palmyra X5writer/palmyra-x5 | 1M | $0.6 | $6 | 2025-04-28 | |||
| Qwen3 30B A3Balibaba/qwen3-30b-a3b | 131.072K | $0.08 | $0.29 | 2025-04-28 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3 235B-A22Balibaba/qwen3-235b-a22b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| o1-proopenai/o1-pro | 200K | $150 | $600 | 2025-03-19 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 | $10 | 2025-03-13 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Sonar Deep Researchperplexity/sonar-deep-research | 128K | $2 | $8 | 2025-02-01 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | 128K | $0.02 | $0.04 | 2024-07-23 | |||
| Mistral Nemomistral/mistral-nemo | 128K | $0.15 | $0.15 | 2024-07-01 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 |