DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Claude model for demanding reasoning and long-horizon agentic work
Flagship GLM model for long-horizon coding, agents, and complex project delivery
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning
Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
xAI's frontier model for long-running agents, coding, knowledge work, and visual projects
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...
2.4-trillion-parameter MoE flagship for coding, professional work, multimodal understanding, and long-horizon agentic workflows
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
xAI's Grok model for chat, coding, agentic tools, and lower hallucination risk
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
Open flagship GLM for long-horizon coding agents and million-token context work
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
MiniMax multimodal model for long-context coding, perception, and agent planning
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Qwen vision-language model for visual reasoning, documents, and agent tasks
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
Initial DeepSeek V4 Flash snapshot for economical reasoning, coding, and million-token agent workloads
DeepSeek V4 Pro initial snapshot with million-token context and support for thinking and non-thinking modes
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
Qwen vision-language model for visual reasoning, documents, and agent tasks
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Flagship Qwen model for complex reasoning, coding, and agentic workflows
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek: DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 1.04858M | $0.15 | $0.6 | 2026-09-10 | |||
| OpenAI: GPT-6 Astraopenai/gpt-6-astra | 1.05M | $10 | $50 | 2026-09-04 | |||
| Google: Gemini 3.8 Flashgoogle/gemini-3.8-flash | 1.04858M | $0.75 | $3.75 | 2026-09-02 | |||
| Claude Fable 5.1anthropic/claude-fable-5-1 | 1M | $10 | $50 | 2026-09-01 | |||
| GLM-5.3zhipuai/glm-5.3 | 1M | $1.4 | $4.4 | 2026-08-14 | |||
| Google: Gemini 3.7 Flashgoogle/gemini-3.7-flash | 1.04858M | $0.75 | $3.75 | 2026-08-13 | |||
| Gemini Flash Latestgoogle/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | 2026-08-13 | |||
| Qwen3.8 2.4T A95Balibaba/qwen3.8-2.4t-a95b | 262.144K | $2 | $6 | 2026-08-12 | |||
| DeepSeek: DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 1.024M | $0.579 | $1.738 | 2026-08-12 | |||
| Grok 4.6xai/grok-4.6 | 500K | $2 | $6 | 2026-08-12 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b | 262.144K | — | — | 2026-08-11 | |||
| NVIDIA: Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning | 262.144K | $0.08 | $0.2 | 2026-08-11 | |||
| Meta: Muse Spark 1.2meta/muse-spark-1.2 | 1.04858M | $1.25 | $4.25 | 2026-08-05 | |||
| Sakana: Sakana Namazusakana/sakana-namazu | 262.144K | $0.95 | $4 | 2026-08-03 | |||
| Qwen3.8 Maxalibaba/qwen3.8-max | 1M | $2 | $6 | 2026-08-03 | |||
| DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1.04858M | $0.065 | $0.18 | 2026-07-31 | |||
| Claude Opus 5anthropic/claude-opus-5 | 1M | $5 | $25 | 2026-07-24 | |||
| Gemini Flash-Lite Latestgoogle/gemini-flash-lite-latest | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash | 1.04858M | $0.75 | $3.75 | 2026-07-21 | |||
| Google: Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| MoonshotAI: Kimi K3moonshotai/kimi-k3 | 1.04858M | $2.34 | $11.7 | 2026-07-16 | |||
| Thinking Machines: Inklingthinkingmachines/inkling | 1.04858M | $1 | $4.05 | 2026-07-15 | |||
| OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol | 1.05M | $2 | $10 | 2026-07-09 | |||
| OpenAI: GPT-5.6 Terraopenai/gpt-5.6-terra | 1.05M | $2 | $12 | 2026-07-09 | |||
| OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna | 1.05M | $0.2 | $1.2 | 2026-07-09 | |||
| Grok 4.5xai/grok-4.5 | 500K | $2 | $6 | 2026-07-08 | |||
| Anthropic: Claude Sonnet 5anthropic/claude-sonnet-5 | 1M | $2 | $10 | 2026-06-30 | |||
| Sakana: Fugu Ultrasakana/fugu-ultra | 1M | $5 | $30 | 2026-06-15 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.71 | $3.5 | 2026-06-12 | |||
| Kimi K2.7 Code Highspeedmoonshotai/kimi-k2.7-code-highspeed | 262.144K | $1.9 | $8 | 2026-06-12 | |||
| Anthropic: Claude Fable 5anthropic/claude-fable-5 | 1M | $10 | $50 | 2026-06-09 | |||
| Qwen3.7 Plusalibaba/qwen3.7-plus | 1M | $0.5 | $3 | 2026-06-02 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| Google: Nano Banana 2 (Gemini 3.1 Flash Image)google/gemini-3.1-flash-image | 131.072K | $0.5 | $3 | 2026-05-28 | |||
| Claude Opus 4.8anthropic/claude-opus-4-8 | 1M | $5 | $25 | 2026-05-28 | |||
| Qwen3.7 Maxalibaba/qwen3.7-max | 1M | $2.5 | $7.5 | 2026-05-21 | |||
| Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | 1.04858M | $0.25 | $1.5 | 2026-05-07 | |||
| Mistral Medium (latest)mistral/mistral-medium-latest | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Qwen3.6 Flashalibaba/qwen3.6-flash | 1M | $0.188 | $1.125 | 2026-04-27 | |||
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 1.024M | $0.087 | $0.173 | 2026-04-24 | |||
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | 1.024M | $0.948 | $1.896 | 2026-04-24 | |||
| DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash-0423 | 1M | $0.139 | $0.278 | 2026-04-23 | |||
| DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro-0423 | 1M | $1.32 | $3.96 | 2026-04-23 | |||
| OpenAI: GPT-5.5openai/gpt-5.5 | 1.05M | $5 | $30 | 2026-04-23 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| MoonshotAI: Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Qwen3.6 Max Previewalibaba/qwen3.6-max-preview | 262.144K | $1.3 | $7.8 | 2026-04-20 | |||
| Grok 4.3xai/grok-4.3 | 1M | $1.25 | $2.5 | 2026-04-17 |