No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...
GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful. It delivers more accurate answers with better contextualization and significantly...
This model always redirects to the latest model in the DeepSeek V4 Flash family.
Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...
Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...
Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.
GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed...
Rocinante 12B is designed for engaging storytelling and rich prose. Early testers have reported: - Expanded vocabulary with unique and expressive word choices - Enhanced creativity for vivid narratives -...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode Note: As of September...
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-Max-Thinkingdeepinfra/qwen/qwen3-max-thinking | 256K | $1.2 | $6 | — | |||
| nvidia.nemotron-super-3-120bbedrock_converse/nvidia.nemotron-super-3-120b | 256K | $0.15 | $0.65 | — | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 198K | $0.43 | $1.75 | — | |||
| OpenAI: o4 Mini (batch)openai/o4-mini:batch | 200K | $0.55 | $2.2 | — | |||
| qwen-coderdashscope/qwen-coder | 1M | $0.3 | $1.5 | — | |||
| au.anthropic.claude-sonnet-4-5-20250929-v1:0bedrock_converse/au.anthropic.claude-sonnet-4-5-20250929-v1:0 | 200K | $3.3 | $16.5 | — | |||
| grok-4.20xai/grok-4.20 | 1M | $1.25 | $2.5 | — | |||
| gemini-omni-1.1-flashgemini/gemini-omni-1.1-flash | 131.072K | $1.5 | $9 | — | |||
| nvidia.nemotron-nano-3-30bbedrock_converse/nvidia.nemotron-nano-3-30b | 262.144K | $0.06 | $0.24 | — | |||
| glm-5.3-flashzai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| meta-models/Muse-Glimmer-30Bdeepinfra/meta-models/muse-glimmer-30b | 131.072K | $0.3 | $1.2 | — | |||
| nvidia.nemotron-nano-9b-v2bedrock_converse/nvidia.nemotron-nano-9b-v2 | 128K | $0.06 | $0.23 | — | |||
| command-r7b-12-2024cohere_chat/command-r7b-12-2024 | 128K | $0.037 | $0.15 | — | |||
| zai-org/GLM-5.3together_ai/zai-org/glm-5.3 | 1.04858M | $1.4 | $4.4 | — | |||
| thinkingmachines/Inkling-Smalldeepinfra/thinkingmachines/inkling-small | 524.288K | $0.45 | $1.2 | — | |||
| nvidia.nemotron-nano-12b-v2bedrock_converse/nvidia.nemotron-nano-12b-v2 | 128K | $0.2 | $0.6 | — | |||
| kimi-k2.7-codemoonshot/kimi-k2.7-code | 262.144K | $0.95 | $4 | — | |||
| gemma-4-31b-itgemini/gemma-4-31b-it | 262.144K | — | — | — | |||
| Qwen/Qwen2-VL-7B-Instructnebius/qwen/qwen2-vl-7b-instruct | 131.072K | $0.02 | $0.06 | — | |||
| command-r-plus-08-2024cohere_chat/command-r-plus-08-2024 | 128K | $2.5 | $10 | — | |||
| gpt-4.1azure/gpt-4.1 | 1.04758M | $2 | $8 | — | |||
| gemma-4-26b-a4b-itgemini/gemma-4-26b-a4b-it | 262.144K | — | — | — | |||
| databricks-glm-5-3-flashdatabricks/databricks-glm-5-3-flash | 1.04858M | — | — | — | |||
| Qwen/Qwen2-VL-72B-Instructnebius/qwen/qwen2-vl-72b-instruct | 131.072K | $0.13 | $0.4 | — | |||
| google/gemma-4-31B-it-turbodeepinfra/google/gemma-4-31b-it-turbo | 262.144K | $0.09 | $0.34 | — | |||
| Qwen/Qwen3-Maxdeepinfra/qwen/qwen3-max | 256K | $1.2 | $6 | — | |||
| Qwen/Qwen2.5-VL-72B-Instructnebius/qwen/qwen2.5-vl-72b-instruct | 131.072K | $0.13 | $0.4 | — | |||
| command-r-pluscohere_chat/command-r-plus | 128K | $2.5 | $10 | — | |||
| XiaomiMiMo/MiMo-V2.5deepinfra/xiaomimimo/mimo-v2.5 | 262.144K | $0.4 | $2 | — | |||
| OpenAI: GPT-5.1 Chatopenai/gpt-5.1-chat | 128K | $1.25 | $10 | — | |||
| Qwen/Qwen2.5-Coder-7Bnebius/qwen/qwen2.5-coder-7b | 32.768K | $0.01 | $0.03 | — | |||
| google/gemini-3.5-flashdeepinfra/google/gemini-3.5-flash | 1M | $1.5 | $9 | — | |||
| Qwen: Qwen3.6 Flashqwen/qwen3.6-flash | 1M | $0.188 | $1.125 | — | |||
| OpenAI: GPT-5.3 Chatopenai/gpt-5.3-chat | 128K | $1.75 | $14 | — | |||
| DeepSeek: DeepSeek V4 Flash Latest~deepseek/deepseek-v4-flash-latest | 1.04858M | $0.04 | $0.1 | — | |||
| MoonshotAI: Kimi K2 0905moonshotai/kimi-k2-0905 | 262.144K | $0.6 | $2.5 | — | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 81.92K | $0.2 | $2.4 | — | |||
| NVIDIA: Nemotron Nano 9B V2 (free)nvidia/nemotron-nano-9b-v2:free | 128K | Free | Free | — | |||
| Nous: Hermes 4 70Bnousresearch/hermes-4-70b | 131.072K | $0.13 | $0.4 | — | |||
| ByteDance Seed: Seed 1.6 Flashbytedance-seed/seed-1.6-flash | 262.144K | $0.075 | $0.3 | — | |||
| ByteDance Seed: Seed 1.6bytedance-seed/seed-1.6 | 262.144K | $0.25 | $2 | — | |||
| OpenAI: GPT-5.6 Luna Proopenai/gpt-5.6-luna-pro | 1.05M | $0.2 | $1.2 | — | |||
| Mistral: Mistral Small 3mistralai/mistral-small-24b-instruct-2501 | 32.768K | $0.05 | $0.08 | — | |||
| TheDrummer: Rocinante 12Bthedrummer/rocinante-12b | 65.536K | $0.25 | $0.5 | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 262.144K | Free | Free | — | |||
| Claude Opus 5 (Fast)anthropic/claude-opus-5-fast | 1M | $10 | $50 | — | |||
| Qwen: Qwen3.7 Flashqwen/qwen3.7-flash | 1M | $0.03 | $0.13 | — | |||
| Poolside: Laguna S 2.1 (free)poolside/laguna-s-2.1:free | 262.144K | Free | Free | — | |||
| Meta: Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | 60K | $0.027 | $0.201 | — | |||
| anthropic/claude-sonnet-4-6deepinfra/anthropic/claude-sonnet-4-6 | 1M | $3 | $15 | — |