Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
No provider description is available for this model yet.
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen: Qwen3.6 Max Previewqwen/qwen3.6-max-preview | 262.144K | $1.027 | $6.162 | — | |||
| mistralai/Mistral-Small-3.2-24B-Instruct-2506deepinfra/mistralai/mistral-small-3.2-24b-instruct-2506 | 128K | $0.075 | $0.2 | — | |||
| apac.anthropic.claude-sonnet-4-20250514-v1:0bedrock_converse/apac.anthropic.claude-sonnet-4-20250514-v1:0 | 1M | $3 | $15 | — | |||
| moonshotai/Kimi-K2-Instruct-0905deepinfra/moonshotai/kimi-k2-instruct-0905 | 262.144K | $0.5 | $2 | — | |||
| nvidia/Llama-3.1-Nemotron-70B-Instructdeepinfra/nvidia/llama-3.1-nemotron-70b-instruct | 131.072K | $0.6 | $0.6 | — | |||
| nvidia/Llama-3.3-Nemotron-Super-49B-v1.5deepinfra/nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.1 | $0.4 | — | |||
| us-gov-west-1/nvidia.nemotron-nano-9b-v2bedrock/us-gov-west-1/nvidia.nemotron-nano-9b-v2 | 128K | $0.072 | $0.276 | — | |||
| openai/gpt-oss-120bdeepinfra/openai/gpt-oss-120b | 131.072K | $0.05 | $0.45 | — | |||
| openai/gpt-oss-20bdeepinfra/openai/gpt-oss-20b | 131.072K | $0.04 | $0.15 | — | |||
| zai-org/GLM-4.5deepinfra/zai-org/glm-4.5 | 131.072K | $0.4 | $1.6 | — | |||
| deepseek-coderdeepseek/deepseek-coder | 128K | $0.14 | $0.28 | — | |||
| deepseek.v3-v1:0bedrock_converse/deepseek.v3-v1:0 | 163.84K | $0.58 | $1.68 | — | |||
| deepseek.v3.2bedrock_converse/deepseek.v3.2 | 163.84K | $0.62 | $1.85 | — | |||
| glm-4-7-251222volcengine/glm-4-7-251222 | 204.8K | — | — | — | |||
| kimi-k2-thinking-251104volcengine/kimi-k2-thinking-251104 | 229.376K | — | — | — | |||
| eu.amazon.nova-lite-v1:0bedrock_converse/eu.amazon.nova-lite-v1:0 | 300K | $0.078 | $0.312 | — | |||
| eu.amazon.nova-micro-v1:0bedrock_converse/eu.amazon.nova-micro-v1:0 | 128K | $0.046 | $0.184 | — | |||
| eu.amazon.nova-pro-v1:0bedrock_converse/eu.amazon.nova-pro-v1:0 | 300K | $1.05 | $4.2 | — | |||
| eu.anthropic.claude-3-5-haiku-20241022-v1:0bedrock/eu.anthropic.claude-3-5-haiku-20241022-v1:0 | 200K | $0.25 | $1.25 | — | |||
| eu.anthropic.claude-haiku-4-5-20251001-v1:0bedrock_converse/eu.anthropic.claude-haiku-4-5-20251001-v1:0 | 200K | $1.1 | $5.5 | — | |||
| eu.anthropic.claude-3-5-sonnet-20240620-v1:0bedrock/eu.anthropic.claude-3-5-sonnet-20240620-v1:0 | 200K | $3 | $15 | — | |||
| eu.anthropic.claude-3-5-sonnet-20241022-v2:0bedrock/eu.anthropic.claude-3-5-sonnet-20241022-v2:0 | 200K | $3 | $15 | — | |||
| apac.anthropic.claude-3-sonnet-20240229-v1:0bedrock/apac.anthropic.claude-3-sonnet-20240229-v1:0 | 200K | $3 | $15 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 262.144K | $0.39 | $0.97 | — | |||
| muse-glimmer-30bfireworks_ai/muse-glimmer-30b | 131.072K | $0.35 | $1.5 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 131.072K | $0.117 | $0.455 | — | |||
| apac.anthropic.claude-haiku-4-5-20251001-v1:0bedrock_converse/apac.anthropic.claude-haiku-4-5-20251001-v1:0 | 200K | $1.1 | $5.5 | — | |||
| Qwen/QwQ-32Btogether_ai/qwen/qwq-32b | 131.072K | $1.2 | $1.2 | — | |||
| qwen3p8-maxfireworks_ai/qwen3p8-max | 262.144K | $2 | $6 | — | |||
| @cf/meta/llama-guard-3-8bcloudflare/@cf/meta/llama-guard-3-8b | 131.072K | $0.484 | $0.03 | — | |||
| apac.anthropic.claude-3-haiku-20240307-v1:0bedrock/apac.anthropic.claude-3-haiku-20240307-v1:0 | 200K | $0.25 | $1.25 | — | |||
| kimi-k3-usfireworks_ai/kimi-k3-us | 1.04858M | $3.3 | $16.5 | — | |||
| DeepSeek: R1 0528deepseek/deepseek-r1-0528 | 163.84K | $0.5 | $2.15 | — | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 131.072K | $0.6 | $2.2 | — | |||
| nvidia/NVIDIA-Nemotron-Nano-9B-v2together_ai/nvidia/nvidia-nemotron-nano-9b-v2 | 131.072K | $0.06 | $0.25 | — | |||
| mistralai/Ministral-3-14B-Instruct-2512together_ai/mistralai/ministral-3-14b-instruct-2512 | 262.144K | $0.2 | $0.2 | — | |||
| FW-Kimi-K2.7-Codeazure_ai/fw-kimi-k2.7-code | 262.144K | $1.05 | $4.4 | — | |||
| Qwen/Qwen3-VL-8B-Instructtogether_ai/qwen/qwen3-vl-8b-instruct | 262.144K | $0.18 | $0.68 | — | |||
| apac.anthropic.claude-3-5-sonnet-20241022-v2:0bedrock/apac.anthropic.claude-3-5-sonnet-20241022-v2:0 | 200K | $3 | $15 | — | |||
| FW-Kimi-K2.6azure_ai/fw-kimi-k2.6 | 262.144K | $1.045 | $4.4 | — | |||
| Qwen: Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | 262.144K | $0.087 | $0.35 | — | |||
| apac.anthropic.claude-3-5-sonnet-20240620-v1:0bedrock/apac.anthropic.claude-3-5-sonnet-20240620-v1:0 | 200K | $3 | $15 | — | |||
| apac.amazon.nova-pro-v1:0bedrock_converse/apac.amazon.nova-pro-v1:0 | 300K | $0.84 | $3.36 | — | |||
| kimi-k3-fastfireworks_ai/kimi-k3-fast | 1.04858M | $4.5 | $22.5 | — | |||
| Qwen/Qwen3-VL-32B-Instructtogether_ai/qwen/qwen3-vl-32b-instruct | 262.144K | $0.5 | $1.5 | — | |||
| apac.amazon.nova-micro-v1:0bedrock_converse/apac.amazon.nova-micro-v1:0 | 128K | $0.037 | $0.148 | — | |||
| us-gov.xai.grok-4.6bedrock_converse/us-gov.xai.grok-4.6 | 500K | $2.64 | $7.92 | — | |||
| Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | 262.144K | $0.07 | $0.28 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 131.072K | Free | Free | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 262.144K | $0.1 | $0.9 | — |