Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

qwen/qwen3.6-max-preview 262.144K context $1.027/M input $6.162/M output

No provider description is available for this model yet.

deepinfra/openai/gpt-oss-120b 131.072K context $0.05/M input $0.45/M output

No provider description is available for this model yet.

deepinfra/openai/gpt-oss-20b 131.072K context $0.04/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-4.5 131.072K context $0.4/M input $1.6/M output

No provider description is available for this model yet.

deepseek/deepseek-coder 128K context $0.14/M input $0.28/M output

No provider description is available for this model yet.

bedrock_converse/deepseek.v3-v1:0 163.84K context $0.58/M input $1.68/M output

No provider description is available for this model yet.

bedrock_converse/deepseek.v3.2 163.84K context $0.62/M input $1.85/M output

No provider description is available for this model yet.

volcengine/glm-4-7-251222 204.8K context Input not listed Output not listed

No provider description is available for this model yet.

volcengine/kimi-k2-thinking-251104 229.376K context Input not listed Output not listed

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:batch 262.144K context $0.39/M input $0.97/M output

No provider description is available for this model yet.

fireworks_ai/muse-glimmer-30b 131.072K context $0.35/M input $1.5/M output

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

qwen/qwen3-8b 131.072K context $0.117/M input $0.455/M output

No provider description is available for this model yet.

together_ai/qwen/qwq-32b 131.072K context $1.2/M input $1.2/M output

No provider description is available for this model yet.

fireworks_ai/qwen3p8-max 262.144K context $2/M input $6/M output

No provider description is available for this model yet.

fireworks_ai/kimi-k3-us 1.04858M context $3.3/M input $16.5/M output

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

deepseek/deepseek-r1-0528 163.84K context $0.5/M input $2.15/M output

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

z-ai/glm-4.5 131.072K context $0.6/M input $2.2/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.7-code 262.144K context $1.05/M input $4.4/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.6 262.144K context $1.045/M input $4.4/M output

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

qwen/qwen3-235b-a22b-2507 262.144K context $0.087/M input $0.35/M output

No provider description is available for this model yet.

fireworks_ai/kimi-k3-fast 1.04858M context $4.5/M input $22.5/M output

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

qwen/qwen3-coder-30b-a3b-instruct 262.144K context $0.07/M input $0.28/M output

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

openai/gpt-oss-20b:free 131.072K context Free input Free output

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.9/M output