Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Native multimodal GLM model for efficient coding and long-horizon agent tasks
Dense 27B vision-language model for coding, agent tasks, and image and video understanding
Flagship GLM model for long-horizon coding, agents, and complex project delivery
Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Open flagship GLM for long-horizon coding agents and million-token context work
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
MiniMax multimodal model for long-context coding, perception, and agent planning
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Qwen vision-language model for visual reasoning, documents, and agent tasks
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Strong GLM coding model for agentic engineering, terminals, and repository generation
Open Gemma instruction model for efficient chat and self-hosted deployments
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open MiniMax flagship for coding agents, office automation, and complex environments
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen multimodal MoE for visual agents and long technical tasks
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
Budget GLM lane for fast coding help, routing, and everyday automation
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Safety model for policy screening, moderation, and risk-aware routing workflows
Efficient open MiniMax model built for coding agents and tool-heavy workflows
Fully open 70B multilingual LLM supporting 1800+ languages with 65K context. Trained on 15T tokens of compliant open data. Apache 2.0, EU AI Act compliant.
Efficient Qwen thinking model for local reasoning, math, and coding agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Dense open Qwen model for self-hosted chat, reasoning, and coding
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
Compact open Llama model for lightweight chat, drafting, and self-hosting
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Flash Nextalibaba/qwen3.8-flash-next | 262.144K | $0.12 | $0.4 | 2026-08-27 | |||
| GLM-5.3-Flashzhipuai/glm-5.3-flash | 1M | $0.075 | $0.25 | 2026-08-26 | |||
| Qwen3.8 27Balibaba/qwen3.8-27b | 262.144K | $0.1 | $0.4 | 2026-08-14 | |||
| GLM-5.3zhipuai/glm-5.3 | 1M | $1.4 | $4.4 | 2026-08-14 | |||
| Qwen3.8 2.4T A95Balibaba/qwen3.8-2.4t-a95b | 262.144K | $2 | $6 | 2026-08-12 | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1M | $0.05 | $0.16 | 2026-07-31 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| Hy3tencent/hy3 | 256K | $0.066 | $0.26 | 2026-07-06 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| GPT OSS Safeguard 120Bopenai/gpt-oss-safeguard-120b | 131.072K | $0.15 | $0.6 | 2025-10-29 | |||
| MiniMax-M2minimax/MiniMax-M2 | 204.8K | $0.3 | $1.2 | 2025-10-27 | |||
| Apertus 70Bswiss-ai/apertus-70b | 65.536K | $0.46 | $2.42 | 2025-09-02 | |||
| Qwen3-Next 80B-A3B (Thinking)alibaba/qwen3-next-80b-a3b-thinking | 131.072K | $0.5 | $6 | 2025-09 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | 128K | $0.02 | $0.04 | 2024-07-23 |