Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding
4 providers
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding
Experimental multimodal DeepSeek V4 Flash model for image understanding, coding, and agentic work
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Dense 1B-class open-source model for on-device and resource-constrained use, with native long-context support, Think / No Think chat modes, and tool calling
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Flash Nextalibaba/qwen3.8-flash-next | 262.144K | $0.12 | $0.4 | 2026-08-27 | |||
| DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp | 1M | $0.15 | $0.6 | 2026-08-21 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1M | $0.035 | $0.07 | 2026-07-31 | |||
| MiniCPM5-1Bopenbmb/minicpm5-1b | 131.072K | $0.124 | $0.743 | 2026-05-19 |