Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Model catalog
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
43 providers
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
15 providers
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1116.0 | 1M | $0.035 | $0.07 | 2026-07-31 | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1114.0 | 1M | $0.26 | $1.56 | — | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1101.0 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1101.0 | 1M | Free | Free | — |