Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
21 providers
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Open Nemotron omni model combining reasoning with text, vision, and audio
Nemotron multimodal model for visual reasoning and agentic AI workflows
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open Whisper checkpoint for robust multilingual transcription and captioning
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Inklingthinkingmachines/inkling | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Nemotron 3 Nano Omni 30B A3B Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 256K | $0.2 | $0.8 | 2026-04-28 | |||
| Nemotron VoiceChatnvidia/nemotron-voicechat | 128K | — | — | 2026-03-16 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Whisper 3 Largeopenai/whisper-large-v3 | 448 | $0.002 | $0.002 | 2024-10-01 |