NVIDIA logo

NVIDIA: NVIDIA: Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Source-linked Weight access not listed
API record Report
InputT
OutputT
Input priceFree
Output priceFree
Context1M
Max output65.536K
Providers0
Inference availability

Providers

Provider-specific identifiers, limits, and listed prices per million tokens. Every row links back to the provider's own documentation.

Report a price
No inference providers are listed

The model record exists, but no source-linked provider offer is available yet.

Submit a source
Specification

Capabilities

Recorded from the source catalog and provider listings.

? Reasoning Unknown
? Tool calling Unknown
? Structured output Unknown
? Attachments Unknown
× Vision input No
? Open weights Unknown
Creator
NVIDIA
Model family
Not documented
Knowledge cutoff
Not documented
License
Not documented
Release date
Not documented
Model ID
nvidia/nemotron-3-ultra-550b-a55b:free
Published evaluations

Benchmarks

Every result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.

Benchmark registry
Benchmark profile
Coding IndexCoding Index Agentic IndexAgentic Index IntelligenceIntelligence Index Design Arena: w…Design Arena: website Design Arena: c…Design Arena: codecategories Design Arena: g…Design Arena: gamedev Design Arena: d…Design Arena: dataviz Design Arena: u…Design Arena: uicomponent This model — Coding Index: 49.3 index (51.0th percentile) This model — Agentic Index: 21.7 index (54.4th percentile) This model — Intelligence Index: 23.4 index (42.7th percentile) This model — Design Arena: website: 1145 elo (32.0th percentile) This model — Design Arena: codecategories: 1154 elo (31.1th percentile) This model — Design Arena: gamedev: 1155 elo (35.5th percentile) This model — Design Arena: dataviz: 1156 elo (32.4th percentile) This model — Design Arena: uicomponent: 1151 elo (32.0th percentile)

Each axis is this model's percentile among the 8 benchmarks it has published results for, measured against every other model with a score on that same benchmark. Percentiles are used because benchmarks do not share a scale — a 60 on one is not a 60 on another. Hover any point for the raw score.

Coding Index index Source ↗
Coding Index #99 of 205
Agentic Index index Source ↗
Agentic Index #87 of 193
Intelligence Index index Source ↗
Intelligence Index #110 of 192
Design Arena: website elo Source ↗
Design Arena: website #118 of 175
Design Arena: codecategories elo Source ↗
Design Arena: codecategories #115 of 167
Design Arena: gamedev elo Source ↗
Design Arena: gamedev #107 of 166
Design Arena: dataviz elo Source ↗
Design Arena: dataviz #111 of 165
Design Arena: uicomponent elo Source ↗
Design Arena: uicomponent #109 of 161
Design Arena: 3d elo Source ↗
Design Arena: 3d #92 of 157
Design Arena: svg elo Source ↗
Design Arena: svg #93 of 116
Design Arena: asciiart elo Source ↗
Design Arena: asciiart #92 of 99

Each strip shows every published score for that benchmark, with this model marked. The lighter shape behind the ticks is the density of results, and the dashed line is the median.

Cost against capability

Price and performance

Listed input price plotted against Coding Index, the benchmark with the widest published coverage that this model appears in.

0 43 86 $0.01 $0.1 $1 $10 inclusionAI: Ling 3.0 Flash VL — $0.06/M, score 57.0 OpenAI: GPT-5.6 Sol (batch) — $1/M, score 77.4 Z.ai: GLM 5.1 — $0.966/M, score 55.8 OpenAI: GPT-5.5 (batch) — $2.5/M, score 74.9 Qwen: Qwen3.6 Plus — $0.325/M, score 54.5 Qwen: Qwen3.5-9B (batch) — $0.17/M, score 28.7 Nemotron 3.5 Lightning 30B A3B — $0.05/M, score 26.8 Google: Gemini 3.6 Flash (batch) — $0.375/M, score 69.2 OpenAI: GPT-5.4 (batch) — $1.25/M, score 71.1 OpenAI: gpt-oss-120b (batch) — $0.15/M, score 30.4 Claude Fable 5 — $10/M, score 76.5 Gemma 3 27B IT — $0.08/M, score 10.1 Gemma 3 12B IT — $0.05/M, score 5.8 OpenAI: GPT-4o-mini (batch) — $0.075/M, score 11.4 OpenAI: o3 Mini High (batch) — $0.55/M, score 16.3 Gemini 3.6 Flash — $0.75/M, score 69.2 Qwen: Qwen3.8 Max (0902) — $2/M, score 71.8 Gemini 3.5 Flash — $1.5/M, score 70.1 Nemotron 3 Super 120B A12B — $0.2/M, score 37.7 GPT-4 Turbo — $10/M, score 21.5 GPT-3.5-turbo — $0.5/M, score 10.7 Kimi K2.5 — $0.3/M, score 46.8 OpenAI: o1 (batch) — $7.5/M, score 39.7 GPT-4o mini — $0.15/M, score 11.4 DeepSeek V4 Flash — $0.15/M, score 52.0 Gemini 3.8 Flash — $0.75/M, score 76.3 LongCat-2.0 — $0.3/M, score 45.3 Claude Sonnet 5 — $2/M, score 71.5 Z.ai: GLM 5.3 Flash — $0.15/M, score 71.5 OpenAI: GPT-5.4 Nano (batch) — $0.1/M, score 56.1 SpaceXAI: Grok 4.5 — $2/M, score 72.4 Kimi K2.7 Code — $0.95/M, score 60.8 Inkling — $1.87/M, score 52.1 Nemotron 3 Ultra 550B A55B — $0.5/M, score 49.3 Anthropic: Claude Opus 4.8 — $5/M, score 74.3 OpenAI: GPT-4 Turbo (batch) — $5/M, score 21.5 Mistral: Mistral Medium 3.1 — $0.4/M, score 20.5 OpenAI: GPT-3.5 Turbo (batch) — $0.25/M, score 10.7 Anthropic: Claude Fable 5.1 (batch) — $5/M, score 81.6 Anthropic: Claude Opus 4.7 — $5/M, score 73.6 MiniMax: MiniMax M3 — $0.3/M, score 58.6 DeepSeek: DeepSeek V4 Pro 0813 (batch) — $0.66/M, score 68.8 Anthropic: Claude Opus 4.8 (batch) — $2.5/M, score 74.3 Gemma 4 31B IT — $0.09/M, score 43.4 Anthropic: Claude Fable 5.1 — $10/M, score 81.6 Google: Gemma 4 31B (batch) — $0.39/M, score 43.4 Qwen: Qwen3.8 2.4T A95B (batch) — $2/M, score 71.9 Mistral: Devstral 2 2512 — $0.4/M, score 31.3 MoonshotAI: Kimi K2.7 Code (batch) — $0.95/M, score 60.8 Gemma 3 4B IT — $0.04/M, score 2.7 MiMo-V2.5 — $0.14/M, score 56.8 Google: Gemini 3.5 Flash Lite (batch) — $0.15/M, score 49.3 Kimi K2 Thinking — $0.4/M, score 21.0 Mistral: Mistral Medium 3.5 (batch) — $0.75/M, score 46.9 Qwen: Qwen3.8 Max (0803) — $2/M, score 68.9 Mistral: Mistral Medium 3.1 (batch) — $0.2/M, score 20.5 OpenAI: GPT-5.6 Terra (batch) — $1/M, score 76.7 Google: Gemini 3.8 Flash (batch) — $0.375/M, score 76.3 Gemma 4 26B A4B IT — $0.042/M, score 39.3 GPT OSS 20B — $0.02/M, score 20.7 Z.ai: GLM 5.3 Flash (batch) — $0.075/M, score 71.5 Gemini 3.5 Flash Lite — $0.3/M, score 49.3 Inkling Small — $0.45/M, score 52.9 Qwen: Qwen3.7 Plus — $0.32/M, score 55.9 Qwen: Qwen3.8 27B — $0.214/M, score 68.1 Anthropic: Claude Opus 4.7 (batch) — $2.5/M, score 73.6 GPT-5.6 Sol — $4/M, score 77.4 Anthropic: Claude Fable 5 (batch) — $5/M, score 76.5 Kimi K3 — $3/M, score 76.2 Gemini 3.7 Flash — $0.75/M, score 76.1 OpenAI: GPT-5.1 (batch) — $0.625/M, score 49.4 MoonshotAI: Kimi K3 (batch) — $3/M, score 76.2 Anthropic: Claude Sonnet 5 (batch) — $1/M, score 71.5 IBM: Granite 4.2 8B — $0.06/M, score 22.4 MiniMax: MiniMax M2.7 — $0.3/M, score 52.6 Google: Gemini 3.7 Flash (batch) — $0.375/M, score 76.1 SpaceXAI: Grok 4.6 — $2/M, score 76.8 Claude Opus 5 — $5/M, score 78.0 Claude Opus 5 (batch) — $2.5/M, score 78.0 MiniMax: MiniMax M3 (batch) — $0.3/M, score 58.6 Gemini 3.1 Flash Lite Preview — $0.25/M, score 34.7 GPT-5.4 — $2.5/M, score 71.1 GPT-5 Mini — $0.25/M, score 15.6 OpenAI: GPT-6 Astra (batch) — $5/M, score 76.9 Qwen: Qwen3 14B — $0.227/M, score 13.8 Google: Gemini 3.1 Pro Preview (batch) — $1/M, score 68.8 Nemotron 3 Nano 30B A3B — $0.05/M, score 14.4 OpenAI: GPT-4.1 Nano (batch) — $0.05/M, score 11.1 Mistral: Ministral 3 8B 2512 — $0.15/M, score 9.7 Mistral: Mistral Large 3 2512 (batch) — $0.25/M, score 20.1 DeepSeek V3.2 — $0.18/M, score 44.2 Anthropic: Claude Sonnet 4.6 — $3/M, score 63.0 Muse Spark 1.1 — $1.25/M, score 71.3 Anthropic: Claude Sonnet 4.5 — $3/M, score 52.1 GPT-5 — $1.25/M, score 37.8 DeepSeek V4 Pro 0813 — $0.442/M, score 68.8 OpenAI: GPT-5 Mini (batch) — $0.125/M, score 15.6 GPT-4.1 nano — $0.1/M, score 11.1 Google: Gemini 3.5 Flash (batch) — $0.75/M, score 70.1 Z.ai: GLM 5.3 (batch) — $0.7/M, score 74.8 Z.ai: GLM 5.3 — $1.4/M, score 74.8 Z.ai: GLM 5.2 (batch) — $0.7/M, score 68.8 SpaceXAI: Grok 4.3 — $1.25/M, score 42.2 DeepSeek V4 Flash 0731 — $0.035/M, score 69.1 GPT-5.6 Terra — $2/M, score 76.7 DeepSeek V4 Pro — $0.435/M, score 59.4 NVIDIA: Nemotron 3 Ultra (batch) — $0.6/M, score 49.3 GPT-5.4 nano — $0.2/M, score 56.1 Kimi K2.6 — $0.95/M, score 61.8 GPT-4 — $30/M, score 13.1 Thinking Machines: Inkling Small (batch) — $0.5/M, score 52.9 Mistral: Mistral Small 4 (batch) — $0.075/M, score 26.6 DeepSeek: DeepSeek V4 Flash 0731 (batch) — $0.11/M, score 69.1 Thinking Machines: Inkling (batch) — $1/M, score 52.1 Anthropic: Claude Sonnet 4.6 (batch) — $1.5/M, score 63.0 Google: Gemini 2.5 Pro (batch) — $0.625/M, score 33.3 Gemini 3.1 Pro Preview — $2/M, score 68.8 Gemini 2.5 Pro — $1.25/M, score 33.3 Mistral: Mistral Large 3 2512 — $0.5/M, score 20.1 DeepSeek: DeepSeek V3.1 Terminus — $0.27/M, score 43.5 OpenAI: gpt-oss-20b (batch) — $0.05/M, score 20.7 SpaceXAI: Grok 4.3 (batch) — $1/M, score 42.2 DeepSeek-R1 — $0.7/M, score 24.6 GPT-4o (2024-05-13) — $5/M, score 24.2 Anthropic: Claude Haiku 4.5 (batch) — $0.5/M, score 43.9 Anthropic: Claude Sonnet 4.5 (batch) — $1.5/M, score 52.1 GPT-4.1 mini — $0.4/M, score 20.2 inclusionAI: Ling 3.0 Flash — $0.021/M, score 50.6 GPT-5.5 — $5/M, score 74.9 Muse Spark 1.2 — $1.25/M, score 72.2 OpenAI: GPT-5.6 Luna (batch) — $0.1/M, score 71.4 OpenAI: GPT-5.4 Mini (batch) — $0.375/M, score 56.1 GPT-6 Astra — $10/M, score 76.9 GPT-5.1 — $1.25/M, score 49.4 Qwen: Qwen3 Next 80B A3B Thinking — $0.15/M, score 17.4 Qwen: Qwen3.8 2.4T A95B — $2/M, score 71.9 Mistral: Ministral 3 8B 2512 (batch) — $0.075/M, score 9.7 Solar Pro 4 — $0.3/M, score 52.7 Hy3 preview — $0.066/M, score 58.8 OpenAI: GPT-4.1 Mini (batch) — $0.2/M, score 20.2 GPT OSS 120B — $0.03/M, score 30.4 Inception: Mercury 2 — $0.25/M, score 31.1 o1 — $15/M, score 39.7 MiMo-V2.5-Pro — $0.435/M, score 60.2 GPT-5.6 Luna — $0.2/M, score 71.4 Z.ai: GLM 4.6 — $0.43/M, score 45.8 OpenAI: GPT-5 (batch) — $0.625/M, score 37.8 Anthropic: Claude Sonnet 4 — $3/M, score 37.6 GPT-5.4 mini — $0.75/M, score 56.1 Trinity Large Thinking — $0.25/M, score 25.8 Qwen: Qwen3 30B A3B Thinking 2507 — $0.2/M, score 12.1 Z.ai: GLM 5.2 — $0.6/M, score 68.8 Qwen: Qwen3.7 Max — $1.475/M, score 66.0 SpaceXAI: Grok Build 0.1 — $1/M, score 51.5 Qwen: Qwen3.6 35B A3B — $0.1/M, score 41.9 Qwen: Qwen3.6 27B — $0.3/M, score 53.7 inclusionAI: Ling-2.6-flash — $0.01/M, score 25.3 Mistral: Mistral Small 4 — $0.15/M, score 26.6 Kwaipilot: KAT-Coder-Pro V2 — $0.3/M, score 59.5 Qwen: Qwen3.5-9B — $0.1/M, score 28.7 Qwen: Qwen3.5-35B-A3B — $0.312/M, score 37.0 Qwen: Qwen3.5-122B-A10B — $0.26/M, score 45.7 Upstage: Solar Pro 3 — $0.15/M, score 16.2 Z.ai: GLM 4.7 — $0.4/M, score 45.3 Amazon: Nova 2 Lite — $0.3/M, score 23.0 Mistral: Ministral 3 3B 2512 — $0.1/M, score 4.8 Google: Gemma 3n 4B — $0.06/M, score 3.2 Meta: Llama 4 Maverick — $0.2/M, score 16.3 OpenAI: o3 Mini High — $1.1/M, score 16.3 Meta: Llama 3.3 70B Instruct — $0.1/M, score 11.9 Meta: Llama 3.1 8B Instruct — $0.05/M, score 5.4 Nex AGI: Nex-N2-Pro — $0.25/M, score 59.1 Mistral: Mistral Medium 3.5 — $1.5/M, score 46.9 inclusionAI: Ring-2.6-1T — $0.075/M, score 42.8 IBM: Granite 4.1 8B — $0.05/M, score 9.5 Qwen: Qwen3.5 397B A17B — $0.55/M, score 48.2 Qwen: Qwen3 Coder Next — $0.12/M, score 36.2 Mistral: Ministral 3 14B 2512 — $0.2/M, score 14.4 Anthropic: Claude Haiku 4.5 — $1/M, score 43.9 Qwen: Qwen3 235B A22B Thinking 2507 — $0.23/M, score 22.1 Qwen: Qwen3 8B — $0.117/M, score 9.0 Qwen: Qwen3 32B — $0.08/M, score 15.3 Meta: Llama 4 Scout — $0.1/M, score 8.2 DeepSeek: DeepSeek V3 0324 — $0.25/M, score 21.2 Cohere: Command A — $2.5/M, score 27.8 Step 3.7 Flash — $0.185/M, score 39.6 Input price per million tokens (log scale) Index

The stepped line is the efficient frontier: at each price, the best score available for that money or less. A model sitting on it is not being beaten by anything cheaper. Price is log-scaled because listed rates span four orders of magnitude. Only models with both a listed price and a score on this benchmark can appear.

Catalog activity

Change log

Field-level changes detected between successful source imports.

Full change log
No changes recorded

This record has not changed within the retained import history.

Public API

Use this record

Fetch the complete source-linked model record. No key, no account, no rate-limited tier.

API documentation
Endpoint
GET https://model.kyssta.lol/api/v1/models/nvidia/nemotron-3-ultra-550b-a55b:free
curl
curl "https://model.kyssta.lol/api/v1/models/nvidia/nemotron-3-ultra-550b-a55b:free"
Common questions

Frequently asked questions

Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.

What is NVIDIA: Nemotron 3 Ultra (free)?

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it. It is published by NVIDIA and catalogued here from OpenRouter.

How much does NVIDIA: Nemotron 3 Ultra (free) cost?

This model is documented as free to use at the listed providers.

What is the context length of NVIDIA: Nemotron 3 Ultra (free)?

NVIDIA: Nemotron 3 Ultra (free) accepts up to 1M tokens of context and returns up to 65.536K output tokens.

More models from NVIDIA

View all →