Reasoning Tools Open weights

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads

nvidia/nemotron-3.5-lightning 2026-08-11 262.144K context $0.05/M input $0.2/M output
10 providers
Reasoning Tools JSON Open weights

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads

nvidia/nemotron-3.5-lightning-30b-a3b 2026-08-11 262.144K context Input not listed Output not listed
2 providers
Reasoning Tools Open weights

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

nvidia/nemotron-3-ultra-550b-a55b 2026-06-04 1M context $0.5/M input $2.5/M output
15 providers
Reasoning Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/nemotron-3.5-content-safety 2026-06-04 128K context $0.2/M input $0.2/M output
3 providers
Reasoning Tools Open weights

Open Nemotron omni model combining reasoning with text, vision, and audio

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning 2026-04-28 256K context $0.2/M input $0.8/M output
5 providers
Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/nemotron-3-content-safety 2026-04-16 128K context Input not listed Output not listed
1 provider
Open weights

Reranking model for improving retrieval quality in search and recommendation systems

nvidia/llama-nemotron-rerank-vl-1b-v2 2026-03-31 128K context Input not listed Output not listed
1 provider
Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/nemotron-cascade-2-30b-a3b 2026-03-24 256K context Input not listed Output not listed
1 provider
Tools Open weights

Nemotron multimodal model for visual reasoning and agentic AI workflows

nvidia/nemotron-voicechat 2026-03-16 128K context Input not listed Output not listed
1 provider
Reasoning Open weights

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

nvidia/nemotron-3-super-120b-a12b 2026-03-11 262.144K context $0.2/M input $0.8/M output
15 providers
Open weights

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

nvidia/llama-nemotron-embed-vl-1b-v2 2026-02-10 32.768K context Input not listed Output not listed
1 provider
Reasoning Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/nemotron-content-safety-reasoning-4b 2026-01-22 128K context Input not listed Output not listed
1 provider
Open weights

Small Nemotron 3 MoE for efficient coding, math, and long-context agents

nvidia/nemotron-3-nano-30b-a3b 2025-12-15 262.144K context $0.05/M input $0.2/M output
11 providers
Reasoning Tools Open weights

Nemotron multimodal model for visual reasoning and agentic AI workflows

nvidia/nemotron-nano-12b-v2-vl 2025-10-28 128K context $0.2/M input $0.6/M output
3 providers
Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/llama-3.1-nemotron-safety-guard-8b-v3 2025-10-28 128K context Input not listed Output not listed
1 provider
Reasoning Tools Open weights

Compact Nemotron model for efficient reasoning and deployable AI agents

nvidia/nemotron-nano-9b-v2 2025-08-18 131.072K context $0.06/M input $0.23/M output
4 providers
Reasoning Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.3-nemotron-super-49b-v1.5 2025-07-25 131.072K context $0.4/M input $0.4/M output
4 providers
Open weights

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

nvidia/mistral-nemotron 2025-06-11 128K context Input not listed Output not listed
Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.1-nemotron-70b-instruct 2025-04-15 128K context Input not listed Output not listed
2 providers
Reasoning Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.3-nemotron-super-49b-v1 2025-04-07 131.072K context Input not listed Output not listed
1 provider
Reasoning Tools Open weights

Flagship Nemotron model for high-throughput reasoning and complex agents

nvidia/llama-3.1-nemotron-ultra-253b 2025-04-07 128K context Input not listed Output not listed
2 providers
Tools Open weights

Compact Nemotron model for efficient reasoning and deployable AI agents

nvidia/nemotron-mini-4b-instruct 2024-08-21 128K context Input not listed Output not listed
1 provider
Open weights

No provider description is available for this model yet.

nvidia/NVIDIA-Nemotron-Parse-2.0 Not documented context Input not listed Output not listed