Reasoning Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/nemotron-3.5-content-safety 2026-06-04 128K context $0.2/M input $0.2/M output
3 providers
Reasoning Tools Open weights

Open Nemotron omni model combining reasoning with text, vision, and audio

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning 2026-04-28 256K context $0.2/M input $0.8/M output
5 providers
Open weights

Reranking model for improving retrieval quality in search and recommendation systems

nvidia/llama-nemotron-rerank-vl-1b-v2 2026-03-31 128K context Input not listed Output not listed
1 provider
Open weights

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

nvidia/llama-nemotron-embed-vl-1b-v2 2026-02-10 32.768K context Input not listed Output not listed
1 provider
Reasoning Tools Open weights

Nemotron multimodal model for visual reasoning and agentic AI workflows

nvidia/nemotron-nano-12b-v2-vl 2025-10-28 128K context $0.2/M input $0.6/M output
3 providers
Open weights

No provider description is available for this model yet.

nvidia/NVIDIA-Nemotron-Parse-2.0 Not documented context Input not listed Output not listed

No provider description is available for this model yet.

nvidia/LocateAnything-3B Not documented context Input not listed Output not listed

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

nvidia/nemotron-3.5-content-safety:free 128K context Free input Free output

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 256K context Free input Free output

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...

nvidia/nemotron-nano-12b-v2-vl:free 128K context Free input Free output