⚡ NVIDIA Nemotron-3-Nano-Omni 30B-A3B Reasoning BF16

Native Multimodal Open Model • Mamba2-Transformer MoE Architecture (30B Total / 3B Active) • 256k Context Window

Architecture
Mamba2 MoE
Active Params
3B / 30B Total
Context Window
256,000 Tokens
Modalities
Text, Vision, Video, Audio
Precision
BF16 / FP8 / NVFP4
Instruction Input
Reasoning & Output Trace
🔍 Chain-of-Thought Execution Trace
Press "Run Nemotron Reasoning" to initialize step-by-step CoT token generation...
💡 Synthesized Output
Generated synthesis response will appear here.
Visual Input
🖼️
Click or Drag & Drop Image / Diagram / Document
Preview
Vision Perception Output
📷 Spatial Grounding & Feature Trace
Awaiting visual input analysis...
Visual analysis synthesis will render here.
Video Sequence Input
🎥
Click or Drag & Drop Video File (MP4/WebM)
Temporal Sequence Output
🎥 Frame Sampling & Action Tracking
Awaiting video temporal processing...
Video event sequence synthesis will render here.
Audio Waveform Input
🎙️
Click or Drag & Drop Audio File (WAV/MP3)
Audio Intelligence Output
🎙️ Phonetic & Acoustic Feature Trace
Awaiting audio waveform ingestion...
Speech transcription & acoustic summary will render here.

🏗️ NVIDIA Nemotron-3-Nano-Omni MoE Architecture

Nemotron-3-Nano-Omni unifies text, vision, video, and audio through a hybrid Mamba2 State-Space Model and Transformer Mixture-of-Experts (MoE). While storing 30 Billion total parameters for comprehensive knowledge retention, it selectively routes only 3 Billion active parameters (30B-A3B) per forward pass.

Total Parameters
30.2 Billion
Active / Token
3.0 Billion
Sequence Complexity
O(N) Linear Time
Context Capacity
256k Tokens