NVIDIA Nemotron
by NvIdia
Open-weight model family aimed at agentic workflows and long-context automation.
Lineage (then → 2026 now)
Nemotron-4 340B (June 2024, synthetic-data model) → Llama Nemotron reasoning models (2025) → Nemotron 3 (Nano Dec 2025; Super and Ultra 2026) → Nemotron 3.5 Lightning (2026-08-11). (Earlier-generation dates from Wikipedia.)
Nemotron 3 family
| Model | Size (total / active) | Release | Licence |
|---|---|---|---|
| Nano | 30B (HF) / 3.5B active (research page: 31.6B / 3.2B, 3.6B with embeddings) | 2025-12-15 (HF card) | NVIDIA Nemotron Open Model License |
| Super | 120B / 12B | 2026-03-11 (HF card) | NVIDIA Nemotron Open Model License |
| Nano Omni | 30B / ~3B active, multimodal text/image/video/audio | 2026-04-28 (NVIDIA blog, press) | NVIDIA open model agreement (press) |
| Ultra | 550B / 55B | 2026-06-04 (HF card; announced at Computex 2026-06-01) | OpenMDW-1.1 (HF card) |
| 3.5 ASR Streaming | 0.6B, cache-aware streaming speech recognition, ~40 language-locales | 2026-06-04 announced, HF release 2026-06-05 | NVIDIA Open Model License |
| 3.5 Lightning | 30B / 3B active, 128 routed + 1 shared experts | 2026-08-11 (HF card, SiliconANGLE) | OpenMDW-1.1 (HF card, NVIDIA docs) |
- Architecture: hybrid Mamba-Transformer MoE; Super and Ultra use LatentMoE with multi-token prediction. Ultra pretrained on 20T tokens (NVIDIA developer blog).
- Context: up to 1M tokens (HF cards; default config 256K for Nano and Super because of VRAM). Super needs at least 8x H100-80GB (card).
- Openness: NVIDIA publishes weights, training recipes and the data it can redistribute (research page; Ultra and 3.5 Lightning “fully open” incl. data and recipes).
- NeMo Switchyard: a model-routing library released alongside Lightning on 2026-08-11; Lightning is meant as the cheap execution target while a larger model (for example Ultra) plans (NVIDIA developer blog, SiliconANGLE).
- Other 2026 members (NVIDIA developer blog tag page): Nemotron 3 Embed (listed #1 on RTEB, 2026-07-16), Nemotron 3 Content Safety (2026-03-20).
- NemoClaw is a separate NVIDIA product (open-source stack that runs OpenClaw assistants with Nemotron models and the OpenShell sandbox, announced at GTC 2026), not a Nemotron model; its former alias here was removed.
Use
Suited to agent pipelines where long context and throughput matter. Ultra needs datacentre-class hardware (NVFP4 checkpoint on Blackwell); Nano and 3.5 Lightning run on a single GPU class. Cloud use via NVIDIA NIM / build.nvidia.com; prices not checked.
Lineage and what changed in 2026
Nemotron-4 340B (June 2024) → Llama Nemotron reasoning models (2025) → Nemotron 3 (Nano Dec 2025; Super March, Ultra June 2026) → Nemotron 3.5 (ASR June, Lightning August 2026). The 2025 note had only Nano; the newest releases moved to the OpenMDW-1.1 licence. Earlier-generation dates are from Wikipedia.
Related: Meta Muse, NvIdia, Ollama.
Open items
- Nemotron 3 Nano Omni licence wording and NIM pricing were not checked on a primary page.
- A creator claim that Lightning and Muse Glimmer 30B were the two headline local models of August 2026 and a Better Stack verdict were not verified and are removed. The earlier “ASR runs locally on CPU” claim is also dropped (unverified).
Sources
- https://research.nvidia.com/labs/nemotron/Nemotron-3/ (accessed 2026-10-02)
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (accessed 2026-10-02)
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
- https://docs.nvidia.com/nemotron/latest/nemotron/lightning35/README.html
- https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/
- https://siliconangle.com/2026/08/11/nvidia-releases-nemotron-3-5-lightning-nemo-switchyard-give-enterprise-ai-capability-options/
- https://blogs.nvidia.com/blog/nemotron-3-nano-omni-multimodal-ai-agents/
- https://www.marktechpost.com/2026/06/06/nvidia-releases-nemotron-3-5-asr-a-600m-parameter-cache-aware-streaming-model-transcribing-40-language-locales-in-real-time/
- https://nvidianews.nvidia.com/news/nvidia-announces-nemoclaw
- https://developer.nvidia.com/blog/tag/nemotron/
- https://en.wikipedia.org/wiki/Nemotron
(all accessed 2026-10-02 via fetch or search results)