NVIDIA Nemotron

by NvIdia

Open-weight model family aimed at agentic workflows and long-context automation.

Lineage (then → 2026 now)

Nemotron-4 340B (June 2024, synthetic-data model) → Llama Nemotron reasoning models (2025) → Nemotron 3 (Nano Dec 2025; Super and Ultra 2026) → Nemotron 3.5 Lightning (2026-08-11). (Earlier-generation dates from Wikipedia.)

Nemotron 3 family

ModelSize (total / active)ReleaseLicence
Nano30B (HF) / 3.5B active (research page: 31.6B / 3.2B, 3.6B with embeddings)2025-12-15 (HF card)NVIDIA Nemotron Open Model License
Super120B / 12B2026-03-11 (HF card)NVIDIA Nemotron Open Model License
Nano Omni30B / ~3B active, multimodal text/image/video/audio2026-04-28 (NVIDIA blog, press)NVIDIA open model agreement (press)
Ultra550B / 55B2026-06-04 (HF card; announced at Computex 2026-06-01)OpenMDW-1.1 (HF card)
3.5 ASR Streaming0.6B, cache-aware streaming speech recognition, ~40 language-locales2026-06-04 announced, HF release 2026-06-05NVIDIA Open Model License
3.5 Lightning30B / 3B active, 128 routed + 1 shared experts2026-08-11 (HF card, SiliconANGLE)OpenMDW-1.1 (HF card, NVIDIA docs)
  • Architecture: hybrid Mamba-Transformer MoE; Super and Ultra use LatentMoE with multi-token prediction. Ultra pretrained on 20T tokens (NVIDIA developer blog).
  • Context: up to 1M tokens (HF cards; default config 256K for Nano and Super because of VRAM). Super needs at least 8x H100-80GB (card).
  • Openness: NVIDIA publishes weights, training recipes and the data it can redistribute (research page; Ultra and 3.5 Lightning “fully open” incl. data and recipes).
  • NeMo Switchyard: a model-routing library released alongside Lightning on 2026-08-11; Lightning is meant as the cheap execution target while a larger model (for example Ultra) plans (NVIDIA developer blog, SiliconANGLE).
  • Other 2026 members (NVIDIA developer blog tag page): Nemotron 3 Embed (listed #1 on RTEB, 2026-07-16), Nemotron 3 Content Safety (2026-03-20).
  • NemoClaw is a separate NVIDIA product (open-source stack that runs OpenClaw assistants with Nemotron models and the OpenShell sandbox, announced at GTC 2026), not a Nemotron model; its former alias here was removed.

Use

Suited to agent pipelines where long context and throughput matter. Ultra needs datacentre-class hardware (NVFP4 checkpoint on Blackwell); Nano and 3.5 Lightning run on a single GPU class. Cloud use via NVIDIA NIM / build.nvidia.com; prices not checked.

Lineage and what changed in 2026

Nemotron-4 340B (June 2024) → Llama Nemotron reasoning models (2025) → Nemotron 3 (Nano Dec 2025; Super March, Ultra June 2026) → Nemotron 3.5 (ASR June, Lightning August 2026). The 2025 note had only Nano; the newest releases moved to the OpenMDW-1.1 licence. Earlier-generation dates are from Wikipedia.

Related: Meta Muse, NvIdia, Ollama.

Open items

  • Nemotron 3 Nano Omni licence wording and NIM pricing were not checked on a primary page.
  • A creator claim that Lightning and Muse Glimmer 30B were the two headline local models of August 2026 and a Better Stack verdict were not verified and are removed. The earlier “ASR runs locally on CPU” claim is also dropped (unverified).

Sources