NVIDIA has completed its open model family Nemotron 3 with the three reasoning sizes Nano, Super and Ultra, and on August 11, 2026 added Nemotron 3.5 Lightning, a new model in the Nano size class built for long-running agents. With the release of Nemotron 3 Ultra (550B) on 4 June 2026, NVIDIA also offers its first frontier-scale model with open weights. Most models are released under the NVIDIA Open Model License; Nemotron 3 Ultra and Nemotron 3.5 Lightning are licensed under OpenMDW-1.1 per their model cards - both permit free download, modification, and commercial operation. innFactory AI Consulting from Germany advises on GDPR-compliant deployment. This overview reflects the state as of October 2026.
What is NVIDIA Nemotron?
Nemotron is NVIDIA’s family of open models (open weights) built specifically for agentic AI and reasoning. Unlike purely proprietary models, NVIDIA openly publishes weights, training data and training techniques – so the community can run, customize and further train the models for their own purposes. The current Nemotron 3 generation uses an efficient hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture, activating only a fraction of the parameters per request.
Current models at a glance
NVIDIA offers three reasoning sizes that scale with task complexity and hardware:
| Model | Parameters (active) | Context | Release | Licence |
|---|---|---|---|---|
| Nemotron 3 Nano | approx. 31.6B (3.6B) | 1M | 15 December 2025 | NVIDIA Open Model License |
| Nemotron 3 Super | 120B (12B) | – | March 2026 (GTC) | NVIDIA Open Model License |
| Nemotron 3 Ultra | 550B (55B) | up to 1M | 4 June 2026 | OpenMDW-1.1 |
| Nemotron 3.5 Lightning | 30B (3B) | 1M | 11 August 2026 | OpenMDW-1.1 |
| Nemotron-3-Diarization | 100M | – | 23 September 2026 | OpenMDW-1.1 |
Nemotron 3 Nano (30B-A3B)
- ~31.6B parameters total, ~3.6B active (hybrid Mamba-Transformer MoE)
- 1M token context window
- Highly efficient: up to 4x higher throughput than Nemotron Nano 2
- Self-hostable on moderate hardware (vLLM, SGLang, LM Studio, llama.cpp, Ollama)
- Release: 15 December 2025
Nemotron 3 Super (120B-A12B)
- 120B parameters total, ~12B active (hybrid Mamba-Transformer MoE)
- Strong agentic, reasoning and tool-calling capabilities
- Optimized for multi-agent systems and high-throughput workloads (e.g. IT ticket automation)
- Release: GTC, March 2026
Nemotron 3 Ultra (550B-A55B)
- 550B parameters total, ~55B active (hybrid LatentMoE: Mamba-2 + MoE + Attention, with Multi-Token Prediction)
- Up to 1M token context for long-context analysis
- Frontier reasoning over code, mathematics and science
- Release: 4 June 2026 (Computex)
Nemotron 3.5 Lightning
On August 11, 2026, NVIDIA released Nemotron 3.5 Lightning, a model in the size class of Nemotron 3 Nano:
- 30B total parameters, ~3B active (hybrid architecture combining Mamba-2, MoE, and attention layers)
- 1M token context window
- Per NVIDIA, up to 4x higher output speed and roughly 30% faster agentic task completion
- Single-GPU deployment (1x DGX Spark/GB10 or 1x H100), also runs on an RTX 5090 workstation
- Open weights, training data, and recipes under the OpenMDW-1.1 license (BF16 and NVFP4 checkpoints), not the NVIDIA Open Model License
- Per the model card: MMLU Pro 81.9, GPQA Diamond (no tools) 75.4
- Available via Hugging Face, ModelScope, OpenRouter, and build.nvidia.com (NIM microservice)
NVIDIA positions Nemotron 3.5 Lightning for long-running autonomous agents and as a sub-agent workhorse in multi-agent systems. It complements the Nano class but does not officially replace Nemotron 3 Nano - per NVIDIA, both remain actively listed.
Nemotron-3-Diarization
Nemotron-3-Diarization (23 September 2026) is a 100M-parameter speaker diarisation model under OpenMDW 1.1: it detects who speaks when (up to eight speakers, streaming and offline, latency from 80 ms to 30.4 s). Per the NVIDIA model card the DER on DIHARD III is 12.73% (predecessor: 19.09%). It adds speaker attribution to transcription models in meetings and call centres, does not identify individuals, and runs on your own infrastructure (NeMo, Transformers or C++ runtime).
Additional models: NVIDIA also maintains Nemotron Nano Omni (multimodal, vision/audio/language, 256K context), Nemotron 3.5 ASR (600M-parameter streaming speech recognition covering 40 language locales, since June 4, 2026), Nemotron 3.5 Content Safety (4B-parameter safety model for multimodal moderation), retriever/RAG models, plus OCR models. The older Llama Nemotron line (Nano 8B, Super 49B, Ultra 253B) is based on Llama 3.1 with 128K context.
Key Strengths
Efficiency through Hybrid MoE
The Mamba-Transformer MoE architecture activates only a fraction of total parameters per token. This significantly reduces inference costs – according to NVIDIA, Nemotron 3 Nano achieves up to 4x higher throughput than the previous generation and reduces the number of reasoning tokens.
Agentic & Tool Use
All Nemotron 3 models are built for agentic workflows: native tool calling, function calls and structured output for multi-agent systems.
Open Weights & Permissive License
The NVIDIA Open Model License is permissive and allows use, modification, distribution and commercial deployment – without an attribution requirement. NVIDIA also publishes training datasets and tools (NeMo Gym, NeMo RL, NeMo Evaluator).
Note: Some individual model cards on Hugging Face list slightly different license names (e.g. OpenMDW). We verify the specific license per model and version as part of our consulting.
EU Availability & GDPR Compliance
Because Nemotron is released as open weights, the cleanest sovereignty path is self-hosting on EU infrastructure – all data stays under your control.
Self-Hosting (recommended for sovereignty)
- Run on your own hardware or with an EU cloud provider (e.g. in Frankfurt)
- Full GDPR compliance, no dependency on US APIs
- Nano can already run on moderate hardware; Ultra requires multiple high-end GPUs (e.g. 8x B200/GB200, 16x H100 or 8x H200)
Managed via Hyperscalers (EU Regions)
- AWS: Per AWS, Amazon Bedrock offers Nemotron Nano 3 30B and Nemotron 3 Super 120B in-region in Ireland (eu-west-1), Milan (eu-south-1), and London (eu-west-2, UK). For Frankfurt (eu-central-1) and Stockholm (eu-north-1), the AWS sources contradict each other: the model cards list both regions for Nano 3 30B, while the region overview lists them for Super 120B – so check the Bedrock console before deploying. Nemotron 3 Ultra is not listed on Bedrock but is available via Amazon SageMaker JumpStart (self-managed endpoint, region freely selectable, e.g. Frankfurt)
- Microsoft Foundry: Nemotron Nano and Super run as a NIM microservice on Managed Compute - the customer selects the Azure region with suitable GPU capacity (e.g. West Europe) themselves; unlike the MAI models, this is not a serverless/Global Standard offering
- Google: Gemini Enterprise Agent Platform (formerly Vertex AI – rebranded in April 2026 at Cloud Next) in the Model Garden - specific EU regions for Nemotron there were not conclusively verified
Directly via NVIDIA
- build.nvidia.com and NVIDIA NIM microservices for hosted endpoints or containerized self-service deployments
- Hosted inference also via providers such as Together AI, Fireworks, DeepInfra, OpenRouter, Baseten
For sensitive data we recommend self-hosting in the EU. For a quick start, EU regions of the hyperscalers or NIM containers in your own cloud environment are suitable.
Integration with CompanyGPT
Thanks to open weights, Nemotron integrates well into our GDPR-compliant solution CompanyGPT. This lets you run a powerful reasoning and agent model entirely within your own or an EU-hosted environment – without company data flowing to third countries. innFactory AI Consulting handles selection, deployment and fine-tuning of the right Nemotron model.
Our Recommendation
Nemotron is an open model family for agentic and reasoning applications – and well suited for sovereign, GDPR-compliant deployments thanks to its open weights.
For most enterprises, we recommend:
- Nemotron 3 Nano or the newer Nemotron 3.5 Lightning for efficient, cost-effective agents and RAG – both self-hostable on moderate to single-GPU hardware; per NVIDIA, Lightning is especially well suited to long-running agentic workloads
- Nemotron 3 Super as a balanced choice for demanding multi-agent systems
- Nemotron 3 Ultra for frontier reasoning at maximum requirements (given appropriate GPU infrastructure, available via self-hosting or Amazon SageMaker JumpStart)
We are happy to advise you on model selection, self-hosting in the EU, and integration into existing workflows.
