↓ Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM NVIDIA USA

NVIDIA Nemotron

NVIDIA Nemotron 3 (Nano, Super, Ultra), Nemotron 3.5 Lightning and Nemotron-3-Diarization: open models for agentic AI, EU self-hosting and GDPR assessment.

License NVIDIA Open Model License (Nemotron 3 Nano/Super, per their Hugging Face model cards, permissive, commercial use allowed); Nemotron 3 Ultra and Nemotron 3.5 Lightning are licensed under OpenMDW-1.1 per their model cards (also permissive, commercial use allowed)
GDPR Hosting Available
Context 128K-1M Tokens
Modality Text, Code → Text, Code

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Nemotron 3.5 Lightning (30B-A3B) Recommended
August 11, 2026
30B total parameters, ~3B active (hybrid Mamba-2 + MoE + Attention) 1M token context window Per NVIDIA, up to 4x higher output speed and roughly 30% faster agentic task completion versus comparably sized models Single-GPU deployment (e.g. 1x DGX Spark/GB10 or 1x H100), also runs on RTX 5090 workstations Open weights, training data, and recipes under OpenMDW-1.1 (BF16 and NVFP4 checkpoints); commercial use permitted Per the NVIDIA model card: MMLU Pro 81.9, GPQA Diamond (no tools) 75.4 (NVFP4 checkpoint)
Available only since August 11, 2026 - less broad real-world track record than Nemotron 3 Nano/Super/Ultra so far Licensed under OpenMDW-1.1 rather than the NVIDIA Open Model License - consider separately in compliance reviews
Current
Nemotron-3-Diarization (100M, speaker diarization)
23 September 2026
Detects who spoke when: up to eight speakers, streaming and offline, latency configurable from 80 ms to 30.4 s 100M parameters (31-layer transformer), 16 kHz mono audio as input, per-frame speaker activity at 10 ms resolution as output Per the NVIDIA model card DER 12.73% on DIHARD III (predecessor 19.09%), trained on about 10,000 hours of real conversations and 82,611 hours of synthetic mixtures Complements transcription models such as Whisper or gpt-transcribe with speaker attribution in meetings and call centres
Maximum eight speakers, no speaker identification (order of appearance only) Training languages per the model card English, Mandarin, Hindi and others – German not explicitly listed, evaluate before use
Current
Nemotron 3 Ultra (550B-A55B)
4 June 2026
550B total parameters, ~55B active (hybrid LatentMoE: Mamba-2 + MoE + Attention) Up to 1M token context window Frontier reasoning over code, mathematics and science Open weights on Hugging Face Available via NVIDIA NIM on build.nvidia.com
Very high hardware requirements for self-hosting (e.g. 8x B200/GB200, 16x H100 or 8x H200)
Current
Nemotron 3 Super (120B-A12B) Recommended
March 2026 (GTC)
120B total parameters, ~12B active (hybrid Mamba-Transformer MoE) Strong agentic, reasoning and tool-calling capabilities Optimized for multi-agent systems and high-throughput workloads Open weights on Hugging Face
Medium to high resource requirements for self-hosting
Current
Nemotron 3 Nano (30B-A3B) Recommended
15 December 2025
~31.6B total parameters, ~3.6B active (hybrid Mamba-Transformer MoE) 1M token context window Highly efficient – up to 4x higher throughput than Nemotron Nano 2 Self-hostable on consumer hardware (vLLM, SGLang, LM Studio, llama.cpp, Ollama) Open weights on Hugging Face
Smallest variant – less suited for highly complex reasoning tasks
Current
Nemotron 3 Nano Omni (30B-A3B)
28 April 2026
Multimodal: unifies vision, audio and language 256K token context window Optimized for efficient multimodal AI agents
Newer model – ecosystem support still maturing
Current
Llama Nemotron Ultra (253B)
2025
Reasoning variant based on Llama 3.1 128K token context Open weights, established Llama ecosystem
Previous generation – superseded by Nemotron 3
Current

Use Cases

Typical applications for this model

Agentic Workflows & Multi-Agent Systems
Reasoning over Code, Mathematics & Science
Tool Calling & Function Calls
RAG & Knowledge Retrieval
Self-Hosted Deployments on EU Infrastructure
Cost-Efficient Inference (Nano)

Technical Details

API, features and capabilities

API & Availability
Availability Public
Features & Capabilities
Tool Use Function Calling Structured Output Reasoning Mode
Training & Knowledge
Knowledge Cutoff May 2026 (Ultra), February 2026 (Super)
Fine-Tuning Available (LoRA, Full, PEFT)
Language Support
Best Quality English
Supported Multilingual (incl. German, Spanish, French, Italian, Japanese)
German is supported; best quality in English. Ultra supports 10 languages beyond English.

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-Hosted
Own Infrastructure (EU)
Open weights - full control, cleanest sovereignty path
AWS
Ireland (eu-west-1), Milan (eu-south-1); Frankfurt/Stockholm depending on model
Amazon Bedrock: per AWS, Nemotron Nano 3 30B and Nemotron 3 Super 120B are available in-region in Ireland (eu-west-1), Milan (eu-south-1), and London (eu-west-2, UK). For Frankfurt (eu-central-1) and Stockholm (eu-north-1), the AWS sources contradict each other: the model cards list both regions for Nano 3 30B, while the region overview lists them for Super 120B – check the Bedrock console before deploying. Nemotron 3 Ultra is not listed on Bedrock but is available via Amazon SageMaker JumpStart (self-managed endpoint, region freely selectable).
Azure
West Europe
Microsoft Foundry as a NIM microservice on Managed Compute - the customer selects the Azure region with suitable GPU capacity themselves; unlike the MAI models, this is not a serverless/Global Standard offering.
Google Cloud
Region selectable via self-deployment (e.g. europe-west3 Frankfurt)
Gemini Enterprise Agent Platform (formerly Vertex AI, rebrand April 2026) - Model Garden as a self-deployed model; Google's documentation does not list a managed Nemotron API (Model as a Service)
License & Hosting
License NVIDIA Open Model License (Nemotron 3 Nano/Super, per their Hugging Face model cards, permissive, commercial use allowed); Nemotron 3 Ultra and Nemotron 3.5 Lightning are licensed under OpenMDW-1.1 per their model cards (also permissive, commercial use allowed)
Security Filters Customizable (Nemotron Safety / Guard models available)
Enterprise Support Yes
SLA Available Yes
On-Premise Edge-capable

Benchmarks

Performance comparison with standardized tests

AIME (Ultra, no tools)
88.6
GPQA (Ultra)
87.0
SWE-Bench Verified (Ultra)
70.7
LiveCodeBench v6 (Ultra)
89.0
MMLU-Pro (Ultra)
86.8

NVIDIA has completed its open model family Nemotron 3 with the three reasoning sizes Nano, Super and Ultra, and on August 11, 2026 added Nemotron 3.5 Lightning, a new model in the Nano size class built for long-running agents. With the release of Nemotron 3 Ultra (550B) on 4 June 2026, NVIDIA also offers its first frontier-scale model with open weights. Most models are released under the NVIDIA Open Model License; Nemotron 3 Ultra and Nemotron 3.5 Lightning are licensed under OpenMDW-1.1 per their model cards - both permit free download, modification, and commercial operation. innFactory AI Consulting from Germany advises on GDPR-compliant deployment. This overview reflects the state as of October 2026.

What is NVIDIA Nemotron?

Nemotron is NVIDIA’s family of open models (open weights) built specifically for agentic AI and reasoning. Unlike purely proprietary models, NVIDIA openly publishes weights, training data and training techniques – so the community can run, customize and further train the models for their own purposes. The current Nemotron 3 generation uses an efficient hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture, activating only a fraction of the parameters per request.

Current models at a glance

NVIDIA offers three reasoning sizes that scale with task complexity and hardware:

ModelParameters (active)ContextReleaseLicence
Nemotron 3 Nanoapprox. 31.6B (3.6B)1M15 December 2025NVIDIA Open Model License
Nemotron 3 Super120B (12B)–March 2026 (GTC)NVIDIA Open Model License
Nemotron 3 Ultra550B (55B)up to 1M4 June 2026OpenMDW-1.1
Nemotron 3.5 Lightning30B (3B)1M11 August 2026OpenMDW-1.1
Nemotron-3-Diarization100M–23 September 2026OpenMDW-1.1

Nemotron 3 Nano (30B-A3B)

  • ~31.6B parameters total, ~3.6B active (hybrid Mamba-Transformer MoE)
  • 1M token context window
  • Highly efficient: up to 4x higher throughput than Nemotron Nano 2
  • Self-hostable on moderate hardware (vLLM, SGLang, LM Studio, llama.cpp, Ollama)
  • Release: 15 December 2025

Nemotron 3 Super (120B-A12B)

  • 120B parameters total, ~12B active (hybrid Mamba-Transformer MoE)
  • Strong agentic, reasoning and tool-calling capabilities
  • Optimized for multi-agent systems and high-throughput workloads (e.g. IT ticket automation)
  • Release: GTC, March 2026

Nemotron 3 Ultra (550B-A55B)

  • 550B parameters total, ~55B active (hybrid LatentMoE: Mamba-2 + MoE + Attention, with Multi-Token Prediction)
  • Up to 1M token context for long-context analysis
  • Frontier reasoning over code, mathematics and science
  • Release: 4 June 2026 (Computex)

Nemotron 3.5 Lightning

On August 11, 2026, NVIDIA released Nemotron 3.5 Lightning, a model in the size class of Nemotron 3 Nano:

  • 30B total parameters, ~3B active (hybrid architecture combining Mamba-2, MoE, and attention layers)
  • 1M token context window
  • Per NVIDIA, up to 4x higher output speed and roughly 30% faster agentic task completion
  • Single-GPU deployment (1x DGX Spark/GB10 or 1x H100), also runs on an RTX 5090 workstation
  • Open weights, training data, and recipes under the OpenMDW-1.1 license (BF16 and NVFP4 checkpoints), not the NVIDIA Open Model License
  • Per the model card: MMLU Pro 81.9, GPQA Diamond (no tools) 75.4
  • Available via Hugging Face, ModelScope, OpenRouter, and build.nvidia.com (NIM microservice)

NVIDIA positions Nemotron 3.5 Lightning for long-running autonomous agents and as a sub-agent workhorse in multi-agent systems. It complements the Nano class but does not officially replace Nemotron 3 Nano - per NVIDIA, both remain actively listed.

Nemotron-3-Diarization

Nemotron-3-Diarization (23 September 2026) is a 100M-parameter speaker diarisation model under OpenMDW 1.1: it detects who speaks when (up to eight speakers, streaming and offline, latency from 80 ms to 30.4 s). Per the NVIDIA model card the DER on DIHARD III is 12.73% (predecessor: 19.09%). It adds speaker attribution to transcription models in meetings and call centres, does not identify individuals, and runs on your own infrastructure (NeMo, Transformers or C++ runtime).

Additional models: NVIDIA also maintains Nemotron Nano Omni (multimodal, vision/audio/language, 256K context), Nemotron 3.5 ASR (600M-parameter streaming speech recognition covering 40 language locales, since June 4, 2026), Nemotron 3.5 Content Safety (4B-parameter safety model for multimodal moderation), retriever/RAG models, plus OCR models. The older Llama Nemotron line (Nano 8B, Super 49B, Ultra 253B) is based on Llama 3.1 with 128K context.

Key Strengths

Efficiency through Hybrid MoE

The Mamba-Transformer MoE architecture activates only a fraction of total parameters per token. This significantly reduces inference costs – according to NVIDIA, Nemotron 3 Nano achieves up to 4x higher throughput than the previous generation and reduces the number of reasoning tokens.

Agentic & Tool Use

All Nemotron 3 models are built for agentic workflows: native tool calling, function calls and structured output for multi-agent systems.

Open Weights & Permissive License

The NVIDIA Open Model License is permissive and allows use, modification, distribution and commercial deployment – without an attribution requirement. NVIDIA also publishes training datasets and tools (NeMo Gym, NeMo RL, NeMo Evaluator).

Note: Some individual model cards on Hugging Face list slightly different license names (e.g. OpenMDW). We verify the specific license per model and version as part of our consulting.

EU Availability & GDPR Compliance

Because Nemotron is released as open weights, the cleanest sovereignty path is self-hosting on EU infrastructure – all data stays under your control.

Self-Hosting (recommended for sovereignty)

  • Run on your own hardware or with an EU cloud provider (e.g. in Frankfurt)
  • Full GDPR compliance, no dependency on US APIs
  • Nano can already run on moderate hardware; Ultra requires multiple high-end GPUs (e.g. 8x B200/GB200, 16x H100 or 8x H200)

Managed via Hyperscalers (EU Regions)

  • AWS: Per AWS, Amazon Bedrock offers Nemotron Nano 3 30B and Nemotron 3 Super 120B in-region in Ireland (eu-west-1), Milan (eu-south-1), and London (eu-west-2, UK). For Frankfurt (eu-central-1) and Stockholm (eu-north-1), the AWS sources contradict each other: the model cards list both regions for Nano 3 30B, while the region overview lists them for Super 120B – so check the Bedrock console before deploying. Nemotron 3 Ultra is not listed on Bedrock but is available via Amazon SageMaker JumpStart (self-managed endpoint, region freely selectable, e.g. Frankfurt)
  • Microsoft Foundry: Nemotron Nano and Super run as a NIM microservice on Managed Compute - the customer selects the Azure region with suitable GPU capacity (e.g. West Europe) themselves; unlike the MAI models, this is not a serverless/Global Standard offering
  • Google: Gemini Enterprise Agent Platform (formerly Vertex AI – rebranded in April 2026 at Cloud Next) in the Model Garden - specific EU regions for Nemotron there were not conclusively verified

Directly via NVIDIA

  • build.nvidia.com and NVIDIA NIM microservices for hosted endpoints or containerized self-service deployments
  • Hosted inference also via providers such as Together AI, Fireworks, DeepInfra, OpenRouter, Baseten

For sensitive data we recommend self-hosting in the EU. For a quick start, EU regions of the hyperscalers or NIM containers in your own cloud environment are suitable.

Integration with CompanyGPT

Thanks to open weights, Nemotron integrates well into our GDPR-compliant solution CompanyGPT. This lets you run a powerful reasoning and agent model entirely within your own or an EU-hosted environment – without company data flowing to third countries. innFactory AI Consulting handles selection, deployment and fine-tuning of the right Nemotron model.

Our Recommendation

Nemotron is an open model family for agentic and reasoning applications – and well suited for sovereign, GDPR-compliant deployments thanks to its open weights.

For most enterprises, we recommend:

  • Nemotron 3 Nano or the newer Nemotron 3.5 Lightning for efficient, cost-effective agents and RAG – both self-hostable on moderate to single-GPU hardware; per NVIDIA, Lightning is especially well suited to long-running agentic workloads
  • Nemotron 3 Super as a balanced choice for demanding multi-agent systems
  • Nemotron 3 Ultra for frontier reasoning at maximum requirements (given appropriate GPU infrastructure, available via self-hosting or Amazon SageMaker JumpStart)

We are happy to advise you on model selection, self-hosting in the EU, and integration into existing workflows.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.