Nemotron 3.5 Lightning is NVIDIA’s August 11, 2026 addition to the Nemotron 3 family — positioned as the highest-efficiency open model for long-running agentic AI workloads. It launches alongside NeMo Switchyard, an open-source library for intelligent multi-model routing.
Where Nemotron 3 Nano targeted smaller deployments, Lightning is optimized for sustained autonomous agent tasks — the kind that run for minutes or hours, not single-turn Q&A.
Key specs
| Spec | Value |
|---|---|
| Model family | Nemotron 3 |
| Predecessor | Nemotron 3 Nano |
| Position | Highest efficiency for long-running agentic workloads |
| Weights | Open on Hugging Face |
| Training data | Published as much as licensing permits — traceability and auditing supported |
NVIDIA has not disclosed detailed parameter counts or full architecture specifications in the launch post. The “highest efficiency in its class” claim is NVIDIA’s marketing position.
NeMo Switchyard
The companion release. NeMo Switchyard is an open-source library that adds smart routing to popular agent frameworks — intelligently directing each request to the most capable and cost-effective model for the job. Enterprise teams can build custom routers based on their specific needs without rewriting applications.
Available on GitHub, with partner platform integrations coming.
Availability
| Channel | Status |
|---|---|
| Hugging Face | Available |
| build.nvidia.com (NIM) | Available |
| OpenRouter | Available |
| Cloud partners | Available |
How it compares
- vs Muse Glimmer: Both target agentic workloads with open weights. Glimmer is 30B for consumer hardware; Lightning is built for NVIDIA enterprise GPU infrastructure.
- vs Muse Spark 1.2: Spark competes at the frontier (AA Index 54, TB 2.1 82.9%). Lightning’s positioning is efficiency over peak capability.
- vs Inkling: Inkling (Apache 2.0, 975B MoE) is a larger open-weight model with published benchmarks. Lightning’s benchmark story is thinner at launch.
Limitations
- Sparse launch details. Parameter count, architecture, and benchmark scores not published in the initial post.
- “Highest efficiency” is unverified. No independent comparison data at launch.
- Enterprise-focused. Not positioned for consumer or laptop deployment.
- License terms matter. NVIDIA open model licenses vary — verify the specific terms for your use case.
Also released: the Nemotron-RL-Agentic-Terminal-Pivot dataset — agentic reinforcement learning data used to post-train Lightning for coding agent capabilities.
For the full AI model landscape, see the 2026 LLM guide.
Best for
- Long-running autonomous coding agents requiring sustained reasoning
- Enterprise agentic AI workloads on NVIDIA infrastructure
- Multi-model routing architectures via NeMo Switchyard
- Fine-tuning for domain-specific agentic tasks
Pros & cons
Strengths
- Highest-efficiency model in its class for long-running agentic workloads (NVIDIA's claim)
- Open weights with training data and technique documentation — enables traceability and auditing
- Ships with NeMo Switchyard — open-source intelligent routing library for multi-model agent architectures
- Broad ecosystem availability — HuggingFace, NIM, OpenRouter, and cloud partners from day one
Limitations
- Detailed parameter count and architecture not fully disclosed in launch post
- Benchmark scores not published in the initial announcement — "highest efficiency" is NVIDIA's marketing claim
- NVIDIA model licenses have historically included restrictions — check the specific terms
- Positioned for enterprise and GPU infrastructure — not a consumer/laptop model
Frequently asked questions
When was Nemotron 3.5 Lightning released?
August 11, 2026, alongside NeMo Switchyard. It follows Nemotron 3 Nano in the Nemotron 3 model family.
What is NeMo Switchyard?
An open-source library for smart routing inside popular agent tools. It intelligently directs each request to the most capable model for the job — allowing developers to use multiple models without rewriting applications.
What is the Nemotron Coalition?
A group of contributors that help advance the model through inference software and datasets. NVIDIA publishes training data and techniques as much as licensing permits for traceability.
How does Nemotron 3.5 Lightning relate to Nemotron 3 Nano?
Both are part of the Nemotron 3 model family. Nano targets smaller/faster deployments; Lightning targets highest efficiency for sustained agentic workloads.