Large Language Models

Nemotron 3.5 Lightning

NVIDIA's highest-efficiency open model for long-running agentic AI workloads — built by the Nemotron Coalition, available with NeMo Switchyard for intelligent multi-model routing.

Nemotron 3.5 Lightning is NVIDIA’s August 11, 2026 addition to the Nemotron 3 family — positioned as the highest-efficiency open model for long-running agentic AI workloads. It launches alongside NeMo Switchyard, an open-source library for intelligent multi-model routing.

Where Nemotron 3 Nano targeted smaller deployments, Lightning is optimized for sustained autonomous agent tasks — the kind that run for minutes or hours, not single-turn Q&A.

Key specs

SpecValue
Model familyNemotron 3
PredecessorNemotron 3 Nano
PositionHighest efficiency for long-running agentic workloads
WeightsOpen on Hugging Face
Training dataPublished as much as licensing permits — traceability and auditing supported

NVIDIA has not disclosed detailed parameter counts or full architecture specifications in the launch post. The “highest efficiency in its class” claim is NVIDIA’s marketing position.

NeMo Switchyard

The companion release. NeMo Switchyard is an open-source library that adds smart routing to popular agent frameworks — intelligently directing each request to the most capable and cost-effective model for the job. Enterprise teams can build custom routers based on their specific needs without rewriting applications.

Available on GitHub, with partner platform integrations coming.

Availability

ChannelStatus
Hugging FaceAvailable
build.nvidia.com (NIM)Available
OpenRouterAvailable
Cloud partnersAvailable

How it compares

  • vs Muse Glimmer: Both target agentic workloads with open weights. Glimmer is 30B for consumer hardware; Lightning is built for NVIDIA enterprise GPU infrastructure.
  • vs Muse Spark 1.2: Spark competes at the frontier (AA Index 54, TB 2.1 82.9%). Lightning’s positioning is efficiency over peak capability.
  • vs Inkling: Inkling (Apache 2.0, 975B MoE) is a larger open-weight model with published benchmarks. Lightning’s benchmark story is thinner at launch.

Limitations

  • Sparse launch details. Parameter count, architecture, and benchmark scores not published in the initial post.
  • “Highest efficiency” is unverified. No independent comparison data at launch.
  • Enterprise-focused. Not positioned for consumer or laptop deployment.
  • License terms matter. NVIDIA open model licenses vary — verify the specific terms for your use case.

Also released: the Nemotron-RL-Agentic-Terminal-Pivot dataset — agentic reinforcement learning data used to post-train Lightning for coding agent capabilities.

For the full AI model landscape, see the 2026 LLM guide.

Best for

  • Long-running autonomous coding agents requiring sustained reasoning
  • Enterprise agentic AI workloads on NVIDIA infrastructure
  • Multi-model routing architectures via NeMo Switchyard
  • Fine-tuning for domain-specific agentic tasks

Pros & cons

Strengths

  • Highest-efficiency model in its class for long-running agentic workloads (NVIDIA's claim)
  • Open weights with training data and technique documentation — enables traceability and auditing
  • Ships with NeMo Switchyard — open-source intelligent routing library for multi-model agent architectures
  • Broad ecosystem availability — HuggingFace, NIM, OpenRouter, and cloud partners from day one

Limitations

  • Detailed parameter count and architecture not fully disclosed in launch post
  • Benchmark scores not published in the initial announcement — "highest efficiency" is NVIDIA's marketing claim
  • NVIDIA model licenses have historically included restrictions — check the specific terms
  • Positioned for enterprise and GPU infrastructure — not a consumer/laptop model

Frequently asked questions

When was Nemotron 3.5 Lightning released?

August 11, 2026, alongside NeMo Switchyard. It follows Nemotron 3 Nano in the Nemotron 3 model family.

What is NeMo Switchyard?

An open-source library for smart routing inside popular agent tools. It intelligently directs each request to the most capable model for the job — allowing developers to use multiple models without rewriting applications.

What is the Nemotron Coalition?

A group of contributors that help advance the model through inference software and datasets. NVIDIA publishes training data and techniques as much as licensing permits for traceability.

How does Nemotron 3.5 Lightning relate to Nemotron 3 Nano?

Both are part of the Nemotron 3 model family. Nano targets smaller/faster deployments; Lightning targets highest efficiency for sustained agentic workloads.

Sources

Last updated: 2026-08-11 · Specs and pricing change fast — verify on the vendor's site before relying on them.