Nvidia Releases Nemotron 3.5 Lightning, an Open Model Built to Run Inside Agent Systems

Nvidia released Nemotron 3.5 Lightning this week, a 30-billion-parameter open-weight model designed for a specific job: handling the high-volume, repetitive steps that keep AI agents running for hours or days at a time. The release comes paired with NeMo Switchyard, an open-source routing library meant to decide, step by step, which model should handle which part of an agent’s workload. Together they outline a concrete answer to a question the AI industry has been circling for months: whether agentic systems should rely on one large model for everything, or a coordinated mix of models each doing what they’re best at.

A Small Model Tuned for Execution Speed

Nemotron 3.5 Lightning extends the Nemotron 3 family, which began with Nemotron 3 Nano earlier this year. It uses a hybrid mixture-of-experts design with 30 billion total parameters but only 3 billion active per token, interleaving Mamba-2 layers with MoE and attention layers rather than relying on a single architecture type. That combination, along with speculative decoding and multi-token prediction, is what lets Nvidia claim up to four times faster output speed and roughly 30 percent faster task completion compared with similarly sized models, according to internal benchmarks run on the company’s PinchBench suite. The model supports up to a one-million-token context window and was trained on more than 20 trillion tokens spanning code, math, science, and general knowledge.

None of those numbers matter much in isolation. What makes Lightning notable is what Nvidia built it to avoid: the latency and cost of sending every small task inside an agent workflow, a tool call, a formatting check, a routine classification, to a large frontier reasoning model that was never designed for that kind of volume.

Why Agent Systems Are Starting to Look Like Model Ensembles

Nvidia frames Lightning as one piece of a broader shift in how agentic AI gets built. Instead of a single model handling planning, reasoning, and execution end to end, the company describes modern always-on agents as systems of models, where a larger frontier model such as Nemotron 3 Ultra or GPT-5.6 handles orchestration and complex planning, while smaller specialized models like Lightning execute the high-volume steps underneath it: tool use, code review, security alert monitoring, or answering routine billing questions.

The logic is straightforward. Frontier reasoning models are expensive and comparatively slow, which makes sense for hard, infrequent decisions but not for thousands of repetitive calls inside a single agent session. A smaller model tuned specifically for that execution layer can run those steps faster and at a fraction of the cost, as long as something in the system knows how to send the right task to the right model.

That “something” is where NeMo Switchyard comes in. Nvidia describes it as an open-source routing library that automatically directs each step of an agent’s workflow to the most suitable model available, open, proprietary, or built by Nvidia, without requiring developers to rewrite their applications around a single provider. Nvidia’s internal testing claims Switchyard preserves frontier-level accuracy while cutting task completion cost to close to a third of running everything through Opus 4.8 alone.

Early partner numbers, while still self-reported and workload-specific, point in a consistent direction. LangChain reported a 74 percent cost reduction across a set of multi-turn agent tasks by routing only a small share of calls to a frontier model, accepting a modest accuracy tradeoff. Ramp said it matched frontier-model performance on its software engineering benchmark while cutting cost by 58 percent and runtime by a third. Classmethod and Cognition each reported cost reductions in the high twenties percent range in internal testing, and Boomi said its evaluation across five routing capabilities reached full accuracy on domain classification while shifting most traffic to a faster, fine-tuned model.

Where the Model Is Being Deployed

Rather than deploying Lightning off the shelf, most of the named early adopters are fine-tuning it for narrow domains. CrowdStrike is applying it to cybersecurity workloads, Harvey is using it alongside its Trajectory product for legal work, and CodeRabbit has paired it with Baseten to route code reviews. Fastino Labs reported strong accuracy after customizing the model for software development, finance, and healthcare tasks, and Lila Sciences is using it to support agentic reasoning across physical and life sciences research.

That pattern, a compact open base model fine-tuned by specialized vendors for individual verticals, says more about Nvidia’s strategy than the headline benchmarks do. The company appears focused on becoming the default execution layer underneath a growing number of narrow, purpose-built agents, regardless of which frontier model sits on top of the stack doing the planning, rather than competing head-on with chatbot-facing frontier labs like OpenAI, Anthropic, or Google.

Because the model is open, organizations can also run it locally, on RTX PCs, DGX Spark, DGX Station, or Jetson devices, in addition to data centers and the cloud, which matters for teams that need to keep specialized agentic workloads on internal hardware for privacy or latency reasons. Nvidia also published a reinforcement learning dataset, Nemotron-RL-Agentic-Terminal-Pivot, used to post-train the model for coding agent capabilities, continuing its practice of releasing enough of the training pipeline for outside teams to audit or retrain the model themselves.

The Business Logic Underneath an Open Release

Nvidia’s willingness to give away model weights fits a business model that depends on inference happening somewhere, ideally on Nvidia hardware. A cheaper, faster, more efficient model ecosystem lowers the barrier to running more agents in more places, and every additional agent deployed anywhere adds to inference demand regardless of which company built the model doing the reasoning on top. CEO Jensen Huang made that connection explicit in public comments over the summer, arguing that open models are ultimately good for chip demand rather than a threat to it, and Lightning is the company’s first fully open model release since he made that argument publicly.

Nvidia positions Lightning as a component sitting underneath larger systems run by Nemotron 3 Ultra or comparable frontier models, handling volume rather than the hardest reasoning steps, and stops short of framing it as a replacement for those frontier models. Whether this layered, multi-model approach becomes the standard architecture for enterprise agents, or whether frontier labs manage to make routing unnecessary by making their models cheap and fast enough to handle everything directly, is likely to be one of the more consequential competitive questions in agentic AI over the coming year.

Leave a Reply

F in WA @

Discover more from LiteRouter Academia

Subscribe now to keep reading and get access to the full archive.

Continue reading