Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to slash agentic AI costs by 66%

Reviewed byNidhi Govil

15 Sources

Share

Nvidia launched Nemotron 3.5 Lightning, a 30-billion-parameter open-source AI model delivering 4x faster token generation, alongside NeMo Switchyard, an open-source routing library that cuts agentic task costs to roughly one-third of frontier models. The dual release addresses rising AI bills while maintaining frontier-level performance across specialized agentic tasks.

Nvidia Tackles Rising AI Costs With Dual Open-Source Release

Nvidia released Nemotron 3.5 Lightning and NeMo Switchyard on August 11, marking the chip giant's first open-source AI model launch since CEO Jensen Huang publicly defended open models in late July

2

. The 30-billion-parameter mixture-of-experts model delivers up to 4x faster token generation and 30% faster task completion compared to open models in its class

3

. Paired with the NeMo Switchyard open-source routing library, internal benchmarks show the combination maintains frontier-level task completion while reducing costs to roughly one-third of running Opus 4.8 alone

3

. This dual release addresses the core challenge enterprises face with agentic AI: balancing performance against escalating token costs without building custom routing infrastructure

5

.

Source: VentureBeat

Source: VentureBeat

Nemotron 3.5 Lightning Powers Specialized Agentic Tasks

The customizable open model targets high-volume specialized agentic tasks within larger multi-agent systems, including code review, tool use, security alert monitoring and answering billing questions

1

. Because Nemotron 3.5 Lightning uses open weights, developers can fine-tune it with their own examples to match specific workflows, writing styles or coding conventions

3

. AI leaders across industries are already customizing the model: CrowdStrike for cybersecurity, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review

4

. The model runs locally on RTX PCs, DGX Spark, Jetson devices and scales to RTX PRO workstations, data centers and cloud environments

3

. Nvidia collaborated with vLLM, Ollama, llama.cpp and LM Studio to optimize local deployment, while Unsloth provides day-one support with quantized models through Unsloth Studio

3

.

Source: Geeky Gadgets

Source: Geeky Gadgets

NeMo Switchyard Automates Model Selection Across Workflows

The open-source routing library automatically directs each step of an agent workflow to the most suitable model based on accuracy, speed and cost requirements

3

. Unlike fixed routing approaches, NeMo Switchyard responds to shifting agent states as tools return results or errors emerge, adjusting model selection dynamically rather than using static task categories

5

. The router evaluates model verbosity—predicting how many tokens a given model produces—and steers work toward cheaper options before calls are made

5

. Integration partners split into two groups: agent frameworks like Cognition, LangChain and Nous Research that call Switchyard directly, and LLM gateways including Kong, LiteLLM and OpenRouter that built native support

5

. Early testing shows dramatic results. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, with only a 6% accuracy tradeoff

5

.

Source: SiliconANGLE

Source: SiliconANGLE

Strategic Positioning Amid Open-Source Competition

The release lands during the busiest open-weight stretch the industry has seen in months, with Alibaba, Moonshot, Zhipu and DeepSeek shipping competitive Chinese models since spring, several approaching frontier performance while undercutting US labs on size or price

5

. Meta added pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer

5

. Jensen Huang has argued that free AI drives hardware sales, telling Axios that "free AI should be great for chips" since models still require GPUs regardless of licensing

2

. Nvidia remains among the few major US firms releasing open-source models, which have drawn increased attention as AI bills balloon and cheap Chinese models near capabilities of top systems from Anthropic and OpenAI

1

. The company also formed a coalition last month with other firms to develop tools for AI safety and cybersecurity, signing an open letter with Microsoft backing open-weight models to prevent innovation from drifting overseas

1

.

What's Next: Nemotron 4 Development Continues

Separately, Nvidia is developing Nemotron 4, a new AI model family targeting top open-source models globally, according to The Information

1

. The largest Nemotron 4 model is expected to have at least 1 trillion parameters, with employees suggesting readiness as early as late fall, though Nvidia has not set a release date or completed final training

1

. Watch for how enterprises adopt model routing strategies as intelligent agents shift from experimental to production workloads, particularly whether cost savings materialize at scale beyond benchmark tests. The competitive landscape will likely intensify as Chinese labs continue releasing capable open models and US firms respond with their own open alternatives, potentially reshaping how organizations balance performance, cost and control in agentic AI deployments.

Today's Top Stories

Ā© 2026 TheOutpost.AI All rights reserved