16 Sources
[1]
Nvidia is developing Nemotron 4 open-source models, The Information reports
Aug 11 (Reuters) - Nvidia (NVDA.O), opens new tab is developing a new AI model family, Nemotron 4, with the goal of rivaling top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major U.S. firms to release
[2]
Nvidia unveils first open-source AI model since CEO Jensen Huang entered the chat
In late July, Nvidia CEO Jensen Huang posted on X for the first time to defend open-source models in artificial intelligence, inserting himself into a debate that was raging across the industry. Less than three weeks later, Nvidia is releasing Nemotron 3.5 Lightning, which the company says is
[3]
NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents
The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools
[4]
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
The new lightweight open model and routing library delivers greater control over AI, data and workflows across edge devices, PCs, workstations, data centers and the cloud. As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs
[5]
Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests
Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow
[6]
Nvidia releases Nemotron 3.5 Lightning open-source AI model
Nvidia $NVDA released Nemotron 3.5 Lightning on Tuesday, a free, open-source AI model the company says can run on a single graphics processing unit on a PC and is designed for autonomous agent workloads. The new model is a 30-billion-parameter mixture-of-experts model built for specialized tasks
[7]
Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options
Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard. As enterprises find themselves drowning in artificial intelligence model options, the question is no longer
[8]
Nvidia Nemotron 3.5 Lightning Runs Locally on a Single GPU
Nvidia has released Nemotron 3.5 Lightning, a customizable open-weight language model designed for local AI agents and high-volume automated workloads. It has 30 billion parameters in total but activates approximately 3 billion parameters per token, helping reduce inference requirements compared
[9]
Nvidia is developing Nemotron 4 open-source models: The Information
The chip giant is among the few major US firms to release open-source models, which have drawn more attention this year as AI bills balloon and cheap Chinese models near the capabilities of top systems from leading American labs Anthropic and OpenAI. Nvidia is developing a new AI model family,
[10]
NVIDIA Nemotron 3.5 Lightning Delivers 4X AI Throughput
NVIDIA's Nemotron 3.5 Lightning is an AI model designed for execution-layer tasks, emphasizing speed and efficiency in workflows that require high throughput. According to Sam Witteveen, the model employs a hybrid member transformer architecture with an active mixture of experts (MoE), using up to
[11]
NVIDIA's Nemotron 3.5 Lightning Accelerates Token Generation By 4x But Agentic Tasks Only Speed Up By 30%, As Orchestration Remains The Real Bottleneck
The race for open-weight models in the US is heating up even as China continues to own this segment, as evidenced by NVIDIA's release of its Nemotron 3.5 Lightning AI model, which has dropped just hours after Meta introduced its Muse Glimmer model on Monday. NVIDIA's Nemotron 3.5 Lightning is
[12]
NVIDIA introduces Nemotron 3.5 Lightning 30B open AI model and NeMo Switchyard for agentic AI
NVIDIA has introduced Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts (MoE) open model designed for high-volume execution tasks in long-running AI agents. The company has also released NeMo Switchyard, an open-source model routing library that directs requests to different AI
[13]
Nvidia building 1-trillion-parameter Nemotron 4 to rival open AI models
Nvidia is developing a new AI model family, Nemotron 4, with the goal of challenging top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major U.S. firms to release open-source models, which have drawn more
[14]
Nvidia readies Nemotron 4, a giant model to bolster its open AI lineup
According to The Information, Nvidia still has to complete the final training phase for Nemotron 4 and has not set an official launch date. Employees involved in the project nevertheless believe the model could be ready as soon as late fall. The group confirms it is working on this new generation
[15]
Nvidia readies Nemotron 4, a giant model to strengthen its open AI offering
According to The Information, Nvidia still needs to complete the final training phase for Nemotron 4 and has not set any official release date. Employees involved in the project nonetheless believe the model could be ready as early as late fall. The company confirms it is working on this new
[16]
Nvidia is developing Nemotron 4 open-source models, The Information reports
Aug 11 (Reuters) - Nvidia is developing a new AI model family, Nemotron 4, with the goal of rivaling top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major U.S. firms to release open-source models, which
Share
Copy Link
Nvidia launched Nemotron 3.5 Lightning, a 30-billion-parameter open-source AI model delivering 4x faster token generation, alongside NeMo Switchyard, an open-source routing library that cuts agentic task costs to roughly one-third of frontier models. The dual release addresses rising AI bills while maintaining frontier-level performance across specialized agentic tasks.
Nvidia released Nemotron 3.5 Lightning and NeMo Switchyard on August 11, marking the chip giant's first open-source AI model launch since CEO Jensen Huang publicly defended open models in late July
2
. The 30-billion-parameter mixture-of-experts model delivers up to 4x faster token generation and 30% faster task completion compared to open models in its class3
. Paired with the NeMo Switchyard open-source routing library, internal benchmarks show the combination maintains frontier-level task completion while reducing costs to roughly one-third of running Opus 4.8 alone3
. This dual release addresses the core challenge enterprises face with agentic AI: balancing performance against escalating token costs without building custom routing infrastructure5
.
Source: VentureBeat
The customizable open model targets high-volume specialized agentic tasks within larger multi-agent systems, including code review, tool use, security alert monitoring and answering billing questions
1
. Because Nemotron 3.5 Lightning uses open weights, developers can fine-tune it with their own examples to match specific workflows, writing styles or coding conventions3
. AI leaders across industries are already customizing the model: CrowdStrike for cybersecurity, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review4
. The model runs locally on RTX PCs, DGX Spark, Jetson devices and scales to RTX PRO workstations, data centers and cloud environments3
. Nvidia collaborated with vLLM, Ollama, llama.cpp and LM Studio to optimize local deployment, while Unsloth provides day-one support with quantized models through Unsloth Studio3
.
Source: Geeky Gadgets
The open-source routing library automatically directs each step of an agent workflow to the most suitable model based on accuracy, speed and cost requirements
3
. Unlike fixed routing approaches, NeMo Switchyard responds to shifting agent states as tools return results or errors emerge, adjusting model selection dynamically rather than using static task categories5
. The router evaluates model verbosity—predicting how many tokens a given model produces—and steers work toward cheaper options before calls are made5
. Integration partners split into two groups: agent frameworks like Cognition, LangChain and Nous Research that call Switchyard directly, and LLM gateways including Kong, LiteLLM and OpenRouter that built native support5
. Early testing shows dramatic results. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, with only a 6% accuracy tradeoff5
.Related Stories
The release lands during the busiest open-weight stretch the industry has seen in months, with Alibaba, Moonshot, Zhipu and DeepSeek shipping competitive Chinese models since spring, several approaching frontier performance while undercutting US labs on size or price
5
. Meta added pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer5
. Jensen Huang has argued that free AI drives hardware sales, telling Axios that "free AI should be great for chips" since models still require GPUs regardless of licensing2
. Nvidia remains among the few major US firms releasing open-source models, which have drawn increased attention as AI bills balloon and cheap Chinese models near capabilities of top systems from Anthropic and OpenAI1
. The company also formed a coalition last month with other firms to develop tools for AI safety and cybersecurity, signing an open letter with Microsoft backing open-weight models to prevent innovation from drifting overseas1
.
Source: SiliconANGLE
Separately, Nvidia is developing Nemotron 4, a new AI model family targeting top open-source models globally, according to The Information
1
. The largest Nemotron 4 model is expected to have at least 1 trillion parameters, with employees suggesting readiness as early as late fall, though Nvidia has not set a release date or completed final training1
. Watch for how enterprises adopt model routing strategies as intelligent agents shift from experimental to production workloads, particularly whether cost savings materialize at scale beyond benchmark tests. The competitive landscape will likely intensify as Chinese labs continue releasing capable open models and US firms respond with their own open alternatives, potentially reshaping how organizations balance performance, cost and control in agentic AI deployments.Summarized by
Navi
[4]
11 Mar 2026•Technology

15 Dec 2025•Technology

19 Aug 2025•Technology

1
Technology

2
Technology

3
Policy and Regulation
