15 Sources
[1]
Nvidia is developing Nemotron 4 open-source models, The Information reports
Aug 11 (Reuters) - Nvidia (NVDA.O), opens new tab is developing a new AI model family, Nemotron 4, with the goal of rivaling top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major U.S. firms to release open-source models, which have drawn more attention this year as AI bills balloon and cheap Chinese models near the capabilities of top systems from leading American labs Anthropic and OpenAI. A spate of recently disclosed hacks involving autonomous ā AI agents has added to the attention, especially because open models do not have curbs on cybersecurity use. The largest Nemotron 4 model is expected to have at least 1 trillion parameters, according to multiple employees working on the project, The Information reported. Nvidia has not set a release date for Nemotron 4 and has yet to complete final training, though employees said the model could be ready as early as late fall, according to the report. The company ā did not immediately respond to a Reuters request for comment on the report. Nvidia last month formed a coalition with other companies to develop and share tools for AI safety and cybersecurity. It also signed an open letter with tech heavyweights such as Microsoft (MSFT.O), opens new tab ā backing open-weight models so that innovation does not drift overseas. Separately on Tuesday, the chip firm unveiled Nemotron 3.5 Lightning, an addition to its offerings aimed at code review, ā tool use, security alert monitoring, answering billing questions and other tasks. It also released NeMo Switchyard, an open-source model-routing library designed to automatically direct AI tasks ā to the most suitable models. Late last year, the chip giant unveiled the third generation of its family of open-source models as offerings from Chinese AI labs proliferated. Reporting by Anhata Rooprai in Bengaluru; Editing by Pooja Desai Our Standards: The Thomson Reuters Trust Principles., opens new tab
[2]
Nvidia unveils first open-source AI model since CEO Jensen Huang entered the chat
In late July, Nvidia CEO Jensen Huang posted on X for the first time to defend open-source models in artificial intelligence, inserting himself into a debate that was raging across the industry. Less than three weeks later, Nvidia is releasing Nemotron 3.5 Lightning, which the company says is "lightweight" and can run on a single graphics processing unit on a PC. It's Nvidia's first open-source model since Huang joined most of his tech peers in urging the U.S. government to support open models while "avoiding premature restrictions" that could push innovation overseas. The new Nemotron offering is free for companies to download, use and modify without getting permission or paying Nvidia. For Nvidia, open-source AI is a boon for chip sales, because the models still need to run on GPUs, and the lower prices can serve to boost usage over proprietary models from the likes of OpenAI and Anthropic. "Free AI should be great for hardware," Huang told Axios in an interview last month. "Free AI should be great for chips." Huang jumped headfirst into a debate that had sprung up in Washington following the announcement of Kimi K3, a model developed by China's Moonshot AI that narrowed the gap with the most powerful American models. Politicians worried that Kimi K3 was potentially troublesome for national security, and that it represented intellectual property theft via a technique called distillation, which involves the use of answers from an advanced AI model's service to train a lighter model.
[3]
NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents
The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. That includes NVIDIA's latest open models, software and developer tools, plus the accelerated computing, libraries and educational resources that help users get started. It's shaping up to be a big month for agents. Follow along for the latest developments in this special-edition NVIDIA Local AI blog series, with new updates added over the coming weeks. Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook -- and stay informed by subscribing to the RTX AI PC newsletter. Follow NVIDIA Workstation on LinkedIn and X. Tuesday, Aug. 11, 6:00 a.m. PT š NVIDIA Introduces Nemotron 3.5 Lightning for Fast, Specialized Agentic Tasks Today, NVIDIA expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts (MoE) model for always-on agents. Nemotron 3.5 Lightning delivers up to 4x faster token generation and 30% faster time to completion compared to open models in its class. And because Nemotron 3.5 Lightning is open weights, AI enthusiasts and developers can fine-tune it with their own examples to better match specific tasks, interests and workflows. For example, they could train the model to: * Write in a preferred style: Follow established tones, formats and terminology when writing emails, reports or other documents. * Learn a specialty: Better understand the language and common tasks associated with areas such as photography, gaming or 3D design. * Code a certain way: Follow preferred coding conventions, frameworks and testing approaches when writing, reviewing or refactoring code. Paired with access to apps, files and other tools, these fine-tuned models can power more personalized local agentic AI experiences -- from an assistant that helps manage email and calendars, to a smart-home agent that handles everyday routines, to a coding companion that works alongside developers on a local codebase. NVIDIA collaborated with vLLM, Ollama, llama.cpp and LM Studio to provide the best local deployment experience for Nemotron 3.5 Lightning models -- offering developers choice of NVFP4 and GGUF format of models. Unsloth also provides day-one support with optimized and quantized models for efficient local deployment via Unsloth Studio. Nemotron 3.5 Lightning runs locally on NVIDIA RTX PCs, NVIDIA DGX Spark and OEM GB10 systems, and NVIDIA Jetson, and scales up to RTX PRO workstations, NVIDIA DGX Station and GB300 deskside systems, data centers and cloud environments. With NVIDIA Blackwell systems available from Acer, ASUS, Dell Technologies, Exxact, GIGABYTE, HP, Lenovo, MSI and Supermicro, users can choose from a wide range of devices and form factors to fit their needs. As generative AI adoption grows, enterprises are looking for ways to keep rising token costs in check without sacrificing access to frontier intelligence. NVIDIA NeMo Switchyard, an open source routing library, automatically directs each step of an agent workflow to the best-fit model based on accuracy, speed and cost. It also gives developers the flexibility to work across models and providers for different tasks. Internal benchmarks show that NeMo Switchyard, by routing each step across a system of models, helped maintain frontier-level task completion while reducing benchmark completion cost to roughly one-third of Opus 4.8 alone. NeMo Switchyard is available on GitHub. Visit the Nemotron 3.5 Lightning, NeMo Switchyard and Jetson AI technical blogs to get started. And to build at the edge, start with Jetson AI Lab tutorials and discover real-world Jetson projects. Nemotron 3.5 Lightning is also available through OpenRouter, on build.nvidia.com as an NVIDIA NIM microservice, and through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on GitHub.
[4]
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
The new lightweight open model and routing library delivers greater control over AI, data and workflows across edge devices, PCs, workstations, data centers and the cloud. As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it's deployed and evolves. Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads. This release follows Nemotron 3 Nano and reflects NVIDIA's commitment to continually improving open models for greater accuracy and speed. Built for specialized tasks within larger multi-agent systems, Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, helps create smarter and more efficient agentic applications. Also, NVIDIA is releasing NeMo Switchyard, an open source library for smart routing inside popular agent tools. Enterprises can use it to build a router based on their specific needs. When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job without requiring developers to rewrite their applications. Together, Nemotron 3.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates -- across PCs, workstations, data centers and the cloud. Always-On Agents Need a System of Models Modern agentic systems -- always-on agents -- increasingly operate as systems of models, or model ensembles, with different models specialized for different tasks. NVIDIA Nemotron open models are designed for this architecture. A frontier reasoning model such as Nemotron 3 Ultra or GPT-5.6 may plan and orchestrate a workflow, while smaller specialized models like Nemotron 3.5 Lightning can perform targeted tasks such as code review, tool use, security alert monitoring and answering billing questions. Powering High-Volume Specialized Tasks With Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is a fully customizable open model built for high-volume tasks powering always-on agents. It was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software and datasets to help advance the model. The model delivers up to 4x faster output speed, leading to 30% faster agentic task completion compared with other models in its class. And because it's open and customizable, Nemotron 3.5 Lightning can be easily post-trained with NVIDIA NeMo on an organization's own domain data, tools and workflows to improve accuracy for specialized tasks. AI leaders across industries are customizing Nemotron 3.5 Lightning for their workloads, including CrowdStrike for cybersecurity, Harvey with Trajectory for legal services and CodeRabbit with Baseten for code review, helping improve accuracy for domain-specific agentic tasks. Additionally, Lila Sciences is helping to improve reasoning capabilities for agentic tasks across physical and life sciences, and Fastino Labs customized the model and is seeing leading accuracies for software development, finance and healthcare workloads. Nemotron 3.5 Lightning also gives organizations control over privacy and deployment. It can run on local AI systems -- including NVIDIA RTX PCs, NVIDIA DGX Spark, NVIDIA DGX Station and NVIDIA Jetson -- to help users maximize existing infrastructure investments, or scale across edge AI devices, NVIDIA RTX PRO workstations, data centers and cloud environments for enterprise use cases. And Nemotron 3.5 Lightning can run locally or on premises for high-volume, specialized tasks that require fast responses. Also, as with every Nemotron launch, NVIDIA publishes as much of the training data and techniques as licensing permits, which allows for traceability, auditing and training of other models. Alongside Lightning, NVIDIA is releasing Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset used to post-train it for coding agent capabilities. More Efficient AI Apps With Model Routing Some models are better for coding, some for reasoning, some for lightweight tasks and some are optimized to run locally for greater privacy and efficiency. If customers rely on one default model, they might either overspend or lose quality; if they manage routing manually, it becomes integration work that can slow down a deployment. NVIDIA NeMo Switchyard is an open source model routing library for AI agents. The technology routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on specific needs. Agent application developers can tune or modify the router with different routing algorithms to match their priorities, such as quality, latency and cost requirements. In a system of models, enterprises can create powerful AI agents with improved tokenomics. Internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone. NVIDIA is working with partners across the AI ecosystem to bring intelligent model routing into the tools and platforms developers already use. * Boomi: Evaluated Switchyard across five routing capabilities, achieving 100% domain-routing accuracy, sending 59% of traffic to a 5x faster fine-tuned model and reducing later-turn latency by 21%. * Cadence: Improved efficiency by 9.9% by using the ChipStack AI Super Agent for a formal verification use case. * Classmethod: Is running opencode and Fireworks workloads using NeMo Switchyard internally, with initial testing showing a 27% cost reduction while maintaining quality. * Cognition: Integrated the NVIDIA NeMo Switchyard staged router into Devin Desktop for NVIDIA internal use, achieving near-frontier performance on FrontierCode Main while reducing mean cost by 28% relative to routing all requests to a single underlying frontier model. * Kong: Delivers routing with NeMo Switchyard natively through Kong AI Gateway. * LangChain: With NeMo Switchyard, achieved 74% lower cost in 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff. * LiteLLM: Is adding NeMo Switchyard as a plug-in into its proxy layer so developers can access these benefits without changing their existing stack. * Nous Research: Integrated NeMo Switchyard into Hermes to provide developers with an easy-to-configure routing system to improve agent efficiency. * Ramp: Used NeMo Switchyard to match a frontier model's performance while cutting costs by 58% and runtime by 33% in Ramp SWE-Bench. * Siemens: Is benchmarking to improve efficiency in its Fuse EDA AI Agent. Nemotron 3.5 Lightning is available on Hugging Face, ModelScope, OpenRouter and build.nvidia.com as an NVIDIA NIM microservice as well as through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on GitHub and coming to partner platforms soon.
[5]
Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests
Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow changes. Nvidia is proposing a fix that touches both ends of that problem at once. The company is out on Tuesday with Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for high-volume, specialized agent tasks, alongside NeMo Switchyard, an open-source library that routes each step of an agent workflow to whichever model fits it best. The headline numbers: According to Nvidia, Lightning delivers up to 4x faster output than comparable models in its class, completing agentic tasks roughly 30% faster than Qwen3.6-35B at matching accuracy. Paired through Switchyard, Nvidia says the combination holds frontier-level task completion while cutting benchmark costs to roughly a third of running Opus 4.8 alone. The timing puts Nvidia in the middle of the busiest open-weight stretch the industry has seen in months. Alibaba, Moonshot, Zhipu and DeepSeek have all shipped competitive open models out of China since the spring, several landing at or near frontier performance while undercutting US labs on size or price. Meta added to that pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer. Open weights have gone from a differentiator to table stakes in a matter of months, and Nvidia's release lands squarely inside that shift rather than ahead of it. The pairing is the point. A model alone doesn't solve the cost problem, and a router alone has nothing efficient to route to. Nvidia is betting that open source, applied at both the model layer and the routing layer, is what actually moves the cost needle on agentic AI, not a single cheaper model and not a smarter router bolted onto someone else's stack. Switchyard's real rivals aren't other open models -- they're Not Diamond, which already powers OpenRouter's Auto mode, and RouteLLM, the open-source framework from UC Berkeley and LMSYS. Neither ships its own model. Nvidia's bet is that owning both sides of the decision, under one open license, is what a router-only or model-only competitor can't match. "That is the power of a system of models, matching the right model to each step of the workflow," Kari Briski, vice president of generative AI at Nvidia, said in a briefing. How the router actually changes the workflow Model routing isn't a new category. OpenRouter, LiteLLM and a handful of standalone routing startups already let developers point traffic across multiple providers. Switchyard plugs into several of them rather than replacing them outright. The core problem Switchyard solves is that the right model changes as an agent moves through a task. An agent's state shifts as tools return results, errors show up, or a step turns out to be routine rather than complex, and a fixed model choice can't adapt to any of that. Briski described routing strategies that respond to that shifting state rather than a static task category. "It has many types of routing strategies," Briski said. "You can have a random router, which is not that great, or you can have an agent state route or a classifier route. Depending on your routing strategy, it wants to choose the best model. In some cases you want to go with a model like Lightning for really efficient tasks, and the router will actually choose Lightning if it's set up in your pool of models." Cost enters the routing decision directly, not as an afterthought. In response to a question from VentureBeat, Briski said Switchyard can evaluate model verbosity, meaning how many tokens a given model tends to produce for a task, and use that prediction to steer work toward the cheaper option before the call is made. The part that keeps this from becoming its own integration project is where Switchyard sits. Nvidia split its partners into two groups: agent frameworks that call Switchyard directly, including Cognition, LangChain and Nous Research, and LLM gateways that have built Switchyard support into their own products, including Kong, LiteLLM and OpenRouter. Kong ships Switchyard natively inside Kong AI Gateway. Briski pointed to that same list of gateway partners when describing how the library fits into the existing routing ecosystem. "We are an ecosystem lover, and we want to make sure that we are integrated," Briski said. "We've partnered with OpenRouter, LiteLLM and Kong, and they've already integrated our routing algorithm, so you can pick it up right where you're already using the best tools." Nvidia shared results from nine companies testing Switchyard, several with specific figures attached. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, at a 6% accuracy tradeoff. Ramp said it matched a frontier model's performance on Ramp SWE-Bench while cutting costs 58% and runtime 33%. Cognition integrated Switchyard's staged router into Devin Desktop for internal use and reported near-frontier performance on FrontierCode Main while cutting mean cost 28% relative to routing everything to a single frontier model. Lightning's architecture and performance gains Nemotron 3.5 Lightning is a standalone open model in its own right, built for high-volume, specialized agent tasks rather than general-purpose use. It extends the hybrid Mamba-Transformer, latent mixture-of-experts architecture Nvidia introduced with the Nemotron 3 family in December 2025, the same line behind Nemotron 3 Super, which Nvidia uses as Lightning's own baseline in its post-training comparisons. Positioned within a routing setup like Switchyard, it's built to sit at the fast, cheap end of the decision rather than the frontier end, but it runs and ships independent of any router. According to the Artificial Analysis Intelligence Index, a general capability benchmark spanning nine evaluations, Lightning scores 24, tied with gpt-oss-120b and behind Nemotron 3 Super, Gemma 4 31B, Claude 4.5 Haiku and Mistral Medium 3.5, all at 30. Lightning isn't a general-intelligence leader in its size class, and Nvidia isn't claiming it is. The actual claim is narrower: according to PinchBench data supplied by Nvidia, Lightning matches Qwen3.6-35B's accuracy roughly 30% faster and beats Gemma 4 26B's accuracy at a similar completion time on PinchBench, a real-world agent task benchmark spanning coding, research and file management. That's a speed-to-accuracy tradeoff, not a capability win. Post-training is where Nvidia says the bigger gains show up. The company shared before-and-after figures from four early-access partners: CrowdStrike's malicious-content recall against a Nemotron 3 Super baseline, CodeRabbit's coding router against a GPT 5.4 Nano baseline, Harvey and Trajectory's legal task completion against an Opus 4.6 baseline, and Lila Sciences' energy simulation work against an Opus 4.8 baseline. CodeRabbit's case is the most specific: Nvidia says the standard NeMo Auto model recipe, trained for one epoch, built into a working router agent for $85 in about two hours. What this means for enterprises There is no shortage of competitive offerings in the growing market for open models. The new Nemotron Lightning release will be yet another option for organizations to consider. On the model side, Lightning's own benchmark chart picks Qwen3.6-35B as its direct comparison point. Asked by VentureBeat directly how Lightning compares to Chinese models more broadly, Briski didn't offer a head-to-head benchmark, pointing instead to openness and customizability as the differentiator. "Our value proposition is not just open and it's very customizable," Briski said. For enterprises building agentic infrastructure, three trends stand out: The routing decision is becoming dynamic instead of static. Enterprises that built agent pipelines around a single default model are being pushed toward per-step routing based on live signals like agent state and token cost, not a fixed assignment set at design time. Open source is now a cost lever at two layers, not one. Pairing an open model with an open router a vendor controls end to end is a newer argument than cheaper weights alone, and worth watching for whether other labs follow the same pattern. The competitive question shifts from best model to best system. As routing libraries mature, the differentiator moves from which model an enterprise defaults to, toward how well its routing layer matches models to tasks in production, a harder thing to benchmark and a harder thing to market.
[6]
Nvidia releases Nemotron 3.5 Lightning open-source AI model
Nvidia $NVDA released Nemotron 3.5 Lightning on Tuesday, a free, open-source AI model the company says can run on a single graphics processing unit on a PC and is designed for autonomous agent workloads. The new model is a 30-billion-parameter mixture-of-experts model built for specialized tasks within larger multi-agent systems, the company said. According to CNBC, companies can use, adapt, and redistribute it at no cost and without seeking Nvidia's approval. Nvidia said the model delivers up to four times faster output speed, leading to 30% faster task completion compared with other models in its class. Nvidia said companies including CrowdStrike $CRWD, CodeRabbit, and Harvey have tested and customized the model. According to CNBC, Nvidia applied distillation -- a technique that involves training a smaller model using outputs from a larger one -- so that Nemotron 3.5 Lightning punches above its weight relative to the company's bigger Nemotron offerings. The model can run on local systems including Nvidia RTX PCs, Nvidia DGX Spark, and Nvidia DGX Station, or scale across data centers and cloud environments, the company said. It is available on Hugging Face, ModelScope, OpenRouter, and Nvidia's website. Alongside Nemotron 3.5 Lightning, Nvidia released NeMo Switchyard, an open-source model routing library that automatically directs requests to the most suitable and cost-efficient model for each step of an agent workflow. Internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to roughly one-third that of Anthropic's Opus 4.8 used alone, the company said. NeMo Switchyard is available on GitHub. Several companies have tested the routing software. Ramp used NeMo Switchyard to match a frontier model's performance while cutting costs by 58% and reducing runtime by 33%, Nvidia said. LangChain achieved 74% lower cost in multi-turn agent tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff. The release comes as Nvidia CEO Jensen Huang has argued publicly that open-source AI is good for chip sales, saying cheaper and more accessible models draw more people into AI ecosystems and ultimately drive up demand for the hardware and data centers Nvidia provides. "Free AI should be great for hardware," Huang told Axios last month. "Free AI should be great for chips." Huang's comments came against the backdrop of a Washington debate over open-weight AI, which intensified after Beijing-based Moonshot AI unveiled Kimi K3, a model whose capabilities brought it closer to the top-performing American systems. On Monday, Meta $META CEO Mark Zuckerberg also published a case for open-source AI as his company released a new coding model.
[7]
Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options
Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard. As enterprises find themselves drowning in artificial intelligence model options, the question is no longer raw power and capability, but fit-for-what-purpose and when. As agents become the norm, the tasks they perform range from sequences of swift, simple tasks to high-complexity reasoning across deep knowledge domains. The former might require small models that can swiftly handle small sorting tasks efficiently and the latter would be best handled by frontier models with high intelligence and reasoning capability. Nemotron 3.5 Lightning joins the open model family to support high-volume tasks when powering always-on agents. It's a 30 billion-parameter mixture-of-experts model that Nvidia claims is capable of delivering up to four times the output speed and 30% faster agentic task completion compared with other models in its weight class. It's also open and customizable, meaning it can be readily post-trained with Nvidia NeMo on an enterprise's own hardware, with proprietary domain data, tools and workflows to improve accuracy for specialized tasks. This alone makes the model a powerful backbone workhorse within the family. Nvidia launched the Nemotron 3 model family in December to act as an open foundation for agentic AI systems. At the time, it consisted of three sizes: Nano, Super and Ultra. The company said Lightning's differentiator is that it can be specialized quickly. A company can reproducibly regenerate its own data and specialization in a newly trained model with minimal equipment and cost. "What we're hearing is that Lightning is remarkably easy to customize," said Vice President of Generative AI Kari Briski. CodeRabbit Inc., according to Briski, used Nvidia's standard auto model recipe, trained for one epoch, and produced a router agent for $85 in around two hours. She also said another partner dropped Lightning directly into an existing post-training stack "with no changes required." In one case she said, a company did training on a single H100 card relatively inexpensively, and another simply set up a training job overnight and came back in the morning to pick up the results. This is turning post-training and fine-tuning experience into loading the dishwasher and hitting the button. Nvidia is also releasing its datasets, its post-training datasets and the recipes/framework specifically used to train the models. Developers could blend Nvidia's own post-training data with their own enterprise data to create powerful hybrids that get the best of both worlds. "That's why we see people getting really great results really quickly," Briski said. Picking the right model for the task, why AI agents need a router Cost, accuracy and efficiency aren't "cheap model versus expensive model"; it's the best fit for the task. When an enterprise agent is doing work, the tasks vary in complexity, expertise, token use, prediction and capability. That means routing to a model that will do the task correctly is similar to providing a task to a person. Some models are better at coding, others are better at inferring context from text, and yet others have specialized knowledge instilled in them. NeMo Switchyard is an open-source model routing library for AI agents that takes into account what numerous AI models are available and their capabilities, and routes prompts to the most capable and efficient model available at each step of the agent workflow based on specific needs. "Depending on your routing strategy, it wants to choose the best model," Briski explained. Developers can customize the router with their own choice of strategy depending on their priority. For example, quality, delay and cost requirements. A system of models could be sorted for tokenomics, speed or particular quirks, expertise or other needs that an industry requires. Nvidia said it's working with a number of partners across the AI ecosystem on intelligent AI model routing alongside tools developers already use, including Boomi LP, Cadence Design Systems Inc., Classmethod Inc., Cognition AI Inc., Kong Inc., Langchain Inc., Nous Research Inc. and Siemens AG.
[8]
Nvidia is developing Nemotron 4 open-source models: The Information
The chip giant is among the few major US firms to release open-source models, which have drawn more attention this year ā as ā AI bills balloon and cheap Chinese models near the capabilities of top systems from leading American labs Anthropic and OpenAI. Nvidia is developing a new AI model family, Nemotron 4, with the goal of rivaling top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major US firms to release open-source models, which have drawn more attention this year ā as ā AI bills balloon and cheap Chinese models near the capabilities of top systems from leading American labs Anthropic and OpenAI. A spate of recently disclosed hacks involving autonomous AI agents has added to the attention, especially because open models do not have curbs on cybersecurity use. The largest Nemotron 4 model is expected to have at least 1 ā trillion parameters, according to multiple employees working on the project, The Information reported. Nvidia has not set a release date for Nemotron ā 4 and has yet to complete final training, though employees said the model could be ready as early as late fall, according to the report. The company did not immediately respond to a Reuters request for comment on the report. Nvidia last month formed a coalition with other companies to develop and share tools for AI safety and cybersecurity. It also signed an open letter with tech heavyweights such as Microsoft backing open-weight models so that innovation does not drift overseas. Separately on Tuesday, the chip ā firm unveiled Nemotron 3.5 Lightning, an addition to its offerings aimed at code review, tool use, security alert monitoring, answering billing questions and other tasks. It also released NeMo Switchyard, an open-source model-routing library designed to automatically direct AI tasks to the most suitable models. Late last year, the chip giant unveiled the third generation of its family of open-source models as offerings from Chinese AI labs proliferated.
[9]
NVIDIA Nemotron 3.5 Lightning Delivers 4X AI Throughput
NVIDIA's Nemotron 3.5 Lightning is an AI model designed for execution-layer tasks, emphasizing speed and efficiency in workflows that require high throughput. According to Sam Witteveen, the model employs a hybrid member transformer architecture with an active mixture of experts (MoE), using up to 3 billion parameters simultaneously. This approach supports operations such as data validation, summarization and classification. Techniques like speculative decoding, including D-Flash and D-Spark, enable throughput improvements of up to 4x, making it suitable for large-scale automation scenarios. Discover how the Nemotron 3.5 Lightning integrates with Nvidia's SwitchYard routing system to optimize task allocation and improve operational workflows. Gain insight into its applications in cybersecurity and AI chaining through examples like CrowdStrike and Code Rabbit. This analysis also examines its open source licensing, customization capabilities and the resources available for fine-tuning, providing a detailed understanding of how organizations can implement and adapt the model effectively. What Distinguishes Nemotron 3.5 Lightning? The Nemotron 3.5 Lightning is tailored to handle the operational backbone of AI processes, often referred to as the "grunt work" of automation. It is optimized for repetitive, well-defined tasks, including chaining tools, validating outputs and summarizing data. Unlike models that focus on creative problem-solving or high-level reasoning, this model emphasizes speed and reliability, making it a practical choice for organizations that depend on long-running AI workflows or require precise task validation. Its design prioritizes efficiency and scalability, allowing businesses to handle high-throughput operations without compromising on accuracy. By focusing on execution-layer tasks, the model ensures that organizations can maintain operational efficiency while reducing computational overhead. Key Technical Features The architecture of Nemotron 3.5 Lightning is a cornerstone of its performance, balancing computational power with resource efficiency. With 30 billion parameters and an active mixture of experts (MoE) using 3 billion parameters at any given time, the model achieves a remarkable combination of precision and speed. Its hybrid member transformer architecture further enhances its ability to deliver consistent results across a variety of tasks. Highlighted features include: * Multi-token prediction: This capability allows the model to generate multiple tokens simultaneously, significantly accelerating output generation and improving overall efficiency. * Speculative decoding: Advanced techniques such as D-Flash and D-Spark enhance decoding speed, delivering up to 4x throughput and achieving 30-35% faster performance compared to similar models. These features make the model particularly effective in environments where high throughput and accuracy are critical, such as large-scale data processing or real-time task execution. Advance your skills in NVIDIA by reading more of our detailed content. Customization and Real-World Applications The Nemotron 3.5 Lightning is open source, providing access to its weights and allowing extensive customization to meet specific organizational needs. Nvidia supports this customization with post-training recipes and datasets, allowing businesses to fine-tune the model for various industries and applications. Notable real-world use cases include: * CrowdStrike: Leveraged the model for cybersecurity tasks, achieving significant reductions in both cost and time while enhancing operational efficiency. * Code Rabbit: Utilized the model to streamline AI tool chaining and validation processes, improving workflow reliability and speed. These examples underscore the model's versatility and its ability to deliver measurable efficiency gains across diverse sectors, from cybersecurity to software development. Strengths and Limitations The Nemotron 3.5 Lightning is particularly well-suited for execution-layer tasks, excelling in areas such as tool chaining, managing long-running processes and handling repetitive operations. Its high-speed performance and task-specific focus make it a valuable asset for organizations prioritizing operational efficiency. However, the model does have limitations: * Not designed for advanced reasoning: The model is not suitable for tasks requiring high-level intelligence or creative problem-solving. * Vulnerability to prompt injection: It lacks robust resistance to prompt injection attacks, making it less ideal for front-line or coding agent roles. Despite these constraints, its strengths in efficiency, scalability and task-specific performance make it a compelling choice for organizations focused on execution-focused applications. Licensing and Accessibility The Nemotron 3.5 Lightning is released under the Open MDW license, which allows for commercial use and customization without requiring attribution. This licensing model ensures that businesses can integrate the model into their workflows without encountering legal or financial barriers. The model is also fully compatible with Nvidia hardware, including RTX graphics cards and DGX Spark systems. For organizations already operating within the Nvidia ecosystem, integrating Nemotron 3.5 Lightning is seamless, further enhancing its appeal as a practical and efficient AI solution. Integration with Nvidia Ecosystem As part of the Nemotron 3 family, the Lightning model is distilled from the more advanced Nemotron 3 Ultra. It is designed to work seamlessly with NVIDIA's SwitchYard routing system, which facilitates model orchestration and task allocation. This integration ensures efficient operation within larger AI frameworks, streamlining workflows and enhancing overall system performance. By using the SwitchYard system, organizations can optimize task distribution and maximize the model's potential, making sure that resources are allocated effectively across various AI processes. Resources for Customization and Learning NVIDIA provides a comprehensive suite of resources to help organizations maximize the potential of Nemotron 3.5 Lightning. These include training recipes and datasets that support advanced strategies such as curriculum learning and on-policy distillation. By using these tools, businesses can fine-tune the model to meet their unique requirements, optimizing its performance for specific tasks and industries. These resources empower organizations to adapt the model to their needs, making sure that it delivers maximum value and aligns with their operational goals. Practical Implications for Businesses The Nemotron 3.5 Lightning offers a practical, high-performance solution for organizations aiming to optimize execution-layer AI tasks. Its combination of speed, customization potential and seamless integration with NVIDIA's ecosystem makes it an invaluable tool for handling repetitive, task-specific operations at scale. While it is not designed for high-level reasoning, its strengths in efficiency, adaptability and scalability make it a compelling choice for businesses looking to enhance their AI capabilities and streamline their workflows. Media Credit: Sam Witteveen Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[10]
NVIDIA's Nemotron 3.5 Lightning Accelerates Token Generation By 4x But Agentic Tasks Only Speed Up By 30%, As Orchestration Remains The Real Bottleneck
The race for open-weight models in the US is heating up even as China continues to own this segment, as evidenced by NVIDIA's release of its Nemotron 3.5 Lightning AI model, which has dropped just hours after Meta introduced its Muse Glimmer model on Monday. NVIDIA's Nemotron 3.5 Lightning is trying to solve the token bottleneck problem when the real constraints lie with the orchestration layer As stated earlier, NVIDIA has just released the Nemotron 3.5 Lightning, an open-weight, mixture-of-experts (MoE) model with 30 billion parameters (3 billion active parameters) that is designed to handle and orchestrate always-on agents that grind through a constant stream of routine tasks. For the benefit of those who might not be aware, orchestration refers to the process of coordinating multiple specialized AI agents to work together efficiently to solve complex, multi-step tasks that a single AI agent cannot handle alone. This agent-level orchestration involves a number of steps, including task intake and decomposition, agent selection, context and state sharing (when Agent A finishes its step, the orchestrator passes that specific data to Agent B so progress never resets), periodic validation and error correction. Do note that hardware-level orchestration occurs within the CPU, which typically manages the flow of data between memory and compute cores to minimize latency and maximize throughput. Architectural elements NVIDIA's Nemotron 3.5 Lightning has 4 key architectural elements: Why NVIDIA's Nemotron 3.5 Lightning offers diminishing returns NVIDIA's Nemotron 3.5 Lightning achieves a score of 24 on the Artificial Analysis Intelligence Index, which is a composite benchmark score designed to evaluate and track the overall capabilities of LLMs as they progress towards Artificial General Intelligence (AGI), and measures a given model's agentic, coding, scientific reasoning, and general-purpose capabilities. This score is a significant improvement over what its predecessor - the Nemotron 3 Nano - achieved. Even so, NVIDIA claims that the Nemotron 3.5 Lightning is able to generate tokens at a rate that is 4x faster than "similar-sized models," but that only speeds up actual agentic tasks by just 30 percent. In other words, the model quadruples token output but delivers a diminishing returns on agentic task acceleration. This suggests that the real bottleneck lies within the orchestration layer and agent harness - the software that decides what runs, where it runs, and in what order. While the Nemotron 3.5 Lightning certainly moves the needle when it comes to incremental utility, it lacks the oomph factor required to propel US open-weight models to an ascendant state, which might partially explain why NVIDIA has already started hyping up the Nemotron 4 as its next big bet. Follow Wccftech on Google to get more of our news coverage in your feeds.
[11]
NVIDIA introduces Nemotron 3.5 Lightning 30B open AI model and NeMo Switchyard for agentic AI
NVIDIA has introduced Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts (MoE) open model designed for high-volume execution tasks in long-running AI agents. The company has also released NeMo Switchyard, an open-source model routing library that directs requests to different AI models based on requirements such as quality, latency and cost. Nemotron 3.5 Lightning is designed for specialized tasks within multi-agent systems, while larger reasoning models can handle planning and orchestration. NVIDIA says it delivers up to 4x faster output than similar-sized models and enables 30% faster agentic task completion. Always-on agents and systems of models Long-running AI agents increasingly use multiple models for different tasks. Larger reasoning models such as Nemotron 3 Ultra can handle planning and orchestration, while smaller models such as Nemotron 3.5 Lightning can perform tasks including code review, tool use, security alert monitoring and answering billing questions. NVIDIA Nemotron 3.5 Lightning Nemotron 3.5 Lightning is the smallest model in the Nemotron 3 family, with 30 billion total parameters and 3 billion active parameters. Its MoE architecture activates only a subset of experts for each token. The model is designed for high-volume agent execution, including tool calls, result validation and subagent delegation. NVIDIA said it was developed with contributions from the Nemotron Coalition, which provided evaluation methodologies, inference software and datasets. Key specifications: * Architecture: 30B MoE * Active parameters: 3B * Formats: BF16, NVFP4 * Primary use: High-volume agent execution * Deployment: Local, on-premises, edge, workstation, data center and cloud Performance and optimization NVIDIA says Nemotron 3.5 Lightning provides up to 4x faster output than other models in its class and completes agentic tasks 30% faster. On PinchBench, it reaches 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at similar accuracy. The model uses multi-token prediction (MTP) and speculative decoding. It includes two draft models: * DSpark: Recommended for DGX Spark inference and low-concurrency data center workloads. * DFlash: An additional draft model for workload evaluation. MTP is intended for medium- to high-concurrency workloads, with the optimal draft length decreasing as concurrency increases. The NVFP4 checkpoint uses specialized NVFP4 kernels also used by Nemotron 3 Ultra and supports NVIDIA Blackwell, Hopper and Ampere GPUs. The model also uses harness-optimized training for popular agent harnesses. Customization and deployment Nemotron 3.5 Lightning can be customized with an organization's own data, tools and workflows. NVIDIA has released its weights, training data and recipes as permissively as licensing permits under OpenMDW-1.1. Developers can use NeMo Automodel and NeMo Megatron Bridge for LoRA or full SFT, while NeMo RL and NeMo Gym support reinforcement learning, evaluations and rollouts. The release also includes Nemotron-RL-Agentic-Terminal-Pivot, an open reinforcement learning dataset used for coding-agent capabilities. NVIDIA said it publishes as much of the training data and techniques as licensing permits for traceability, auditing and training other models. Industry use cases NVIDIA said companies are customizing the model for: * CrowdStrike: Cybersecurity * Harvey with Trajectory: Legal services * CodeRabbit with Baseten: Code review * Lila Sciences: Physical and life sciences * Fastino Labs: Software development, finance and healthcare Nemotron 3.5 Lightning can run on NVIDIA RTX PCs, DGX Spark, DGX Station, Jetson, GeForce RTX 5090 and RTX PRO workstations. It is also supported by LM Studio, llama.cpp, Ollama and Unsloth. NVIDIA said it worked with teams including EXO Labs to evaluate the model on DGX Spark. The model supports agent harnesses including OpenClaw and Hermes Agent, which are supported by NVIDIA's NemoClaw open-source security and management stack. Accuracy and efficiency Nemotron 3.5 Lightning reaches the accuracy-speed Pareto frontier on the Artificial Analysis Intelligence Index, according to NVIDIA. The index combines nine evaluations covering agentic tasks, coding, scientific reasoning and general intelligence. On PinchBench, the model reaches 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at similar accuracy. NVIDIA says its inference throughput and token efficiency place it on the efficiency frontier for high-volume workloads. NeMo Switchyard NeMo Switchyard is an open-source model routing library for AI agents. It can route prompts to different models at each step of an agent workflow based on requirements such as quality, latency and cost. Nemotron 3.5 Lightning can be used as a routing target alongside other open and closed models, allowing larger models to handle planning while smaller models handle execution. NVIDIA's internal benchmarks show that Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of using Opus 4.8 alone. Companies including Boomi, Cadence, Classmethod, Cognition, Kong, LangChain, LiteLLM, Nous Research, Ramp and Siemens are evaluating or integrating Switchyard across AI tools and platforms. Partner ecosystem Nemotron 3.5 Lightning is supported across post-training, inference, agent frameworks and cloud platforms. * Post-training: AgileRL, Applied Compute, Deep Cogito, Fastino Labs, Locai Labs, Prime Intellect, Reasonable, Thinking Machines Lab, Thoughtworks, Trajectory, Uniphore * Inference: Ollama, Exo, Canonical, LM Studio, Unsloth * Agent frameworks: Aible, Cline, Factory AI, Hermes Agent, Kilo Code, LangChain, OpenClaw, OpenCode, OpenHands, Pi * Cloud platforms: Amazon SageMaker JumpStart, Google Cloud Gemini Enterprise Agent Platform, MSFT Foundry, OCI Enterprise AI * GSI: Accenture, Tata Consultancy Services, Tech Mahindra, Wipro * Hosted inference: Baseten, CoreWeave, Crusoe, DeepInfra, Fireworks AI, FriendliAI, GMI Cloud, Modal, Nebius, Together AI Availability Nemotron 3.5 Lightning is available through Hugging Face, ModelScope, OpenRouter and build.nvidia.com, including as an NVIDIA NIM microservice. It is also available through NVIDIA Cloud Partners and other inference, post-training and cloud platforms. NeMo Switchyard is available on GitHub and is coming to partner platforms. NVIDIA also provides deployment guides for vLLM, SGLang and TensorRT-LLM, along with documentation for Nemotron 3.5 Lightning and Switchyard.
[12]
Nvidia building 1-trillion-parameter Nemotron 4 to rival open AI models
Nvidia is developing a new AI model family, Nemotron 4, with the goal of challenging top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major U.S. firms to release open-source models, which have drawn more attention this year as AI bills balloon and cheap Chinese models near the capabilities of top systems from leading American labs Anthropic and OpenAI. A spate of recently disclosed hacks involving autonomous AI agents has added to the attention, especially because open models do not have curbs on cybersecurity use. The largest Nemotron 4 model is expected to have at least 1 trillion parameters, according to multiple employees working on the project, The Information reported. The company has not set a release date for Nemotron 4 and has yet to complete final training, though employees said the model could be ready as early as late fall, the report said. Nvidia pointed to previous remarks that it was working on Nemotron 4 but did not confirm the details in The Information's report. "Nvidia is investing in Nemotron because we believe every company and every country needs accessible frontier open models to strengthen safety and security, accelerate innovation, and provide a foundation they can rely on from one generation to the next," Kari Briski, vice president of generative AI, said in an emailed statement. The chip giant formed a coalition last month with other companies to develop and share tools for AI safety and cybersecurity. It also signed an open letter with tech heavyweights such as Microsoft MSFT.O backing open-weight models so that innovation does not drift overseas. Separately on Tuesday, the chip firm unveiled Nemotron 3.5 Lightning, an addition to its offerings aimed at code review, tool use, security alert monitoring, answering billing questions and other tasks. It also released NeMo Switchyard, an open-source model-routing library designed to automatically direct AI tasks to the most suitable models. (Reporting by Anhata Rooprai in Bengaluru; Editing by Pooja Desai)
[13]
Nvidia readies Nemotron 4, a giant model to bolster its open AI lineup
According to The Information, Nvidia still has to complete the final training phase for Nemotron 4 and has not set an official launch date. Employees involved in the project nevertheless believe the model could be ready as soon as late fall. The group confirms it is working on this new generation without validating the reported specifications. Nvidia says its investments reflect a desire to provide companies and governments with high-performing open models, notably to support innovation, safety, and cybersecurity. This strategy comes as spending on AI is rising sharply and lower-cost Chinese models are closing in on the performance of Anthropic and OpenAI systems. Interest in open models is also being fueled by cybersecurity concerns tied to autonomous agents. Nvidia recently joined a coalition focused on AI safety and, with Microsoft and other groups, signed a letter backing open-weight models to help preserve technological innovation. In parallel, Nvidia introduced Nemotron 3.5 Lightning, aimed in particular at code review, tool use, security alert monitoring, and processing billing requests. The chipmaker also launched NeMo Switchyard, an open source library that can automatically route AI tasks to the most suitable models. These initiatives broaden Nvidia's ecosystem beyond its chips and strengthen its presence in AI models and tools.
[14]
Nvidia readies Nemotron 4, a giant model to strengthen its open AI offering
According to The Information, Nvidia still needs to complete the final training phase for Nemotron 4 and has not set any official release date. Employees involved in the project nonetheless believe the model could be ready as early as late fall. The company confirms it is working on this new generation without validating the reported specifications. Nvidia says its investments reflect a desire to provide companies and governments with high-performing open models, notably to support innovation, safety, and cybersecurity. This strategy comes as AI spending is rising sharply and lower-cost Chinese models are closing in on the performance of systems from Anthropic and OpenAI. Interest in open models is also being fueled by cybersecurity concerns tied to autonomous agents. Nvidia recently joined a coalition focused on AI safety and, with Microsoft and other groups, signed a letter defending open-weight models in order to preserve technological innovation. In parallel, Nvidia introduced Nemotron 3.5 Lightning, aimed in particular at code review, tool use, security alert monitoring, and handling billing requests. The chipmaker also launched NeMo Switchyard, an open source library capable of automatically routing AI tasks to the most suitable models. These initiatives broaden Nvidia's ecosystem beyond its chips and deepen its footprint in artificial intelligence models and tools.
[15]
Nvidia is developing Nemotron 4 open-source models, The Information reports
Aug 11 (Reuters) - Nvidia is developing a new AI model family, Nemotron 4, with the goal of rivaling top open-source models globally, The Information reported on Tuesday, citing people who work on the project. The chip giant is among the few major U.S. firms to release open-source models, which have drawn more attention this year as AI bills balloon and cheap Chinese models near the capabilities of top systems from leading American labs Anthropic and OpenAI. A spate of recently disclosed hacks involving autonomous AI agents has added to the attention, especially because open models do not have curbs on cybersecurity use. The largest Nemotron 4 model is expected to have at least 1 trillion parameters, according to multiple employees working on the project, The Information reported. Nvidia has not set a release date for Nemotron 4 and has yet to complete final training, though employees said the model could be ready as early as late fall, according to the report. The company did not immediately respond to a Reuters request for comment on the report. Nvidia last month formed a coalition with other companies to develop and share tools for AI safety and cybersecurity. It also signed an open letter with tech heavyweights such as Microsoft backing open-weight models so that innovation does not drift overseas. Separately on Tuesday, the chip firm unveiled Nemotron 3.5 Lightning, an addition to its offerings aimed at code review, tool use, security alert monitoring, answering billing questions and other tasks. It also released NeMo Switchyard, an open-source model-routing library designed to automatically direct AI tasks to the most suitable models. Late last year, the chip giant unveiled the third generation of its family of open-source models as offerings from Chinese AI labs proliferated. (Reporting by Anhata Rooprai in Bengaluru; Editing by Pooja Desai)
Share
Copy Link
Nvidia launched Nemotron 3.5 Lightning, a 30-billion-parameter open-source AI model delivering 4x faster token generation, alongside NeMo Switchyard, an open-source routing library that cuts agentic task costs to roughly one-third of frontier models. The dual release addresses rising AI bills while maintaining frontier-level performance across specialized agentic tasks.
Nvidia released Nemotron 3.5 Lightning and NeMo Switchyard on August 11, marking the chip giant's first open-source AI model launch since CEO Jensen Huang publicly defended open models in late July
2
. The 30-billion-parameter mixture-of-experts model delivers up to 4x faster token generation and 30% faster task completion compared to open models in its class3
. Paired with the NeMo Switchyard open-source routing library, internal benchmarks show the combination maintains frontier-level task completion while reducing costs to roughly one-third of running Opus 4.8 alone3
. This dual release addresses the core challenge enterprises face with agentic AI: balancing performance against escalating token costs without building custom routing infrastructure5
.
Source: VentureBeat
The customizable open model targets high-volume specialized agentic tasks within larger multi-agent systems, including code review, tool use, security alert monitoring and answering billing questions
1
. Because Nemotron 3.5 Lightning uses open weights, developers can fine-tune it with their own examples to match specific workflows, writing styles or coding conventions3
. AI leaders across industries are already customizing the model: CrowdStrike for cybersecurity, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review4
. The model runs locally on RTX PCs, DGX Spark, Jetson devices and scales to RTX PRO workstations, data centers and cloud environments3
. Nvidia collaborated with vLLM, Ollama, llama.cpp and LM Studio to optimize local deployment, while Unsloth provides day-one support with quantized models through Unsloth Studio3
.
Source: Geeky Gadgets
The open-source routing library automatically directs each step of an agent workflow to the most suitable model based on accuracy, speed and cost requirements
3
. Unlike fixed routing approaches, NeMo Switchyard responds to shifting agent states as tools return results or errors emerge, adjusting model selection dynamically rather than using static task categories5
. The router evaluates model verbosityāpredicting how many tokens a given model producesāand steers work toward cheaper options before calls are made5
. Integration partners split into two groups: agent frameworks like Cognition, LangChain and Nous Research that call Switchyard directly, and LLM gateways including Kong, LiteLLM and OpenRouter that built native support5
. Early testing shows dramatic results. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, with only a 6% accuracy tradeoff5
.
Source: SiliconANGLE
Related Stories
The release lands during the busiest open-weight stretch the industry has seen in months, with Alibaba, Moonshot, Zhipu and DeepSeek shipping competitive Chinese models since spring, several approaching frontier performance while undercutting US labs on size or price
5
. Meta added pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer5
. Jensen Huang has argued that free AI drives hardware sales, telling Axios that "free AI should be great for chips" since models still require GPUs regardless of licensing2
. Nvidia remains among the few major US firms releasing open-source models, which have drawn increased attention as AI bills balloon and cheap Chinese models near capabilities of top systems from Anthropic and OpenAI1
. The company also formed a coalition last month with other firms to develop tools for AI safety and cybersecurity, signing an open letter with Microsoft backing open-weight models to prevent innovation from drifting overseas1
.Separately, Nvidia is developing Nemotron 4, a new AI model family targeting top open-source models globally, according to The Information
1
. The largest Nemotron 4 model is expected to have at least 1 trillion parameters, with employees suggesting readiness as early as late fall, though Nvidia has not set a release date or completed final training1
. Watch for how enterprises adopt model routing strategies as intelligent agents shift from experimental to production workloads, particularly whether cost savings materialize at scale beyond benchmark tests. The competitive landscape will likely intensify as Chinese labs continue releasing capable open models and US firms respond with their own open alternatives, potentially reshaping how organizations balance performance, cost and control in agentic AI deployments.Summarized by
Navi
[4]
15 Dec 2025ā¢Technology

11 Mar 2026ā¢Technology

18 Oct 2024ā¢Technology

1
Technology

2
Science and Research

3
Technology
