3 Sources
[1]
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026
Local agents get easier to install, faster to run and able to tap multiple RTX PCs at home with NVIDIA PAIR -- plus, new NVIDIA RTX Spark Windows PCs arriving in October. Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely. Today's announcements include: * Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer. * Up to 1.9x faster local inference -- new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama. * NVIDIA PAIR -- a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user's local network. * NVIDIA RTX Spark arrives in October -- with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark. Also, August was a busy month for local AI: * Nemotron 3.5 Lightning -- which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson -- is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today. * Z.ai's GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that's bringing agentic AI to DGX Station. * Qwen has released Qwen3.8-Flash-Next, an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B, a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs. * LTX's LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation. * MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI. FastVideo teamed up with NVIDIA researchers to improve this further by releasing FastH3 -- an open-weight, four-step distilled version that improves performance by 7x. Optimized FastVideo recipes for NVIDIA RTX GPUs and DGX Spark are coming soon. * Meta's Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment. * DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on 2x DGX Spark cluster and DGX Station. A Simpler Start for Local Agents Getting a local agent up and running with local models required some effort -- choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems. Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA's latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running. Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience. Perplexity Portable Computer is available on NVIDIA RTX GPUs with at least 24GB VRAM running on Linux, with support on Windows coming soon, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here's some example use-cases: * Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes. * Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page -- all without a document ever reaching a chatbot. * Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team's Slack channel. Try Portable Computer today. Hermes Agent -- developed by Nous Research -- is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit. Configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on Windows. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Support for Linux is coming soon. Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system. One-click local model setup is available now on Windows, with support coming soon to Linux. Learn more about Hermes Agent. OpenClaw has become one of the defining projects of the open-agent movement -- the largest AI project on GitHub, with more than 380K stars and a fast-growing community that's building tools and skills across research, engineering, project management and everyday productivity. NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM. Learn more in the OpenClaw blog. Faster Inference Gives Local Agents a Boost Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms. llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill. vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms. These gains are available on the llama.cpp and vLLM inferencing backends. Users can also experience these via the LM Studio and Ollama applications. Tap Idle PCs for More Local AI Compute With NVIDIA PAIR More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI. Agentic workflows often break complex tasks into smaller jobs that can run in parallel, but performance can slow when every request is competing for the same GPU. PAIR automatically discovers compatible PCs on a local network and routes independent inference requests to whichever system has capacity. It works with Ollama and LM Studio and can adapt as devices join or leave the network. For example, a user could ask Hermes to create a "Sunday Reset" plan by sorting through a cluttered inbox and prioritizing what needs attention now, what can wait and what can be skipped. Hermes can split that work across multiple subagents, while PAIR distributes those jobs across available PCs instead of having them all wait on a single GPU. The result is more compute for local agents, with more tasks running in parallel and the flexibility to move AI workloads to another PC while the main system is being used for gaming, creating or other work. The NVIDIA PAIR beta is available for Windows, macOS and Linux through both graphical and terminal interfaces, supporting NVIDIA GeForce RTX 20 Series GPUs and newer, NVIDIA RTX PRO workstation GPUs (Turing architecture and newer), NVIDIA DGX Spark and Apple M4 or newer silicon. Check out the NVIDIA tech blog to get started with NVIDIA PAIR. Powerful On Device Photo Editing With Cyberlink PhotoDirector AI PC Mode on RTX Spark Open image and video models enable artists to experiment with Creative AI models on PCs. This enables artists to iterate and explore concepts and ideas, without the dreaded token anxiety and keep more of their creative work private and on-device. CyberLink's new PhotoDirector AI PC Mode is one of the first applications to integrate these diffusion models directly into a creative software, and turn them into a creative tool at the finger tips of the artists. Coming to PhotoDirector 365 and optimized for NVIDIA RTX Spark when it launches, AI PC Mode users are getting AI-powered editing tools for generative editing, image enhancement, object and distraction removal, background removal and replacement, portrait refinement and the creation of entirely new visuals -- with the flexibility to choose between local or cloud processing, depending on the task. On NVIDIA GPUs, PhotoDirector uses TensorRT-RTX and FP8 to accelerate local AI. Start using Cyberlink's PhotoDirector 365 photo editing software and learn more about PhotoDirector AI PC Mode, launching with RTX Spark in October. NVIDIA RTX Spark Windows PCs Arrive October 2026 NVIDIA RTX Spark is coming this October -- and at IFA 2026, partners are showing off their hardware. At IFA, newly announced designs join the existing six OEMs shipping in October. Acer showed its compact desktop RTX Spark concept, and Lenovo announced its Yoga Pro 9n and Yoga 9n 2-in-1. RTX Spark is a new beginning for Windows PCs. One PC built for creators, gamers and AI agents. With a powerful 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a highly efficient 20-core Grace CPU, RTX Spark delivers incredible performance and efficiency. This superchip enables high performance thin laptops with all day battery life and compact desktops to power always-on agents. Paired with the new Windows Agent framework, it enables agents that run safely in the background under OS level control. Last week at Gamescom, Electronic Arts, Embark and Ubisoft were among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark Windows PCs. They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX. Read more. Sign up to be notified when RTX Spark laptops and desktops are available. #ICYMI: More Updates From NVIDIA Local AI 🎮 NVIDIA Brings New RTX Tech and Games to Gamescom -- NVIDIA released DLSS 4.5 Ray Reconstruction, featuring a new second-generation transformer model for improved image quality in ray-traced and path-traced games. Gamescom also brought new RTX announcements for titles including 007 First Light, CONTROL Resonant and Gears of War: E-Day, plus expanded game support for the upcoming NVIDIA RTX Spark. 🐋Introducing DeepSeek Harness -- DeepSeek's new open source harness pairs with DeepSeek-V4-Flash to power local agentic coding workflows on NVIDIA DGX Station and multi-DGX Spark setups. 📊MLPerf Client v2.0 Expands AI PC Benchmarking -- MLCommons released MLPerf Client v2.0, developed in collaboration with NVIDIA and other industry leaders. The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads. Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook -- and stay informed by subscribing to the NVIDIA Local AI newsletter. Follow NVIDIA Workstation on LinkedIn and X. See notice regarding software product information.
[2]
NVIDIA Brings Simplified Local AI Support To NVIDIA GPUs Carrying 24+ GB VRAM While vLLM & & llama.cpp Optimizations Boost Compute By Up To 1.9x
NVIDIA is bringing simpler local AI capabilities and adding various optimizations to its GPUs on RTX and DGX platforms. Local Agents To See Up To 1.9x Faster Performance Through Latest Optimizations Across NVIDIA RTX/DGX Platforms The first announcement is faster local agents, which are delivered through continued optimizations that NVIDIA has collaborated on with the open-source llama.cpp and vLLM communities. The latest results were measured on SpeedBench-Coding 8K Throughput with AIPerf, and the results are as follows. Starting with Llama.cpp, NVIDIA RTX platforms such as the GeForce RTX 5090 now offer up to a 50% boost in Token throughput (tok/s) in Qwen3.6-27B, and a 90% acceleration in Qwen3.6-35B. The NVIDIA RTX PRO 6000 Blackwell also sees a nice boost in vLLM with up to 20% speedup across both models. Lastly, NVIDIA's DGX Spark platforms see a 1.4x increase in Qwen3.6-27 B in vLLM & a 20% increase in AI inferencing on DeepSeek v4 Flash. With these kernel optimizations, NVIDIA RTX and DGX platforms will see enhanced speculative decoding techniques and faster prefill. On vLLM, new XQA attention kernels in FlashInfer and backend optimizations help accelerate local AI inference across both platforms. In addition to faster local AI, NVIDIA is also making local AI simpler through new one-click AI support across Hermes Agent, OpenClaw, and Perplexity Computer. These will be available for NVIDIA GPUs featuring 24 GB of memory or higher. Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration, and tools packaged into a single app experience. In September, Perplexity Portable Computer will be available on NVIDIA RTX GPUs with at least 24GB VRAM running Linux or Windows, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Hermes Agent, developed by Nous Research, is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations, and DGX Spark a natural fit. Coming soon, configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on both Windows and Linux. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions, and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system. One-click local model setup is coming soon. Learn more about Hermes Agent. OpenClaw has become one of the defining projects of the open-agent movement -- the largest AI project on GitHub, with more than 380K stars and a fast-growing community that's building tools and skills across research, engineering, project management and everyday productivity. NVIDIA, Microsoft, and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM. Follow Wccftech on Google to get more of our news coverage in your feeds.
[3]
IFA 2026: NVIDIA is making your home into a connected experience with PAIR, RTX Spark and one-click AI agents
Hermes Agent, OpenClaw and Perplexity are adding one-click local AI setup for NVIDIA-powered PCs. NVIDIA is belting out announcements back to back with the dust not having settled on their Gamescom announcements, NVIDIA used IFA 2026 to lay out a broader vision for local AI on the PC, one that goes well beyond simply running a language model on a graphics card. The company announced NVIDIA PAIR, new details about the N1X processor powering its RTX Spark platform, new RTX Spark systems from Lenovo and Acer, performance improvements for popular local inference frameworks, and closer integration with AI agents including Hermes, OpenClaw and Perplexity. The common thread running through almost every announcement is agents. NVIDIA believes the PC is gradually shifting from a machine where users manually operate applications to one where users tell an AI agent what they want accomplished, with the agent coordinating the applications, tools and models needed to do it. According to NVIDIA, these agents can combine models running locally with cloud models, allowing tasks involving private data to remain on the PC while more demanding workloads can still be escalated to larger cloud-based systems. One-click local AI is coming to more agents One of NVIDIA's immediate priorities is making local AI substantially easier to set up. Running an LLM locally can still involve choosing a model, finding the appropriate quantisation, selecting an inference engine, downloading the required components and configuring them correctly. NVIDIA said it has been working with agent developers to reduce much of this process to a single click. Hermes Agent will add a local-model option during onboarding that can recommend a suitable model, download it and configure the required environment automatically. NVIDIA said the capability will support RTX GPUs and NVIDIA systems running Windows and Linux. OpenClaw is getting a similar capability. Its Windows application will be able to automatically select an appropriate local model, download it and configure it for NVIDIA GPUs. NVIDIA said the feature is due later this month. Perplexity is also expanding its local-agent ambitions to RTX systems. The company already offers an agent capable of switching between local and cloud processing, and NVIDIA said this experience is now coming to RTX GPUs across Windows and Linux. The local mode uses a custom model trained by Perplexity and is designed to provide private, unmetered conversations, while harder requests can automatically escalate to the cloud when additional capability is required. NVIDIA positioned the hybrid approach as particularly relevant to professional workloads. A tax specialist, for instance, could use Perplexity's online capabilities to research changes to tax law while processing sensitive client information locally. Similarly, financial professionals could combine online financial datasets with confidential data that cannot legally be uploaded to external services. It is a useful illustration of where NVIDIA sees local AI fitting into the broader AI ecosystem. Rather than positioning local models as replacements for cloud AI, the company increasingly appears to see them as complementary. Up to 1.9x faster local LLM inference NVIDIA is also working with the open-source ecosystem to improve the performance of models already running locally. The company said recent llama.cpp optimisations can deliver up to 1.9x higher throughput, enabled by kernel improvements, enhanced speculative decoding techniques and other inference optimisations. NVIDIA also highlighted improvements to vLLM, claiming a 1.2x performance improvement on the RTX Pro 6000 Blackwell and up to 1.4x on configurations using two DGX Spark systems. Those gains come from new XQA attention kernels and backend optimisations. The improvements are intended to feed into existing local AI ecosystems rather than requiring NVIDIA-specific applications, with NVIDIA pointing to llama.cpp, LM Studio and vLLM as routes through which users can access the optimisations. NVIDIA PAIR turns multiple PCs into a local AI pool Arguably the most interesting announcement at NVIDIA's IFA briefing was NVIDIA PAIR, short for Personal AI Router. PAIR is software rather than a physical networking device. It is designed to discover computers on a local network and distribute AI inference requests between them, effectively allowing the idle AI compute available across several PCs to be treated as a shared resource. NVIDIA's argument is that modern households frequently contain several powerful computers, but those systems spend much of the day underutilised. Gaming desktops are particularly interesting in this context because their GPUs may offer substantial AI performance while sitting idle whenever the owner is not gaming or running a GPU-heavy application. PAIR installs on each participating PC and discovers devices over mDNS. Machines are paired using a six-digit security code, while communication between devices uses mutual TLS. Instead of splitting a single model across several GPUs, PAIR distributes complete inference requests between available machines. What's the difference you ask? Technologies such as tensor parallelism and model sharding attempt to make several GPUs cooperate on one model, often to accommodate models too large for the memory of a single GPU. PAIR instead routes separate inference jobs between machines. That makes it particularly suited to agentic workloads in which several requests may be running simultaneously. PAIR routes AI jobs without affecting your current work PAIR monitors queue depth and GPU utilisation when deciding where an inference request should run. If the GPU in one machine is busy running a game or rendering a video, PAIR can avoid sending an AI workload to it and instead route the request to another available machine. NVIDIA has designed this around existing local inference tools. PAIR proxies requests sent to Ollama and LM Studio, meaning an AI application can continue behaving as though it were communicating with a normal local inference server. PAIR intercepts the request and decides which machine should actually process it. The software supports Windows, Linux and macOS, making mixed-platform networks possible. One of the examples NVIDIA cited was that of a MacBook using Apple Silicon could participate, provided the underlying inference software supports the hardware. NVIDIA said it has tested configurations containing as many as 18 GPUs, although the company expects a much simpler two-machine setup to be among the more common configurations. Another possible example would be running an AI agent on a laptop while directing inference to a more powerful gaming desktop elsewhere in the house. And to help adoption, PAIR is being released as a beta and is open source under the Apache 2.0 licence. Multi-agent workloads are where PAIR could make the most difference NVIDIA highlighted three major use cases for PAIR: multi-agent workflows, users running several AI sessions simultaneously, and offloading inference when the primary PC becomes busy. Multi-agent systems are particularly well suited to the technology. A primary AI agent increasingly acts as an orchestrator, assigning specific jobs to specialised sub-agents. One might perform research, another might manipulate a spreadsheet, while others handle coding or analysis. These jobs can theoretically execute simultaneously, but when all of them use the same GPU, their inference requests can end up waiting in a queue. PAIR can instead send those requests to different computers. In NVIDIA's demonstration, Hermes Agent spawned five sub-agents that processed household information such as emails and smart-home application logs. Their inference requests were distributed between an RTX 5090 gaming PC, an RTX Spark laptop and a DGX Spark system, with PAIR dynamically routing work according to queue depth and GPU utilisation. NVIDIA said the distributed setup could approach a 2x improvement for this type of multi-agent workload compared with processing the requests sequentially on a single node. The current beta scheduler primarily considers queue depth and GPU utilisation. NVIDIA said more signals will be incorporated into future versions to improve how the software accounts for differences between the performance of individual machines. RTX Spark N1X systems arrive in October NVIDIA also used IFA to reveal considerably more information about RTX Spark, the PC platform it first introduced at Computex. RTX Spark is intended to span gaming, content creation and local AI workloads, with NVIDIA particularly emphasising its ability to run the company's wider AI software stack. The first chip in the family continues to be the N1X. NVIDIA had announced two N1X configurations. The higher-end version combines a 6,144-core Blackwell RTX GPU with a 20-core Grace CPU and between 24GB and 128GB of unified memory. A second configuration combines a 5,120-core Blackwell RTX GPU with an 18-core Grace CPU and between 24GB and 32GB of unified memory. The 6,144-core version will appear in both laptops and compact desktops, while the 5,120-core model is initially intended for laptops. We even got our hands on one of the upcoming RTX Spark machines - Lenovo Yoga Pro 9n and Lenovo Yoga 9n 2-in-1. NVIDIA confirmed that both variants will carry the N1X name. Systems are scheduled to begin reaching the market in October, although regional availability will depend on individual OEM partners. Lenovo and Acer add new RTX Spark designs Not just Lenovo but even Acer also announced RTX Spark hardware for 2026. Like we mentioned previously, Lenovo is joining the platform with the Yoga 9N 2-in-1, while Acer has announced a small-form-factor RTX Spark system. NVIDIA said these machines will become part of the 2026 line-up, with more systems expected from manufacturers in 2027. NVIDIA did not disclose detailed specifications for the individual machines during the briefing, instead directing questions around configurations, design and availability to its OEM partners. The company did, however, reiterate its broader performance targets for RTX Spark, including 1440p gaming, demanding content-creation workloads and thin laptop designs, alongside its emphasis on local AI agents. Gaming support for RTX Spark continues to grow NVIDIA also recapped several RTX Spark gaming announcements made around Gamescom. EA and Embark have announced anti-cheat support for the platform. NVIDIA said that it should allow titles including Apex Legends, F1, Arc Raiders and The Finals to run on RTX Spark systems. Ubisoft is also bringing games from its portfolio to the platform, with Anno 117: Pax Romana among the titles NVIDIA highlighted. Those companies join publishers including Epic Games, Riot Games and Krafton that have already committed to supporting RTX Spark. The anti-cheat announcements are particularly important because compatibility layers and new processor architectures can otherwise run into problems with kernel-level anti-cheat systems, even when the underlying game itself functions correctly. And from the moment that the RTX Spark was announced, these have been one of the most spoken about topics. NVIDIA's AI PC strategy is casting a wider net Taken individually, NVIDIA's IFA announcements cover several different products and software projects. Taken together, however, they point towards a more ambitious idea of what a personal AI computer could become. RTX Spark expands the hardware available for running local models. Hermes, OpenClaw and Perplexity are attempting to make those models easier to deploy. llama.cpp and vLLM optimisations are intended to make them faster. PAIR then expands the available compute beyond the machine directly in front of the user. We honestly found the PAIR announcement to be the most interesting. It takes us right back to the early Folding@Home era. The traditional PC model assumes that applications primarily use the compute resources inside the computer on which they are running. NVIDIA PAIR instead treats nearby computers as an available pool of AI acceleration, while leaving those machines free to continue serving their normal purposes. Whether mainstream users accumulate enough demanding agentic workloads for this approach to become necessary remains an open question. But the company is clearly preparing for a world where a single person may have multiple agents, sub-agents and inference sessions operating at once. At IFA 2026, NVIDIA's message was therefore less about another isolated AI feature and more about constructing the infrastructure around local agents: easier models, faster inference, purpose-built hardware and, now, a way for the PCs already sitting around a home or office to work together.
Share
Copy Link
NVIDIA unveiled major local AI advancements at IFA 2026, introducing PAIR for multi-PC AI routing, one-click setup for Hermes Agent, OpenClaw and Perplexity Computer, plus up to 1.9x faster inference through vLLM and llama.cpp optimizations. The announcements signal a shift toward agent-driven computing on RTX and DGX platforms.
NVIDIA partnered with Microsoft at IFA 2026 to accelerate local AI adoption, unveiling tools that simplify agent deployment and dramatically boost AI inference performance. The announcements center on making AI agents easier to configure while enabling households to pool computing resources across multiple devices. Central to this vision is NVIDIA PAIR (Personal AI Router), software that discovers computers on a local network and intelligently distributes AI inference requests between them
3
. PAIR treats idle GPU capacity across gaming desktops and workstations as a shared resource, with machines pairing through six-digit security codes over mDNS3
. This approach addresses a common scenario where powerful gaming PCs sit underutilized while owners work on less capable laptops.
Source: Digit
Three widely used AI agents will offer simplified local model setup on NVIDIA hardware, eliminating the manual configuration that previously required users to choose models, configure quantization settings, and maintain inference servers
1
. Hermes Agent, developed by Nous Research and used by millions, will provide one-click setup across RTX and DGX systems on Windows and Linux2
. The agent automatically detects NVIDIA GPUs, selects appropriate models, and runs them through integrated llama.cpp with NVIDIA inference optimizations already configured2
. OpenClaw, the largest AI project on GitHub with over 380,000 stars, is launching a Windows app that simplifies local model setup on any RTX GPU with at least 24GB of VRAM2
.
Source: Wccftech
Perplexity Computer, initially available on Linux systems like DGX Spark, is expanding to NVIDIA RTX GPUs with 24+ GB VRAM running Windows in September
2
. The agent enables hybrid AI processing, running workflows locally without consuming credits while selectively escalating tasks to one of 15+ frontier models in the cloud when additional research or reasoning is needed1
. Portable Computer requests permission before sending content to the cloud, keeping sensitive information on local devices1
. Use cases span engineering teams reviewing GitHub pull requests, finance professionals analyzing brokerage statements without uploading documents to chatbots, and startups investigating activation funnel drop-offs1
.NVIDIA collaborated with open-source communities to deliver substantial inference performance improvements across RTX and DGX platforms. Measured on SpeedBench-Coding 8K Throughput with AIPerf, the GeForce RTX 5090 now achieves up to 50% higher token throughput on Qwen3.6-27B and 90% acceleration on Qwen3.6-35B through llama.cpp optimizations
2
. The RTX Pro 6000 Blackwell sees a 20% speedup in vLLM across both Qwen models, while DGX Spark platforms achieve 1.4x faster performance on Qwen3.6-27B and 20% improvement on DeepSeek v4 Flash2
. These gains stem from kernel optimizations, enhanced speculative decoding techniques, faster prefill, and new XQA attention kernels in FlashInfer2
. The improvements are accessible through existing tools like LM Studio and Ollama rather than requiring NVIDIA-specific applications3
.NVIDIA RTX Spark, the company's compact Windows PC platform, is launching in October with new systems from Lenovo and Acer
1
. Electronic Arts, Embark, and Ubisoft are among the game publishers bringing blockbuster titles to RTX Spark1
. The platform aims to provide AI enthusiasts, developers, and creators with capable local agent capabilities in a compact form factor.
Source: NVIDIA
Related Stories
August saw multiple model releases optimized for NVIDIA hardware. Nemotron 3.5 Lightning, a 30-billion parameter model, launched for RTX PCs, RTX PRO Workstations, DGX Spark, and Jetson platforms
1
. Qwen released Qwen3.8-Flash-Next, an open-weight multimodal mixture-of-experts model for DGX Spark and DGX Station, alongside Qwen3.8-27B optimized for local agentic and coding workloads1
. Meta's Muse Glimmer, a 30-billion-parameter model for coding and agentic workloads, can run locally on GeForce RTX PCs with NVFP4 quantization support1
. DeepSeek v4 Flash, a 284-billion-parameter MoE model with 13 billion active parameters, runs on 2x DGX Spark clusters1
.NVIDIA positions these developments as part of a broader shift where PCs evolve from machines requiring manual application operation to systems where users instruct AI agents on desired outcomes
3
. AI-agent-driven workflows combine local LLM inference for private data with cloud models for demanding tasks, allowing tax specialists to research law changes online while processing confidential client information locally, or financial professionals to merge public datasets with data that cannot legally leave their systems3
. Rather than replacing cloud AI, NVIDIA sees local AI as complementary, with agents coordinating applications, tools, and models to accomplish user goals3
.Summarized by
Navi
06 Jan 2026•Technology

01 Jun 2026•Technology

07 Jan 2025•Technology
