4 Sources
[1]
Perplexity's Hybrid Compute splits sensitive tasks between cloud and local AI - Engadget
Back in February, debuted Perplexity Computer. Like Claude Cowork, it's a suite of AI agents that can autonomously complete tasks using the web, as well as files and apps on your PC. Since then, the platform has evolved to encompass a few different products, including Personal Computer for the Mac, and today Perplexity is announcing yet offshoot called Hybrid Compute. The new tools allows you to split a task between a frontier, cloud-based model like Opus 5 or GPT-5.6 Sol and a local LLM running on your computer -- the idea being that the local model can handle any sensitive information so that it remains safe and secure on your machine. Perplexity suggests a few different use cases where Hybrid Compute would be a good fit. For instance, a lawyer might want to prepare a brief that compares the case they're working on against existing case law. In that scenario, the tool would allow them to keep their client's data confidential. Perplexity is also pitching Hybrid Compute as a way for thrifty users to save on inference costs by offloading some of the work from the (typically pricier) frontier models. "This is integrated directly into the Mac app. Any time you try to upload files or send information, we're going to automatically check for sensitive content and make sure that you want to share that data to the cloud," said Perplexity's Jon Staff, who oversees all of the company's Mac products. As part of this release, the company trained a new privacy classifier that will automatically suggest files and information the user should keep on their computer. Before Hybrid Compute starts delegating your task between different models, you'll be able to look over what files it wants to gate away to double check it didn't miss something important. At this stage, you can also decide what models you want to tackle the work. On the local side, your options are Gemma E4B and two flavors of Qwen's 35-billion parameter 3.6 model. One of those variants was post-trained by Perplexity. The company says it will offer more local models in the future. Whatever option you choose, installing a local model does not require you to open your Mac's terminal, and Perplexity's app handles the installation process. Once the system is underway, you'll see a visualization displaying your local CPU, GPU and memory usage. Next to that, a sidebar shows how many tokens the task has consumed. You won't be charged for any tokens a local model generates on your personal machine. Once the system generates an output, you can write follow-up instructions as usual. You can also use an iPhone to queue up tasks. I asked Staff if Perplexity benchmarked the outputs Hybrid Compute produced against that of a fully cloud system. "The short answer is that a fully frontier output is going to almost always be better in terms of raw artifact creation. It's more expensive, it's more capable," he said. However, Staff added that some users don't need access to the best, most powerful AI models to do their work, and in those cases, data privacy and cost might be more important factors. "I think it's got to be a sliding scale, and we want the user to have control over where they are on that sliding scale for the particular type of work they want to do. Obviously we will try to suggest the best thing given the context that we have, but ultimately the user needs to be able to decide what solution fits for them," Staff said. For the time being, Hybrid Compute is only available on Apple Silicon Macs running macOS 15, and Perplexity recommends a machine with at least 32GB of unified memory. The feature is available to Pro and Max subscribers, as well as the company's enterprise customers.
[2]
Perplexity launches privacy-minded 'hybrid compute' AI feature for Mac
Aside from being featured by Apple in the M6 Mac mini launch, Perplexity has been relatively quiet on the Mac front this summer. That changes today with hybrid compute for Mac. Perplexity first announced work on splitting AI tasks between local and cloud models in June. Today, Hybrid Compute is launching as part of Perplexity's Mac app. "Computer now splits a task between cloud models and a local model on your Mac," the company says. "Computer starts each task in the cloud. Trigger one from your iPhone, and your Mac accesses your files and runs sensitive steps locally." Perplexity's website explains how the new feature works: Our on-device PII classifier reads each task on the Mac before it is sent. Names, addresses, and account numbers are swapped for stand-ins, then restored when the answer returns. We've open-sourced the on-device PII classifier, trained in collaboration with Perplexity's Secure Intelligence Institute. Macs with 8GB and 16GB RAM won't be able to run hybrid compute, according to Perplexity. 24GB memory is the minimum, and 32GB is recommended as the floor: Hybrid Compute runs on Apple silicon, macOS 15+ -- 24 GB unified memory minimum, 32 GB for best results. Set up the local model once: download PPLX Qwen 3.8 27B in one click. No Ollama or manual runtime setup. Work it handles uses no cloud credits. No API key. MacStories also has coverage of the new Perplexity feature today. You can learn more from Perplexity's announcement page here. Hybrid compute is just the latest AI tool from Perplexity that takes advantage of the power of the Mac. In May, the company overhauled its Mac app with a whole new version. Apple specifically cites Perplexity Personal Computer as a productivity use case for the new M6 Mac mini as well.
[3]
Your files stay put: Perplexity's hybrid AI keeps confidential data off the cloud
Perplexity today launched hybrid compute for its agentic platform, Computer, a system that lets a single AI agent split its work between frontier models running in the cloud and smaller open-weight models running locally on Apple silicon Macs -- routing sensitive data to the local machine so it never leaves the device. The company says it is the first time an AI agent can begin a task in the cloud and dynamically hand off the confidential portions of that same task to a model running on the user's own hardware, without restarting the job or losing context. The feature becomes available today through Perplexity's desktop app for enterprise customers that opt in, as well as Pro and Max subscribers, on any Apple silicon Mac running macOS 15 or later. "Hybrid is really compelling because it's often the work that requires confidentiality that is the most important to get right, and so the accuracy really, really matters," Jon Staff, who leads Perplexity's macOS and iOS engineering teams, said during a press briefing attended by VentureBeat. "By combining these two together, we can get that maximum intelligence from the frontier models, but we also get the security and the privacy that comes with local." How Perplexity's on-device privacy gate keeps sensitive data off the cloud The architecture works like a dispatcher. A frontier model in the cloud breaks a task into subtasks and routes each one to the appropriate place. Web research, long-horizon planning and heavy reasoning run in the cloud, while anything touching private files, local data or actions on the device gets delegated down to a subagent running on the Mac itself. The linchpin is what Perplexity calls a Privacy Gate: a company-trained classifier that runs on the device and scans for personally identifiable information -- names, addresses, account numbers, secrets -- before anything is transmitted to the cloud. When the gate flags sensitive content, the user chooses whether that portion of the task runs locally or gets shared. "What we wanted to do is make sure anything that's shared to that cloud orchestrator is safe," Staff said. "We built and trained our own PII classifier that integrates directly into the Mac app." He described the handoff in detail: "The cloud orchestration will break down the task based on the prompt and figure out how to route it to different subagents... it's going to delegate that down to a sub-agent running on your Mac, and then that portion of the task is run entirely local. None of those tokens go to the cloud." The economics matter, too, for a company that meters cloud usage through credits. Tokens generated locally cost nothing. "You're paying for the electricity, you're paying for the hardware, so we're not charging you for that," Staff said. "The only thing the credits are used for is the orchestration and the delegation." Lawyers, private equity firms and a founder in an Uber: hybrid compute in action Perplexity built its demonstrations around exactly the kind of work most professionals would never hand to a cloud-only agent. In the first, a lawyer on deadline updated a draft brief against privileged case files stored on a Mac while a cloud agent simultaneously pulled public case law from the open web -- sending out, Perplexity says, only anonymized legal questions. "At no point did their privileged information get shared to the cloud," Staff said. "It never left the Mac." In the second demo, a private equity associate's agent reworked a financial model against confidential management projections, benchmarked the deal against public comparables and produced a fifth iteration of an investment committee deck. The task ran roughly 40 minutes in the background with no human input -- work that would have taken hours of manual stitching between local spreadsheets and cloud research. The third demo emphasized continuity across devices. The founder of a pottery shop, riding in the back of an Uber, kicked off a marketing analysis from her iPhone. Computer asked permission to reach her Mac at the studio, fired up the local subagent to process her customer interviews and revenue data, and combined that with cloud research on competitors' public pricing. "It doesn't matter how far away she is from her computer," Staff said. "Tasks like this aren't possible in a fully local or a fully cloud setup," he added. "You need that security of the local and the privacy, but you also need the intelligence of the frontier." Why a Chinese-made Qwen model on enterprise Macs is raising eyebrows The launch model lineup immediately raised a pointed question. At launch, users can choose among three local models: Google's Gemma E4B, Alibaba's Qwen3.6 35B-A3B, and a Perplexity post-trained version of Qwen3.6 35B -- the company's recommended option. Asked by VentureBeat whether enterprise or government customers had raised concerns about giving a Chinese-developed model access to their machines, Staff argued that local inference neutralizes the geopolitical risk. "The great thing about these models is that they are open weight. We're able to evaluate them ourselves," he said. "When that model is running locally on your computer, the data is not going outside of your computer itself... You're not actually sending those tokens to some cloud provider that's hosted in another country. In fact, all of Perplexity's models are U.S. hosted." He added that macOS's built-in sandboxing framework, known as Seatbelt, constrains what the agent can actually do on a machine: "If local execution is trying to do something that it shouldn't, it'll just point blank stop it and it'll request permission from the user." Perplexity does not currently allow unrestricted "YOLO mode" execution, he said, though "I wouldn't be surprised at some point if we allow certain people to do this." For enterprises, admins can set a single organization-wide sensitivity policy and audit a full record of what leaves each device -- a feature aimed squarely at compliance teams in law, finance and healthcare. Questions remain on the consumer side, however. Pressed on how usage data feeds model training, Staff pointed to Perplexity's incognito mode and a long-standing opt-out toggle, and said enterprise contracts can include zero-data-retention terms. A company spokesperson said Perplexity is "not using it for post training" globally and promised to follow up with specifics on non-enterprise accounts. The enterprise privacy problem hybrid AI is trying to solve The announcement lands amid a broader industry reckoning with a stubborn problem: the most valuable enterprise work involves exactly the data companies are least willing to send to someone else's servers. NIST's generative AI risk profile flags data privacy and information leakage among the technology's central risks, and McKinsey's research on the state of AI has consistently found that organizations struggle to move from experimentation to value capture, with data governance among the chief obstacles. Gartner, for its part, named hybrid computing among its top strategic technology trends for 2025, anticipating architectures that blend compute across environments. Perplexity is betting that the answer is not choosing between cloud intelligence and local privacy, but building the orchestration layer that arbitrates between them in real time. It is a defensible position for a company that has always styled itself as a neutral broker -- "Perplexity is like Switzerland in that we work with everyone," a company representative said at the briefing -- sitting at the application layer above whichever models happen to lead at any given moment. "Anytime one of these gets better, Perplexity gets better," Staff said of the interplay among local models, frontier models and Apple's chips. "That's the really cool nature of where we sit in this application layer, orchestrating all the different pieces together." From $520 million startup to $20 billion agent platform in three years Hybrid compute caps an extraordinarily aggressive product run. Perplexity launched its Comet AI browser in July 2025, initially for $200-a-month Max subscribers -- an early bid to make agents, not chat, the interface to computing. Computer, its full agentic platform, arrived in March 2026, followed by desktop apps for Mac and Windows. Just last week, the company launched a local-first version of Computer on NVIDIA's DGX Spark hardware, which starts on the user's device and escalates to cloud models only with permission. Today's launch inverts that flow: cloud-first, delegating down. The business trajectory has been equally steep. Perplexity was valued at $520 million in January 2024; by September 2025, the company had finalized a funding round at a $20 billion valuation. Along the way it made an audacious $34.5 billion bid for Google's Chrome browser during Google's antitrust remedies fight, and Bloomberg reported that Apple executives held internal talks about acquiring the company -- a striking backdrop for a product now built to showcase Apple silicon. The strategy is not without headwinds. Reuters reported in July that Reddit's data-scraping lawsuit against Perplexity survived a motion to dismiss, part of a wave of copyright and data litigation facing the company -- context that makes its privacy-forward positioning both commercially savvy and reputationally necessary. And practical constraints remain: Perplexity recommends at least 32GB of unified memory for the better tier of local models, Staff was candid that the smallest option "significantly underperforms" the larger Qwen models, and Windows and Linux support will come only later. The deeper question is one users cannot easily inspect. The Privacy Gate is itself a machine learning classifier, and classifiers miss things; a false negative means sensitive data reaches the cloud anyway. Perplexity's answer is transparency -- users can expand and review exactly what the gate flagged before anything is sent, and enterprises get device-level audit logs. But the pitch, at bottom, asks professionals to trust one AI to decide what another AI is allowed to see. For an industry that has spent three years telling lawyers, bankers and doctors to keep their most sensitive work away from the cloud, Perplexity's wager is that the fix was never to build a higher wall -- it was to build a smarter gate.
[4]
Perplexity CEO: Perplexity CEO announces rollout of hybrid compute feature for Mac application
Artificial intelligence firm Perplexity introduced hybrid compute for all users of its Mac application, enabling tasks to be divided between cloud infrastructure and local devices to protect sensitive records. Artificial intelligence firm Perplexity is introducing hybrid compute for all users of its Mac application, enabling tasks to be divided between cloud infrastructure and local devices to protect sensitive records. The announcement comes directly from Perplexity co-founder and chief executive officer Aravind Srinivas, who outlined the company's push toward on-device processing on Tuesday. "We're introducing hybrid compute for all users of the Perplexity Mac app," Srinivas said on X. "This will allow Computer to orchestrate local models that can run locally on Mac, particularly for agent steps involving sensitive and private files (eg your bloodwork, tax returns, litigation, etc)," he added. Srinivas noted the distinction between existing setups and the new mechanism, highlighting technical and privacy considerations across hardware environments. "Most AI agent apps don't differentiate between native and web-based runtimes when using cloud-based agents," Srinivas stated. "Mac offers a big opportunity to move token consumption to Apple Silicon (with no price paid for tokens consumed locally) and protect user privacy," he added. "Hybrid compute combines the best of cloud-based frontier models and a privacy-protecting local runtime." Srinivas also confirmed that the underlying safety infrastructure is being made public. "We're also open-sourcing the PII classifier that we use for deciding when to send the workload to the local model in the hybrid compute setup," he stated. According to the company, frontier reasoning, web searches, and planning execute in the cloud, while local models handle confidential records, protected files, and on-device actions. The division operates through a dedicated privacy gate on the machine, which scans files for personal identifiers such as government identity cards, financial account numbers, and credentials before information moves outwards. This gate can mask details, retain records on the device, decline tasks, or prompt the user for explicit approval. The system addresses regulatory and confidentiality requirements across several commercial sectors, including finance, legal services, and advertising. For instance, an investment team can gather public filings in the cloud while cross-referencing non-public transaction data on the local hardware, whereas legal practitioners can examine case law online while maintaining client-privileged files solely on the local drive. For organizational deployments, enterprise administrators gain centralized controls to establish rules governing file containment, data masking, and user consent, alongside audit mechanisms that track data movements. Users can also issue queries remotely via an iPhone to trigger tasks on their Mac hardware. The hybrid capability is available to Pro, Max, and Enterprise subscribers. The deployment launches with three compact local models: Gemma 4 E4B, Qwen3.6 35B-A3B, and a proprietary Perplexity model. Operational requirements specify an Apple silicon Mac operating on macOS 15 or later, equipped with a minimum of 24 gigabytes of unified memory.
Share
Copy Link
Perplexity launched Hybrid Compute for Mac, enabling AI tasks to split between cloud-based frontier models and local AI models running on Apple silicon Macs. The system uses a PII classifier to keep sensitive data off the cloud while maintaining access to powerful AI capabilities for non-confidential work.

Perplexity has launched Hybrid Compute, a feature that allows its agentic AI platform Computer to split AI tasks between cloud and local models, keeping confidential data off the cloud while leveraging powerful frontier capabilities. Available today for Pro subscribers, Max subscribers, and enterprise customers, the system runs on Apple silicon Macs with macOS 15 or later and requires a minimum of 24GB of unified memory, though 32GB is recommended for optimal performance
1
2
.Perplexity co-founder and CEO Aravind Srinivas announced the rollout, emphasizing how the feature addresses a critical gap in AI privacy. "This will allow Computer to orchestrate local models that can run locally on Mac, particularly for agent steps involving sensitive and private files (eg your bloodwork, tax returns, litigation, etc)," Srinivas stated
4
. The company claims this marks the first time an AI agent can begin a task in the cloud and dynamically hand off confidential portions to on-device processing without losing context or restarting the job3
.At the core of Hybrid Compute sits Privacy Gate, a company-trained PII classifier that scans for personally identifiable information before any data leaves the device. The classifier automatically detects names, addresses, account numbers, and other sensitive content, then prompts users to decide whether that portion of the task should run locally or be shared with cloud-based frontier models
3
."Any time you try to upload files or send information, we're going to automatically check for sensitive content and make sure that you want to share that data to the cloud," said Jon Staff, who oversees Perplexity's Mac products
1
. The company trained this privacy classifier in collaboration with Perplexity's Secure Intelligence Institute and has open-sourced it for transparency2
4
.When Privacy Gate flags sensitive content, it can mask details with stand-ins, keep files entirely on the device, or request explicit user approval before proceeding. For enterprise deployments, administrators gain centralized controls to establish rules governing file containment, data masking, and audit mechanisms that track data movements
4
.Hybrid Compute operates like a dispatcher, with a frontier model in the cloud breaking tasks into subtasks and routing each to the appropriate environment. Web research, long-horizon planning, and heavy reasoning run in the cloud, while anything touching private files or local data gets delegated to a subagent running on the Mac itself
3
.Beyond AI privacy benefits, the system offers significant cost savings. Users pay nothing for tokens generated by local AI models on their personal machines, with cloud credits consumed only for orchestration and delegation tasks. "You're paying for the electricity, you're paying for the hardware, so we're not charging you for that," Staff explained, noting this allows thrifty users to reduce inference costs by offloading work from typically pricier frontier models
1
3
.Srinivas highlighted this advantage: "Mac offers a big opportunity to move token consumption to Apple Silicon (with no price paid for tokens consumed locally) and protect user privacy"
4
.At launch, users can choose among three local models: Google's Gemma E4B, Alibaba's Qwen3.6 35B-A3B, and a Perplexity post-trained version of Qwen3.6 35B, which the company recommends
1
3
. Perplexity says it will offer more local models in the future. Installing these models requires no terminal access, with Perplexity's app handling the entire setup process automatically1
.Once running, users see a visualization displaying local CPU, GPU, and memory usage, alongside a sidebar showing token consumption. Tasks can be initiated from an iPhone, which triggers the Mac to access files and run sensitive steps locally while maintaining cloud connectivity for non-confidential operations
2
3
.Related Stories
Perplexity demonstrated Hybrid Compute through scenarios addressing regulatory and confidentiality requirements in finance, legal services, and other sectors. A lawyer could prepare a brief comparing privileged client data against public case law, with the system keeping confidential information on the Mac while pulling public legal precedents from the cloud
1
3
.In another demonstration, a private equity associate's agent reworked a financial model against confidential management projections while benchmarking the deal against public comparables. The task ran approximately 40 minutes in the background with no human input, work that would have required hours of manual coordination between local spreadsheets and cloud research
3
."Hybrid is really compelling because it's often the work that requires confidentiality that is the most important to get right, and so the accuracy really, really matters," Staff said. "By combining these two together, we can get that maximum intelligence from the frontier models, but we also get the security and the privacy that comes with local"
3
.Staff acknowledged that fully cloud-based frontier models will generally produce superior outputs for raw artifact creation, being both more expensive and more capable. However, he emphasized that many users prioritize data privacy and cost over maximum capability. "I think it's got to be a sliding scale, and we want the user to have control over where they are on that sliding scale for the particular type of work they want to do," Staff explained
1
.This positions Hybrid Compute as a flexible solution for professionals who need powerful AI assistance but cannot compromise on confidentiality. The feature integrates directly into Perplexity's Mac app, which Apple specifically cited as a productivity use case for the M6 Mac mini
2
. For organizations watching regulatory developments around AI and data protection, Perplexity's approach offers a pathway to adopt agentic AI platforms while maintaining compliance requirements and keeping sensitive data under direct control.Summarized by
Navi
[1]
02 Jun 2026•Technology

12 Mar 2026•Technology

12 Mar 2026•Technology

1
Technology

2
Policy and Regulation

3
Health