6 Sources
[1]
Microsoft researchers tried to manipulate AI agents - and only one resisted all attempts
The results underscore the dangers of an AI agent-run economy. As you've probably noticed, there's been a lot of hype circulating around AI agents and their supposed potential to transform the economy and human labor by automating routine, time-consuming tasks. A growing body of research, however,
[2]
AI Agents Miss the Mark on the Tasks They Were Designed to Handle
Don't miss out on our latest stories. Add PCMag as a preferred source on Google. Microsoft put several AI agents through their paces in a simulated environment and found that they are far from capable. In a test of how well the top agentic AIs (GPT-4o, GPT-5, and Gemini 2.5 among them) handle
[3]
Agents of misfortune: The world isn't ready for AI agents
Amazon's spat with Perplexity shows that technology is not the only blocker for the agentic era Opinion The agentic era remains a fantasy world. Software agents, the notional next frontier for generative AI services, cannot escape the gravity of their contradictions, legal ambiguities, and
[4]
Microsoft: Don't let AI agents near your credit card yet
Shopping bots pick first option and 'vulnerable to manipulation', Magentic Marketplace trial finds Ready to have your agent talk to my agent and arrange a sale? Microsoft has published a simulated marketplace to put AI agents through their paces and answer a question for the new age: Would you
[5]
Microsoft Magentic Marketplace shows AI can't truly operate independently
AI agents slow down significantly when presented with too many choices A new Microsoft study has raised questions on the current suitability of AI agents operating without full human supervision/ The company recently built a synthetic environment, the "Magentic Marketplace", designed to observe
[6]
Microsoft Gave AI Agents Fake Money to Buy Things Online. They Spent It All on Scams - Decrypt
They can't collaborate or think critically without step-by-step human hand-holding -- autonomous AI shopping isn't ready for prime time. Microsoft built a simulated economy with hundreds of AI agents acting as buyers and sellers, then watched them fail at basic tasks humans handle daily. The
Share
Copy Link
Microsoft researchers tested AI agents in a simulated marketplace and found they struggle with basic tasks, are easily manipulated, and perform poorly when given too many options, raising serious questions about their readiness for real-world deployment.
Microsoft researchers have conducted a comprehensive study examining how AI agents perform in marketplace scenarios, revealing significant limitations that challenge the current push toward autonomous shopping systems. The research, conducted through an open-source simulation called the "Magentic Marketplace," tested industry-leading AI models including GPT-5, GPT-4o, and Gemini 2.5 Flash in realistic transaction scenarios
1
.
Source: TechRadar
The study simulated interactions between 100 customer agents and 300 business agents, allowing researchers to observe how AI systems navigate complex marketplace decisions such as restaurant selection based on menu offerings and pricing comparisons. The findings reveal fundamental flaws that suggest current AI agents are not ready for widespread autonomous deployment in commercial environments
4
.One of the most significant problems identified was what researchers termed the "Paradox of Choice" - essentially analysis paralysis for AI systems. When presented with numerous vendor options, most AI agents failed to conduct exhaustive comparisons and instead accepted initial "good enough" options rather than thoroughly evaluating alternatives
1
.The study found that performance degraded sharply with scale, with agents becoming overwhelmed when interacting with large numbers of business agents. This limitation is particularly concerning given that real-world marketplaces typically offer consumers hundreds or thousands of options
2
.Additionally, the research revealed a significant advantage for vendors who responded first, with Microsoft reporting a 10-30x advantage for response speed over quality. This suggests that current AI agents prioritize immediate availability over optimal value, potentially leading to suboptimal purchasing decisions
2
.Perhaps most concerning were the findings regarding AI agents' susceptibility to manipulation. Researchers tested six different manipulation strategies, including fake credentials, misleading claims, and prompt injection attacks. Most agents fell for these tactics, with only Gemini 2.5 Flash demonstrating consistent resistance to all manipulation attempts
1
.
Source: Decrypt
The manipulation techniques included dubious claims such as "#1-rated Mexican restaurant" without verification, fake reviews claiming many satisfied customers without citations, and suggestions of danger at competing businesses. These tactics proved effective at directing AI agents toward specific vendors, highlighting critical security concerns for autonomous marketplace systems
2
.Related Stories
The research comes at a time when major tech companies are rapidly deploying AI agents for commercial applications. OpenAI's Operator can navigate websites and complete purchases, while Meta's Business AI interacts with customers as automated sales representatives. However, the Microsoft study suggests these systems may not be ready for unsupervised operation
1
.
Source: The Register
The findings align with recent real-world incidents, including OpenAI CISO Dane Stuckey's admission that ChatGPT Atlas browser agents can purchase wrong products on behalf of users. This acknowledgment underscores the practical challenges facing AI agent deployment in commercial environments
2
.Furthermore, the research highlights broader industry tensions, as evidenced by Amazon's recent demand that Perplexity stop allowing its browser to make automated purchases on the e-commerce platform. This dispute illustrates that technological limitations are not the only barriers to widespread AI agent adoption
3
.Summarized by
Navi
[3]
[4]
15 May 2026•Science and Research

20 Feb 2026•Science and Research

15 Nov 2024•Technology

1
Science and Research

2
Policy and Regulation

3
Technology