Cloudflare Introduces AI Crawling Controls as 36.6% of Traffic Now Mixed-Use Bots

5 Sources

Share

Cloudflare launched new controls separating search, AI training, and AI agents access for website owners. The company blocks AI agents from ad-supported pages by default while Apple, Google, and Microsoft earned Accountable designation by honoring no-training preferences without penalizing search rankings.

News article

Cloudflare Splits AI Bot Controls Into Three Independent Settings

Cloudflare introduced an Accountable designation for AI crawling alongside a new Disallow AI Training setting that allows website owners to refuse AI training while maintaining full search visibility

1

2

. The company replaced its single "Block AI Bots" switch with three independent controls covering search indexing, AI training, and AI agents

5

. This change addresses a critical problem: mixed-use crawlers that collect content for both search indexes and AI training now make up 36.6% of verified crawler traffic on Cloudflare's network, the single largest category

2

. Until now, website owners who wanted to block AI training had to sacrifice search visibility entirely when dealing with these mixed-use crawlers

3

.

Apple, Google, Microsoft Earn Accountable Status for Honoring Website Owner Control

Apple, Google, and Microsoft have earned Cloudflare's Accountable label by meeting or committing to meet four specific criteria

2

. These requirements include giving site owners a clear opt-out method through robots.txt, allowing opt-outs from AI-generated summaries, providing URL-level visibility into content usage, and publicly confirming that opting out of training will not affect search rankings

5

. Google offers Google-Extended, a control enabling sites to opt out of training without leaving Search

2

. Matthew Prince, co-founder and CEO of Cloudflare, stated this approach preserves the openness that makes search valuable while giving content creators meaningful control over how their work is used

2

3

.

Ad-Supported Pages Face Default AI Agent Blocks

Cloudflare now blocks AI agents from ad-supported pages by default, reasoning that an advertisement proves a page was built for human visitors

1

. When an AI agent fetches a page on a shopper's behalf, it reads prices and reviews but leaves ads unseen, undermining the revenue model

1

. New websites on Cloudflare receive tailored recommended settings: sites carrying advertising get search crawling enabled, AI training disallowed, and AI agents blocked on ad-carrying pages

4

5

. All other sites default to allowing search, AI training, and AI agents, consistent with how most non-ad-supported websites operate

2

. The default applies to new domains, new sites from existing customers, and free-tier accounts, reaching approximately a fifth of the web

1

.

Crawler Traffic Patterns Drive Timing of Control Changes

Cloudflare's July bot report revealed 52% of crawler requests were for AI training as of June, up from 22% in spring 2025

1

. Mixed-use crawlers blending search, agent use, and training made up more than 36% of activity, while pure search crawling represents a shrinking share

1

. Current data shows fewer than 1% of website owners block search crawlers, while 17% restrict AI training

2

4

. This gap between search and training restrictions explains why Cloudflare split the single control into three separate switches

1

. The company also introduced Bot Preference Sync, replacing its "Managed Robots.txt" feature to let site owners set crawling preferences once across all supported crawlers automatically

2

5

.

Retailers Navigate AI Agent Economics Differently Than Publishers

The block falls hardest on review sites, comparison guides, and publisher content that shopping AI agents read before reaching checkout

1

. Retailers face different economics: a merchant may prefer fewer visitors who convert at higher rates, with AI referrals converting at 3 to 5 times the rate of traditional search

1

. Amazon sued Perplexity in November to stop its Comet browser agent from shopping inside customer accounts, accusing the startup of hiding when a bot was acting for a real person

1

. Walmart and Target tested the opposite approach, exploring ways to work with AI shopping platforms while maintaining their transaction role

1

. Nearly 132 million U.S. adults have bought retail products with AI assistance, and 59% of AI-assisted purchases still end at Amazon

1

. Only 23% of merchants can identify both AI-driven traffic and resulting purchases, while another 21% recognize agent traffic but cannot connect it to sales

1

.

Cloudflare Sets Aggressive Targets for AI Summary Control

Cloudflare aims to drive mixed-use crawler traffic on its network to zero within a year

1

. By early next year, the company plans to let site owners control how much of their content appears in AI summaries in one place, rather than setting preferences with each operator individually

2

5

. Every operator Cloudflare designates as Accountable must give site owners a way to opt out of AI summaries

2

. Cloudflare is participating in open standards work at the Internet Engineering Task Force (IETF), including "ai-prefs," a specification under development that would allow any website to express AI access preferences in a standardized, portable way

2

5

. The company previously launched Kitesurf in August, a cloud-hosted browser built for AI agents rather than people, showing it serves both sides of the AI access debate

1

. Nearly 90% of organizations report managing bot activity as a challenge, with weak controls allowing malicious bots through while blocking legitimate users

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved