Cloudflare launches Disallow AI Training to separate search indexing from model data use
Cloudflare has introduced a new security setting called "Disallow AI Training" to resolve the conflict between website discoverability and AI data usage. Previously, site owners using mixed-use crawlers from major tech firms had to choose between allowing their content to be indexed for search or permitting it to be used for training large language models. The new feature allows websites to…
Key points
- Cloudflare introduces "Disallow AI Training" to allow search indexing while blocking AI model training data usage.
- Apple, Google, and Microsoft are designated "Accountable" operators committed to honoring the new training opt-out preference.
- The legacy "Block AI Bots" setting is deprecated in favor of granular Search, Training, and Agent controls.
The company has designated Apple, Google, and Microsoft as "Accountable" operators, meaning they have committed to honoring this distinction. This new control replaces the broader "Block AI Bots" setting, which is being deprecated. Cloudflare also plans to introduce granular controls for AI summaries by early next year, allowing site owners to specify how much of their content appears in generated answers. This move aims to give publishers more agency over how their digital assets are utilized in the evolving AI ecosystem.
Stay discoverable in search while disallowing AI training
blog.cloudflare.com · 16 September 2026
Without proper controls, website owners have long faced a difficult tradeoff: allow your content to be used for AI training, or risk losing discoverability in search. That tradeoff exists because some of the largest organizations on the Internet use mixed-use crawlers: a single crawler serving both search and AI training. Refuse one, and you refuse the other.
Today, Cloudflare is announcing a new Disallow AI Training setting that lets you easily stay indexed for search while refusing to let that same crawler train on your content. Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting.
Mixed-use crawlers were the hard part of the training question. AI Summaries are next. A site-wide yes or no is too blunt: how much of your content appears in a summary matters as much as whether it appears at all. An opt-out for AI summaries is already one of the requirements we've set for mixed-use crawler operators. By early next year, our goal is to let you control how much of your content is included — set once on Cloudflare, rather than with each operator separately.
Why asking isn’t enough
Most site owners want to be found: by humans, agents, and (good) bots. But a significant portion of the open Internet is funded by advertising, subscriptions, or direct relationships with visitors, and those models only pay when someone actually arrives.
Almost every site owner considers Search beneficial: less than 1% of Cloudflare sites choose to block Search bots. Training, however, is a different story: 17% of sites choose to enable some mechanism to block training. This is exactly why we decided site owners needed more granular controls, rather than a one-size-fits-all “Block AI.”
A robots.txt directive alone cannot solve this problem. Anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it.
A network can solve it, however: we publish the preference, identify who is crawling, classify why they are crawling, and block the ones that ignore it – then report what each operator actually does on Radar.
But blocking removes a crawler. It doesn't change how crawlers behave. The better outcome is operators that don't make you choose at all. So since July, we've been talking to them directly. The response has been encouraging: almost all agreed that site owners should have control and transparency into how their content is used, and reassurance that their choices will be respected. To help site owners understand that, we created a designation: Accountable.
The Accountable designation recognizes both capabilities available today and concrete commitments to deliver them. To qualify, a bot operator must meet or commit to meeting the following requirements:
- A mechanism for site owners to opt out of AI training, through robots.txt or a similar standard.
- A mechanism for site owners to opt out of AI summaries set with the operator directly, and next year through Cloudflare (see section below for more detail).
- URL-level visibility into which pages were made available for training, along with metrics showing how content appeared in search.
- Assurance that opting out of AI training will not affect traditional search results.
Apple, Google, and Microsoft all demonstrate that they meet the qualifications to be Accountable. Each combines capabilities available today with time-bound commitments for those still in development. The details of each of these companies’ crawlers are shared below.
New security setting options
Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior. Three behaviors are available as controls:
- Search - crawling to build a search index.
- Training - crawling to train or fine-tune a model.
- Agent - user-directed agents visiting a page on behalf of a human, such as chat fetch bots and browser-use agents.
A mixed-use crawler is a single crawler doing both Search and Training. Without controls, that combination creates the tradeoff described above: site owners cannot refuse one use without refusing the other.
To avoid blocking Accountable mixed-use crawlers — the ones that don't force that tradeoff on website owners — we are introducing a new setting: Disallow AI Training. Disallow AI Training is named for the Disallow: directive it publishes in your robots.txt.
“Block” setting now means something different
Block and “Block on pages with ads” previously did not apply to mixed-use crawlers because blocking them could also affect search discoverability. Now that we have the new Disallow AI Training setting, Block and “Block on pages with ads” apply to all training crawlers, including mixed-use crawlers.
Training, Search, and Agent controls are applied at the domain level. With the addition of Disallow AI Training, the available settings are:
- Allow : All crawlers are allowed, unless blocked by another setting or a WAF rule.
- Disallow AI Training : Bot Preference Sync publishes the applicable no-training preference in robots.txt. Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search. Disallow AI Training is only available as a setting for Training, not Search or Agent.
- Block on pages with ads : Crawlers, including mixed-use crawlers, are blocked only on pages detected to be serving an ad.
- Block : All crawlers, including mixed-use crawlers, are blocked.
Disallow AI Training works by publishing a preference in robots.txt. An ads-only preference cannot be expressed that way: Cloudflare can detect which pages serve ads, but that list is too large and changes too frequently to enumerate in robots.txt. That's why there's no Disallow AI Training on pages with ads.
Agents do not create the same search-discoverability tradeoff as mixed-use crawlers, and the Internet does not yet have a well-established directive for expressing Disallow preferences to agents. For now, we’re not including a Disallow setting for Agents. As standards such as ai-prefs mature, we will revisit this approach.
What changes on September 15?
We are making the following changes to Bot Management and AI Crawl Control:
- Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training. To stop training and keep search, use Disallow AI Training.
- “Block AI Bots” will be deprecated in favor of the more granular Search, Training, and Agent controls.
- Managed Robots.txt will be deprecated in favor of Bot Preference Sync. Customers who enabled Managed Robots.txt will migrate to the new system.
- Disallow AI Training will become part of the recommended configuration for certain new domains.
- Existing customers will have their preferences migrated to the new controls as described below.
What you need to do
Nothing, in almost every case. Your current settings carry over on their own.
If you want mixed-use crawlers gone entirely, you now have to say so. Select Block. It will stop Applebot, Bingbot, and Googlebot from reaching your site — search included.
Existing domains that never used the Search/Training/Agent controls
Site owners that never configured the more granular controls will be migrated to the new settings based on their legacy Block AI Bots setting:
Existing domains that previously configured the Search/Training/Agent controls
For domains that previously configured the granular controls, we will preserve the practical effect of their selections under the new definitions. Previous Training selections of Block or Block on pages with ads will migrate to Disallow AI Training.
Recommendations for new domains
Beginning September 15, customers onboarding a new domain will be offered one of two preset configurations, depending on whether the site earns money from advertising. Ad revenue depends on a human actually seeing the page. Training replaces that visit with an answer; agents fetch the page with nobody there to see the ads. So the presets for ad-supported sites are more restrictive. You can change any of these settings during onboarding, or at any time afterward.
This text was published by blog.cloudflare.com . It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
2 sources- Hacker News discussion · 41 points news.ycombinator.com
- Cloudflare Lets Sites Disallow AI Training Without Blocking Googlebot Press · searchenginejournal.com ·
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Policy & Regulation
All →- Meta may cap employee AI token usage as AI-driven costs rise · 1 src
- Critics warn AI lab coordination proposals risk antitrust violations and safety · 15 src
- Anthropic CEO urges government oversight as company eyes IPO · 29 src
- OpenAI unveils framework for public disclosure of AI misalignment incidents · 4 src
- Critics label Big AI's proposed development slowdown a market lockout strategy · 14 src
Comments
via GitHub Discussions