Cloudflare AI Crawler Defaults Block Training Bots on Ad Pages
Cloudflare’s September 15, 2026 AI crawler defaults are live: for new domains—and for Free zones that never customized settings—Training and Agent bots are blocked by default on pages
PromptCrates Editorial
Staff Writer

Cloudflare’s September 15, 2026 AI crawler defaults are live: for new domains—and for Free zones that never customized settings—Training and Agent bots are blocked by default on pages that display ads, while Search crawlers remain allowed, according to Cloudflare’s Block AI Bots docs and Content Independence Day blog.
The change also hardens how mixed-purpose crawlers are judged. Bots that combine Search and Training—explicitly including Googlebot, Bingbot, and Applebot in Cloudflare’s examples—follow the most restrictive applicable rule. If Training is blocked, those mixed crawlers can be blocked too, including under the legacy one-click Block AI Bots control that is being deprecated on the same timeline.
Search, Agent, and Training explained
Cloudflare’s July 1, 2026 update reframed “AI bots” into behavior presets every plan can manage, including Free tier customers who previously had fewer levers:
- Search: crawlers that index content to answer questions later, with an expectation of referrals or other equitable compensation back to the site
- Agent: real-time automation acting for a person, such as chat fetch bots and browser-use agents completing a task while a human waits
- Training: crawlers that absorb content to train or fine-tune models, including mixed Search-plus-Training bots under the new taxonomy
Each preset can be set to block on all pages, block only on pages with ads, or allow. Cloudflare’s rationale for ad-page defaults is commercial: ads signal human attention as the monetization goal, so Training and Agent traffic that may siphon that attention is restricted by default, while Search remains the behavior most likely to send visitors back.
That taxonomy matters beyond Cloudflare’s dashboard. Agent traffic is rising alongside tools covered in our reporting on the OpenAI Agents API public beta and broader agent infrastructure such as OmniRoute’s AI gateway surge. Publishers already worrying about distillation and unpaid training—see coverage of Anthropic’s distillation fights involving Alibaba, Moonshot, and DeepSeek—now have finer knobs, but also sharper footguns if they leave legacy training blocks enabled without reviewing mixed crawlers.
Cloudflare also used the July announcement to preview BotBase visibility for Enterprise Bot Management customers and to argue that verified status should mean a bot is honestly labeled and allowable within a category, not automatically welcomed everywhere. Those enterprise features sit beside the Free-tier defaults that most small publishers will feel first.
Why Googlebot warnings matter for SEO teams
The headline risk for marketers is accidental search-engine blocking. Because mixed-purpose bots inherit Training blocks, a site that enabled “Block AI bots” months ago and never revisited Security settings may start refusing Googlebot, Bingbot, or Applebot on ad pages—or sitewide—once the September 15 rule interpretation is enforced for their zone type.
Cloudflare says customers could opt out of the new defaults before the deadline and can still change Search, Agent, and Training independently afterward. The operational checklist is blunt: open Security settings, read the three presets instead of assuming yesterday’s toggle still means what it meant in 2025, and decide whether Training blocks should spare pure search crawlers. Independent explainers and Search Engine Journal-style coverage have hammered the same point for weeks because casual readings of “new domains only” understate Free-tier migration described in related Cloudflare communications.
For teams building AI search indexes of their own, the defaults also reshape crawling economics. Projects that need large open-web corpora—adjacent to indexes discussed in KeenAble’s multi-million page agent search work—will face more zones that refuse Training by default on monetized pages. Expect more robots.txt negotiations, pay-per-crawl experiments, and crawler identity splitting as operators try to keep Search access while disclosing Training separately.
Practical takeaways for publishers and Free-tier sites
First, inventory whether your zone still uses the legacy Block AI Bots toggle and what that implies for mixed crawlers after September 15. Second, decide a deliberate policy for Search versus Training rather than inheriting a one-click choice from last year. Third, monitor crawl stats after the cutover for sudden Googlebot or Bingbot drops on ad templates and article pages that serve display units. Fourth, document the business reason if you intentionally block mixed crawlers: some publishers will accept SEO risk to starve training; others will prioritize discoverability even if models learn from their pages.
Cloudflare positions the shift as Content Independence Day maturity: not all automation is equal, Free customers deserve the same taxonomy as paid plans, and honesty from bot operators should unlock access. Whether that bargain holds depends on how accurately classifications map to real crawler behavior—and how many site owners actually open the settings panel before traffic charts move.
As of mid-September 2026, the authoritative sources remain Cloudflare’s developer documentation, the July 1 changelog on AI traffic options, and the Content Independence Day post. Secondary SEO coverage is useful for operational warnings but should not replace the live dashboard state of each zone.


