AI Crawlers: Search OK, Training NO
Alps Wang
Sep 16, 2026 · 1 views
Reclaiming Control Over AI Data
Cloudflare's new 'Disallow AI Training' setting is a crucial step towards empowering website owners with granular control over their content's use in AI training, directly addressing the inherent conflict between search discoverability and data privacy. By partnering with major players like Apple, Google, and Microsoft, Cloudflare is establishing a de facto standard for 'accountable' AI crawlers, ensuring that publishers can maintain their search presence while opting out of training. The distinction between 'Search,' 'Training,' and 'Agent' behaviors is a sophisticated approach to bot management, moving beyond blunt blocking mechanisms. The introduction of 'Accountable' as a designation for crawlers that respect these controls is particularly noteworthy, fostering transparency and trust in the ecosystem.
However, the implementation details reveal some nuances. While Apple, Google, and Microsoft are largely covered, Bingbot's current reliance on the NOARCHIVE meta tag for opting out of training, with robots.txt support targeted for 2027, presents a temporary gap. This means that for Bing, the 'Disallow AI Training' setting in Cloudflare might not immediately convey the intended preference via robots.txt, requiring manual intervention or reliance on older methods. Furthermore, the blog post touches upon AI summaries as the 'next frontier' but stops short of detailing immediate granular controls for them, focusing instead on an opt-out for now. The long-term implications of how AI summaries might cannibalize website traffic, despite potentially higher conversion rates for referred users, warrant deeper exploration and more immediate, fine-grained control mechanisms for content inclusion in summaries.
The implications for developers and site owners are significant. This feature simplifies complex bot management, allowing for a more strategic approach to AI data utilization. The 'Accountable' designation provides a clear framework for evaluating crawler behavior. For businesses, especially those reliant on advertising revenue, this offers a vital tool to protect their monetization models while still benefiting from search engine visibility. The move towards standardized directives for AI preferences, like the nascent ai-prefs, is a positive sign for the future of web standards in the AI era. Cloudflare's initiative is instrumental in shaping this future by actively engaging with major crawler operators and driving adoption of these principles.
Key Points
- Cloudflare introduces a 'Disallow AI Training' setting to allow website owners to prevent AI training while remaining discoverable in search.
- Major crawlers from Apple, Google, and Microsoft have committed to honoring this setting.
- The new system categorizes bots into 'Search,' 'Training,' and 'Agent' behaviors, enabling granular control.
- A new 'Accountable' designation recognizes crawlers that offer transparency and respect site owner preferences.
- Existing 'Block' and 'Block on pages with ads' settings now apply to mixed-use crawlers, impacting search discoverability.
- Bingbot currently has a temporary gap, with full robots.txt support for AI training opt-out expected in 2027.
- Future focus includes more granular controls for AI summaries.

📖 Source: Have it both ways: stay discoverable in search while disallowing AI training
Related Articles
Comments (0)
No comments yet. Be the first to comment!
