Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Technology

Cloudflare Splits AI Crawler Controls, but Googlebot Exposes the Trade-Off

|Updated: |Author: QUASA Editorial Team|5 min read| 479
Cloudflare Splits AI Crawler Controls, but Googlebot Exposes the Trade-Off

Cloudflare’s separate controls for Search, Agent and Training traffic are now live across all customer tiers. The July 1 Cloudflare release replaced a single broad AI-bot decision with policies based on what an automated visitor does, but that distinction does not guarantee an independent outcome for every crawler.

The most consequential part of the rollout is still pending. Cloudflare’s current configuration documentation schedules new defaults for September 15, 2026, when new domains will block Agent and Training traffic on pages displaying ads while continuing to allow Search. The same change will make training restrictions apply to multipurpose crawlers, including Googlebot, Applebot and BingBot.

The controls separate purposes, not necessarily crawler identities

The revised system classifies automated traffic by intended behavior. Search covers collection or indexing used to answer later queries; Agent covers real-time activity performed for a person, such as fetching a page during a chat or completing a browser task; and Training covers collection used to train or fine-tune a model.

This structure gives a publisher meaningful control when an operator assigns different identities to different jobs. A dedicated indexing crawler can be admitted while a separate model-training crawler is rejected, even when the same company operates both.

The boundary appears when one bot performs more than one role. Cloudflare can assign several classifications to the same crawler, so the dashboard presents separate policy choices without being able to split a multipurpose request into its search and training components.

What the pending defaults actually cover

The scheduled defaults are narrower than a network-wide ban on AI automation. They apply to new domains and target Agent and Training requests on pages where Cloudflare detects advertising; Search remains allowed by default. Customers can instead select a zone-wide block, limit blocking to ad-bearing pages or allow the category.

Those choices affect verified bots carrying the relevant classification as well as additional unverified traffic that Cloudflare places in the same behavioral group. Existing customers can configure the policies before the defaults change, and the older Block AI Bots setting is due to be deprecated as part of the transition.

Allow therefore means that the selected category does not add a blocking rule. It does not promise that a request will pass every other applicable policy, security rule or classification. This distinction is central to understanding why a crawler associated with Search can still be denied.

Why a Training block can catch Googlebot

Cloudflare treats Googlebot, Applebot and BingBot as crawlers with both Search and Training purposes. When a site blocks Training, the more restrictive decision will govern the request even if its Search setting remains on Allow.

For publishers, that creates a choice the new taxonomy cannot resolve on its own. They may want conventional discovery while refusing collection for model development, yet a shared crawler identity prevents Cloudflare from granting one purpose and denying the other at request time.

The possible cost is reduced discoverability rather than merely less AI access. A crawler used to discover and refresh pages cannot perform that work when its requests are rejected, and a TechRadar examination of the publisher trade-off identifies declining search visibility as the principal risk of blocking a multipurpose crawler.

That outcome is not automatic for every site that blocks Training. It depends on the requesting crawler, Cloudflare’s classifications, the selected mitigation and whether the request reaches a page covered by the policy. The precise conclusion is that allowing the Search category alone cannot guarantee access for a search crawler that also carries a blocked classification.

Granular policies still depend on bot operators

The new controls are genuinely more precise than one switch when automated services identify their functions separately. Publishers can make different decisions about advance indexing, user-directed retrieval and model development instead of treating every automated visit as equivalent.

For mixed-purpose identities, however, Cloudflare deliberately gives priority to the restriction. Otherwise, an operator could retain access for a prohibited activity by attaching it to an allowed one. The corresponding drawback is that the permitted function may be lost with the prohibited function.

Cleaner separation ultimately requires bot operators to send distinct crawlers for distinct purposes. That would let a site’s policies map to observable behavior and remove the forced choice between training access and search discovery. Until then, the dashboard offers more granular intent than the underlying crawler ecosystem can always enforce.

The lasting change is control with an explicit limit

Cloudflare’s system recognizes that search indexing, real-time agents and model training create different relationships with a website. Giving every customer separate policies for those behaviors is a substantive improvement over a binary AI-bot setting.

The qualification matters just as much: site owners control Cloudflare’s response to classifications, but they do not control whether an outside operator combines several activities under one crawler identity. For Googlebot and other multipurpose crawlers, blocking training can therefore carry a search-access cost even though Search appears as a separate allowed category.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0