Beehiiv Adds Enforced AI Crawler Controls—but Only on Higher Tiers

In its August 20, 2026 AI Discovery release, Beehiiv combined automatic structured data and llms.txt support with AI Crawl Control, which stops selected crawlers at Cloudflare’s network edge. Crawler analytics are available to every publisher, but activating a block requires a custom domain and a Max or Enterprise plan.
The controls themselves predate the wider AI Discovery package. Inc.’s coverage of the release places the original crawler-control announcement in June and says bots operated by Anthropic, Google, Microsoft and OpenAI made more than 500 million requests to Beehiiv-hosted sites in July.
AI Discovery combines visibility with access control

The package puts four technically different mechanisms under one discovery-and-protection story. Structured data and llms.txt help automated systems interpret or locate accessible material; robots.txt asks crawlers to stay away; Cloudflare enforcement can prevent a recognized crawler from retrieving the page.
Every Beehiiv post now receives a machine-readable summary. The platform also adds author, profile and paywall markup, while Max and Enterprise publishers can insert multiple custom JSON-LD blocks for material such as articles, tutorials, videos and FAQs. These fields describe content and attribution, but they do not guarantee indexing, inclusion in an AI answer or a citation.
Beehiiv also generates a dynamic llms.txt file that maps a publication’s content. It can make pages easier for systems supporting the emerging convention to discover, but it is not an access-control list and does not compel a crawler to cite, exclude or avoid training on the material.
The four mechanisms produce different outcomes

The relevant distinction is whether a feature describes content, requests behavior or technically denies access:
- AEO structured data: may improve how an accessible page is interpreted for indexing or an AI response, but it does not guarantee either indexing or citation. It places no restriction on model-training access and has no technical blocking or retroactive effect. Automatic post markup is broadly included; custom blocks are restricted to Max and Enterprise plans.
- llms.txt: provides a publication map for systems that choose to use it. It may support discovery, but citation remains discretionary, training access is unchanged and compliance is not enforced. It cannot affect copies already collected, and Beehiiv’s release does not state a higher-tier or custom-domain requirement for the generated file.
- robots.txt: Beehiiv’s site-wide discoverability switch asks search engines and AI crawlers not to index the publication, with page-level settings available for individual pages. The request can affect indexing and downstream citation when honored, but it is not a network barrier and depends on crawler compliance. It is prospective rather than a way to erase previously retrieved material.
- AI Crawl Control: blocks future requests from selected recognized crawlers before they reach the website. Blocking a training crawler restricts later collection; blocking an AI-search or assistant crawler can prevent the publication from appearing in that product’s answers; blocking a conventional search crawler can reduce ordinary search visibility. This is the only mechanism in the package with Cloudflare-backed technical enforcement, and it requires the Website Builder, a custom web domain and a Max or Enterprise plan.
Publishers can therefore allow an AI-search crawler for possible discovery while denying a separate model-training crawler. The same trade-off appears in Cloudflare’s granular crawler controls, where search, assistant and training access are treated as distinct decisions.
Beehiiv exposes 22 recognized crawlers

Beehiiv’s crawler-control documentation lists 22 recognized crawlers across four categories and states that blocking affects only future access, not material collected before a rule was enabled. The catalog separates model training, AI search, on-demand assistants and conventional search engines.
Training options include GPTBot, ClaudeBot and CCBot. OAI-SearchBot, PerplexityBot and Claude-SearchBot sit in the AI-search category, while ChatGPT-User, Claude-User and Perplexity-User retrieve pages in response to user activity. Googlebot, Bingbot, Applebot and DuckDuckBot are grouped as search-engine crawlers.
The categories matter because blocking does not have one universal outcome. Denying a training bot is intended to stop later training-related retrieval without removing prior copies. Denying an AI-search or assistant bot can reduce citation opportunities in that service, while denying a conventional search crawler prevents index refreshes and can eventually remove pages from the corresponding results.
The dashboard shows total and blocked requests, crawler share, detected crawler types, frequently requested pages and response codes. Access can be changed by individual crawler, but a new bot will not receive its own control until Beehiiv adds it to the maintained catalog.
Plan, domain and timing define the boundary
Network enforcement is limited to publications on Max or Enterprise that use Beehiiv’s Website Builder with a custom web domain. A publication hosted on the default beehiiv.com subdomain cannot activate blocking, even though its publisher can view crawler analytics.
The rules apply only to the custom website domain. They do not govern an email-sending domain, a branded-link domain, newsletter delivery or inbox placement. Saved choices take effect as Cloudflare processes the update.
The controls are also prospective, not a deletion tool. A recognized crawler can be denied later requests, but Beehiiv cannot use the block to retrieve material already collected or remove it from an existing model; any removal request must go to the relevant AI company.
Beehiiv has not set out a timetable for extending enforced blocking to lower-priced plans or default subdomains. The documented boundary remains analytics for all publishers, with active network-level controls reserved for eligible higher-tier custom-domain publications.
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.