You likely assume that a “Verified” label on a crawler grants it default access to your site. That assumption is now obsolete. Under the 2026 updates, “Verified” no longer means “allowed by default.” It simply means the bot is allowable within its specific category: Search, Agent, or Training. If you block a category, even a verified bot in that lane is blocked.
This shift has a direct AEO impact. If an AI answer engine’s crawler is classified as “Training” rather than “Search,” and you block Training, your content disappears from retrievable sources. The bot cannot cite your site because it never accessed it. With Cloudflare AI bot blocking now category-based, visibility in generative answers depends on precise classification, not just verification status. Non-verified bots remain default-blocked, but verified ones are no longer a green light—they are a tag that must match your access rules.
The three lanes: Search, Agent, and Training

Cloudflare has moved beyond a binary view of bot traffic, classifying AI activity into three distinct behavioral lanes. These categories are not just labels; they represent fundamentally different interactions with your content. Search bots collect and index information to answer questions later, typically with an expectation of referral traffic. Agent bots act in real-time on a user’s behalf, such as ChatGPT-User or browser-use agents driving Chrome. Training bots take content to fine-tune models, where data is permanently absorbed into the AI’s architecture.
The critical mechanism here is the “most restrictive rule.” If a crawler is classified as both Search and Training—like Googlebot or BingBot—blocking the Training category blocks the entire bot. This logic has a direct AEO impact: if you block Training, you also block the Search component, making your content invisible to AI answer engines that rely on that indexing. This is why AI crawler management requires a more granular approach than simple blanket blocks.
Starting September 15, 2026, new default configurations will apply to new domains. On ad-supported pages, the default status shifts to reflect this taxonomy:

| Category | Primary Purpose | Default Status (Ad-Pages) |
|---|---|---|
| Search | Indexing for answers | Allowed |
| Agent | Real-time user action | Blocked |
| Training | Model fine-tuning | Blocked |
Understanding these distinctions is the first step in adjusting your Cloudflare AEO settings to ensure the right bots retain access while preventing unnecessary data extraction.
Why the ‘Verified’ redefinition matters for AEO

Under the previous model, a crawler labeled “Verified” received default access to your site, regardless of its specific function. That safety net is gone. In the updated framework, “Verified” now means a bot is allowable only within its assigned category. If a bot is classified for Training, it does not automatically inherit permission for Search.
Consider a concrete scenario involving GPTBot. If this crawler is verified but categorized strictly as Training, and your site blocks the Training category, the bot is denied entry. The “Verified” label offers no override for this block. This shift turns bot verification from a blanket green light into a granular category tag.
For AEO, the consequences are direct. AI answer engines rely on retrieving your content to generate citations. If the crawler serving these engines is miscategorized as Training, or if you have blocked the Agent lane, your content simply cannot be retrieved. You disappear from AI-generated answers not because your content is low quality, but because the retrieval path is blocked. This is where careful AI crawler management becomes a visibility issue, not just a security one.
The role of content use levels
Cloudflare has also introduced a “content use” setting that defines how bots may handle retrieved data. The three levels are:
- Immediate: The bot stores nothing beyond the session.
- Reference: The bot indexes content, creates excerpts, and links back to the source.
- Full: The bot summarizes and reproduces the content in its entirety.
There is a critical penalty here: bots that reproduce content in full are currently ineligible for “Verified” status. Since AEO depends on excerpting and referencing, this restriction directly gates the behavior AI engines need to cite you. If you want to appear in generative answers, ensuring your Cloudflare AEO settings allow the “Reference” level for verified Search bots is essential. Relying on the old assumption that verification guarantees access will leave your brand invisible to the next wave of users.
BotBase: seeing which crawlers actually reach you
Cloudflare AI bot blocking has evolved from a simple on/off switch to a granular visibility tool. BotBase is the new Enterprise feature that lists every verified bot and its specific classification within the Search, Agent, and Training taxonomy. It includes 11 bot classifications, allowing you to see exactly where each crawler sits in the hierarchy.
This changes how you approach AI crawler management. Previously, blocking “AI bots” was often a blanket action taken without full context. Now, you can filter for specific bots to check if those critical to your AEO strategy—such as BingBot or Googlebot—are correctly classified as ‘Search’. If a crawler that should be indexing your content for answers is mislabeled as ‘Training’, you can identify that discrepancy before it impacts your visibility.
The tool also provides detection IDs, which allow for precise targeting in Security rules. This moves your defense from a broad block to a surgical allow/deny strategy. You can copy a specific bot’s detection ID and apply it directly to your rules, ensuring that only the crawlers you intend to permit access are granted it. This precision is essential for balancing security with the need to remain visible in generative AI results.
Setting the right defaults in Cloudflare AEO settings
The new default configuration for Cloudflare AEO settings is built on a simple economic logic: ad revenue depends on human attention, not agent attention. On pages displaying ads, the system now blocks Training and Agent categories by default, while leaving the Search category open. This distinction matters because it protects the ad-supported business model from being drained by bots that consume content without generating the human engagement needed to serve ads.
Auditing your current tier
To ensure your site isn’t inadvertently blocking the wrong crawlers, start by identifying your access level. New options to manage AI traffic based on Search, Agent, and Training are available to all Cloudflare customers, including those on the Free tier. If you are on the Enterprise plan, you also have access to BotBase, which provides a detailed directory of verified bots. Use this tool to filter traffic by specific bots and verify if the crawlers essential for your AEO impact are correctly classified as “Search.”
The opt-out window
A critical deadline exists for those relying on multi-purpose crawlers. Multi-purpose crawlers, such as Googlebot, Applebot, and BingBot, combine Search with Training. Under the new rules, if you select to block Training, these combined bots will also be blocked. If your strategy relies on these specific bots for both indexing and other purposes, you must opt out of the new default configurations in your Security settings before September 15, 2026. Failing to do so may result in a sudden loss of visibility from major search and AI engines that operate under these combined protocols.
Does robots.txt or llms.txt still work after this?
The short answer is yes, but with a critical caveat: these files are preference signals, not hard blocks. A common misconception is that listing a crawler in an llms.txt file guarantees access. That is no longer true under the new Cloudflare AI bot blocking framework. The actual allow or deny decision is determined by the bot’s Verified status and its assigned category within the Search, Agent, or Training lanes.
Consider the risk if you rely solely on your llms.txt file to allow AI crawlers. If a specific crawler is classified as “Training” and you have blocked that category, the platform-level control wins. Your instruction in the file is ignored, and the crawler is blocked. This disconnect between your site’s intent and the platform’s enforcement is the primary AEO impact risk here. You may think you are welcoming citation engines, but they are still locked out because their behavior fits the “Training” or “Agent” definition rather than “Search.”
So, should you delete your llms.txt file? No. Keep it as a best-effort signal for bots that respect these conventions. However, do not treat it as your primary defense or access control mechanism. Instead, ensure your Cloudflare AEO settings explicitly allow the “Search” category for the specific bots you want to cite you. This dual approach ensures that your intent is communicated to polite crawlers while your infrastructure enforces the correct access rights for all others.
The shift in Cloudflare’s AI bot blocking policy marks a distinct phase in AI crawler management. AEO is no longer just about crafting AI-ready content; it requires ensuring the right infrastructure lets the right bots in. The ‘Verified’ label has evolved from a universal green light into a specific category tag, meaning your content’s visibility depends on precise configuration rather than default allowance.
The next practical step is to audit which bots are currently reaching your site. Verify that your Cloudflare AEO settings accurately reflect your actual AEO impact goals, ensuring that answer engines categorized as ‘Search’ are not inadvertently blocked alongside ‘Training’ or ‘Agent’ traffic. It is a routine check, but an essential one for maintaining visibility in generative search. When was the last time you checked if your top AI answer engines were still allowed to read your site?
