GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Cloudflare's Second Content Independence Day: Search, Agent, and Training Bots Now Get Different Rules

Cloudflare’s second annual Content Independence Day post, published July 26, 2026, replaces the blunt “block AI bots” toggle it launched a year ago with a three-category bot taxonomy. Website owners running on Cloudflare can now distinguish between search crawlers, agent browsers, and training scrapers, and set different access policies for each. A new ad-page protection mode is also available to all customers.

Last year’s announcement was reactive: AI was “taking everything and sending back nothing,” Cloudflare’s post says, and the industry needed a defensive tool fast. The one-click block and a Pay-Per-Crawl marketplace addressed the most egregious version of the problem. What surfaced over the following year was that blanket blocking creates its own traps.

The Problem With One-Size-Fits-All

The new post identifies the core tension for small publishers: “If you run a small site, the problem isn’t just that someone could train models on your content — it’s that nobody can find you in the first place. So you have to make a Faustian bargain: either show up in search and let AI train on you, or risk losing discoverability.”

That framing explains why the new approach moves from binary to categorical. The three categories are operationally meaningful:

Search bots — crawlers whose primary function is indexing content for retrieval. Publishers typically want these; they drive referrals. The concern is when the same bot that indexes is also training on the content.

Agent bots — automated processes executing multi-step tasks: filling forms, extracting data, operating interfaces. These are the AI agents interacting with web infrastructure directly. A publisher selling something may want agents to transact; a news site may not want agents scraping and summarizing.

Training bots — crawlers whose purpose is feeding model pre-training or fine-tuning pipelines. This is what last year’s block-all was primarily aimed at. The new framework lets publishers allow search while blocking training, which was not cleanly possible before.

Ad-Monetized Page Protection

The new ad-page protection feature addresses the revenue cannibalization problem directly. Publishers relying on display advertising earn revenue when humans visit pages. When AI agents visit the same pages — to summarize, extract, or train on content — no ad impression fires. The publisher bears the bandwidth cost with no revenue return.

Cloudflare’s solution lets publishers flag high-value ad-monetized pages and apply stricter access controls to agent and training traffic specifically on those URLs, while leaving search indexing intact. The feature is included at all customer tiers.

Why Cloudflare Holds This Position

Cloudflare handles traffic for a large fraction of the web. Its Radar data from June 2026 put bot traffic at 57.5% of all web requests — the majority of web activity is now automated. That traffic flows through Cloudflare’s infrastructure, which gives it both visibility and enforcement capability that individual site operators lack.

The classification challenge Cloudflare’s post identifies — “We could debate the cutoff for what qualifies as ‘AI’ today, just to find that the standard changes tomorrow” — is why the new taxonomy focuses on bot behavior rather than AI identity. A bot storing content and resharing it in AI summaries is functionally different from a bot indexing for search results, regardless of whether the underlying model is technically “AI.”

The Pay-Per-Crawl marketplace from last year remains available alongside the new controls. For publishers who want to monetize AI access rather than block it, the metered-access path is still there.

Key Numbers

  • Cloudflare’s second Content Independence Day: July 26, 2026
  • Bot share of web traffic: 57.5% (Cloudflare Radar, June 2026)
  • New taxonomy: Search bots, Agent bots, Training bots — each with configurable policies
  • Ad-page protection: available to all customers, agent/training specific
  • Pay-Per-Crawl marketplace: continues from Year 1 for monetization path