On 15 September, Your CDN Decides Who Can Read Your Site
On 1 July, Cloudflare published its second “Content Independence Day” announcement. Most of the coverage focused on the money — Pay Per Crawl becoming Pay Per Use, publishers getting paid when their content shapes an AI answer rather than only when it’s fetched.
That’s the interesting part strategically. It is not the part that will affect you on 15 September.
Buried further down the same announcement is a change to how Cloudflare handles crawlers that do more than one job. It is technical, it is easy to miss, and for a specific group of site owners it means that on 15 September Googlebot stops being able to crawl their website.
If you have ever clicked a button labelled “Block AI bots”, you are potentially in that group. This post is how to find out, in about ten minutes, and what to do about it.
What Cloudflare actually announced
Two things, on 1 July 2026.
First, new controls, live immediately, for every customer including the free tier. Instead of one blunt “block AI bots” switch, Cloudflare now classifies bot traffic by what the bot actually does on your site. Three of those classifications are directly configurable:
- Search — crawling that indexes your content so it can appear in search results later. The behaviour that sends visitors back to you.
- Agent — automation acting in real time on a person’s behalf, right now. ChatGPT’s fetch bot, or Gemini or Claude driving a browser. Usually there is a human waiting at the other end.
- Training — crawling that takes your content to train or fine-tune a model. Your content is absorbed permanently and you get nothing back.
Cloudflare tracks eight further behaviours beyond those three — Transact, Data Collection, Security Testing, SEO, Ads Verification, Social/Link Preview, Feed Fetching, and Monitoring & Operations — but only Search, Agent and Training are configurable by everyone today.
That taxonomy is a genuine improvement. “AI bot” was never one thing, and treating a shopping agent completing a booking for a real customer the same as a scraper harvesting your text for a training run was always crude.
Second, new defaults, arriving on 15 September 2026. This is the part with a date on it, and it has two limbs that affect different people.
Limb one: the ads default
For all new domains onboarding to Cloudflare, Training and Agent will be blocked by default on pages that display ads. Search stays allowed.
Cloudflare’s reasoning is straightforward: an ad is a signal that the site owner intended a human to land on that page and see it. So on those pages, human attention is treated as the goal, and bots that might displace it are kept out.
If you run a service business with no advertising on your site, this limb does nothing to you. If you monetise through display advertising and you are moving onto Cloudflare, it matters a great deal — and it applies to new domains, not to your existing zones.
That is the limb the press covered. It is not the dangerous one.
Limb two: multi-purpose crawlers, and why Googlebot is in scope
Also on 15 September, Cloudflare starts evaluating crawlers that serve multiple purposes against all of their behaviours rather than just one. And the defaults are enforced by the most restrictive rule that applies.
Googlebot crawls for Search. Googlebot also crawls for Training. Under the old model it was categorised as a search crawler and waved through. Under the new model it is both, and if you have told Cloudflare to block Training, the most restrictive applicable rule wins.
Cloudflare names the consequence explicitly: multi-purpose crawlers including Googlebot, Applebot and Bingbot will be blocked for customers who have selected to block Training. That includes people who selected it through the legacy one-click “Block AI bots” service — a button a great many site owners pressed at some point in the last two years, felt good about, and never thought about again.
To be fair to Cloudflare, this is not a trap. They are notifying customers ahead of the date, and there is a setting you can use to opt out of the change. But the notification will arrive by email, to whichever address set up the account, which in a lot of small businesses is a developer who left in 2023.
Six checks to run before 15 September
Check 1 — Are you behind Cloudflare at all?
How: Open your site, then DevTools → Network → click the first document request → Response Headers. Look for a server: cloudflare header, or a cf-ray header. Either confirms it. If you’d rather not open DevTools, a DNS lookup on your domain showing Cloudflare nameservers does the same job.
Why it matters: Plenty of site owners are behind Cloudflare without knowing, because their host put them there. Over 20% of web domains sit behind Cloudflare, so the odds are not trivial.
Pass: No Cloudflare headers. None of this applies to you. Skip to the closing section, which still applies. Fail: Cloudflare headers present. Continue.
Check 2 — Have you ever blocked AI bots?
How: Log into the Cloudflare dashboard, select your zone, and go to Security → Settings. Look at the AI traffic controls and at whether the legacy “Block AI bots” toggle is on.
Why it matters: This is the single check that determines whether 15 September is a non-event or a serious problem for you. If Training is blocked, Googlebot is in scope.
Pass: Training is not blocked, and the legacy toggle is off. Fail: Training is blocked, or the legacy “Block AI bots” toggle is on. Go to Check 3 and treat this as urgent.
Check 3 — Do you actually want to block Training?
How: Answer honestly, in commercial terms rather than principled ones. Does your business depend on people finding you through search and through AI assistants? Or do you sell the content itself?
Why it matters: Blocking Training is a defensible position for a publisher whose text is the product. For a Bedford accountancy practice or a tile retailer, blocking Training buys you almost nothing and, from 15 September, may cost you Googlebot. The asymmetry is enormous.
Pass: You have a specific reason to block Training and you accept the search consequence, having opted out of the multi-purpose change in Security settings. Fail: You blocked it because the button was there and it sounded prudent. Turn it off, or use the opt-out.
Check 4 — What does your robots.txt say now?
How: Visit yourdomain.co.uk/robots.txt in a browser. Look for a Content-Signal: line.
Why it matters: If you enabled Cloudflare’s managed robots.txt, you already have a Content Signals line — typically search=yes,ai-train=no, which says search indexing is fine but training is not. Cloudflare is now adding a fourth field, use=reference, expressing what a crawler may keep and reshare. The three levels are immediate (interact, store nothing), reference (index, excerpt, link back — the default) and full (summarise and reproduce).
Worth being clear about what this is: a signal, not a block. It states your preference. Well-behaved crawlers respect it. Badly-behaved ones ignore it entirely, and no amount of robots.txt will stop them. Enforcement is a separate job.
Pass: Your robots.txt says what you actually mean, and you know whether it is being enforced or merely announced. Fail: You have never looked, or it contradicts what your Cloudflare settings are about to do.
Check 5 — Does your site carry advertising?
How: You know the answer. If you’re unsure whether a third-party widget counts, check whether it serves paid placements.
Why it matters: Limb one only bites on ad-bearing pages, and only for new domains. Most business websites carry no advertising, which makes this limb irrelevant to them — a useful thing to establish so you stop worrying about the wrong half of the announcement.
Pass: No ads, or ads and you’ve made a deliberate choice about Agent traffic on those pages. Fail: You have ads, you’re onboarding a new domain, and you haven’t thought about whether you want shopping and booking agents locked out of your monetised pages.
Check 6 — Are you blocking agents you actually want?
How: In Cloudflare’s AI traffic settings, look specifically at the Agent category, separately from Training.
Why it matters: This is the check most people get wrong, because Agent sounds like the scary one. It isn’t. An Agent visit usually means a real human is sitting there right now, waiting for an answer, a price or a booking. Blocking Agent traffic is closer to unplugging your phone than to protecting your intellectual property.
If you have read our 10-point agent readiness check or looked at what Lighthouse’s Agentic Browsing score measures, you have been working to make your site usable by exactly this traffic. It would be an expensive irony to spend six months on that and then block it at the CDN.
Pass: Agent is allowed, deliberately. Fail: Agent is blocked, and nobody decided that on purpose.
How to read your results
All six pass. Nothing to do. Make a note to re-check after 15 September, because defaults have a way of drifting.
Check 1 fails. You’re not on Cloudflare. The underlying question — what do you want AI systems to do with your content — still deserves an answer, but you have no deadline.
Check 2 fails and Check 3 fails. This is the urgent case, and it is the most common one we expect to see. You blocked AI bots at some point as a precaution, you have no commercial reason to block Training, and on 15 September that precaution starts blocking the crawler your entire search presence depends on. Fix it this month. It is a settings change, not a project.
Check 2 fails and Check 3 passes. You have made a real decision. Use the opt-out in Security settings so the multi-purpose change doesn’t take Googlebot with it, and document why, so the next person to look doesn’t undo it.
Check 6 fails. Separate your thinking about Training from your thinking about Agent. They are genuinely different, which is the whole point of Cloudflare’s new taxonomy — and the reason the old single switch was doing you harm.
The thing worth noticing
The headline story here is about AI, payment models and the future of the open web. The actual risk to your business is that a toggle you clicked eighteen months ago, for reasons you no longer remember, is about to mean something different from what it meant when you clicked it.
That is not really a story about AI. That is a story about configuration drift — the slow accumulation of settings nobody owns, made by people who have moved on, in dashboards nobody opens. The agentic web is going to keep producing changes like this, on dates chosen by infrastructure providers rather than by you. The businesses that cope will not be the ones with the cleverest AI strategy. They will be the ones who know what their own settings currently say.
Ten minutes with your Cloudflare dashboard is worth more this month than any amount of thinking about the future of content licensing.
What to do next
If you’d like someone to run these checks properly — across your CDN settings, your robots.txt, your crawler access and whether AI systems can actually read and use your site — that is part of what our Technical Business Benchmark covers. It’s £750 + VAT, we credit the full amount against any work you go on to commission, and you’ll have the report within 72 hours of our discovery meeting.
If your crawler policy needs keeping current as these standards keep moving — and they will — that’s included in Supercharge Your Site, from £400 + VAT per month.
Or just schedule a call and we’ll tell you in fifteen minutes whether 15 September is your problem or not.
Sources
- Cloudflare, Your site, your rules: new AI traffic options for all customers, 1 July 2026
- Content Signals
- Cloudflare Bot Management documentation on blocking AI bots
Read more