The site is open for people, closed for AI | mrpopular
log in sign up

Shop

home task marketplace cart subscriptions orders

Account

add funds activate promo code affiliate program

Help

support FAQ information reviews blog

Developers

public API reseller API
Dark theme

mrpopularblogSayt Otkryt Dlya…

$0

mrpopularblogSayt Otkryt Dlya…

← Blog

The site is open for people, closed for AI

Aleksandr Dolgopolov
Aleksandr DolgopolovSeptember 10, 2026 · 10 min read
The site is open for people, closed for AI

The most expensive breakage of September doesn't look like a breakage. The site opens, the design is in place, the manager checks it on a phone and sees everything. Meanwhile rankings drop, Google Ads says "Destination not accessible", and the site is gone from AI answers. In a case study published on 7 September 2026 the situation is described literally: in July 2026 a contractor switched on a block for all bots in Cloudflare, Googlebot got caught in it, and ads went down together with organic. Below is who has to be explicitly let back in and why the second channel of losses, scripts instead of links, is completely invisible in reports.

Where the wave of 403s came from

On 1 July 2026 Cloudflare removed the single "block AI bots" toggle and split bots into three categories: Search, Agent, Training. The options are available to everyone, including the free plan. It sounds like an improvement, and it is one, with a single caveat that few people read to the end.

The caveat is that Cloudflare treats some crawlers as mixed purpose: they are search and training at the same time. Googlebot, Bingbot and Applebot are in the mixed group. Any configuration that blocks training, including the old "Block AI bots" option, blocks them too. The warning about this was published on 5 August 2026, but people were clicking those settings back in June and July.

On 4 August 2026 Search Engine Journal described a post from r/SEO: with AI Training = Block plus Bot Fight Mode enabled, Googlebot and Bingbot were getting HTTP 403 when trying to fetch the sitemap, and in the AI Crawlers section of the Cloudflare dashboard both were shown as blocked. In a Cloudflare Community thread a site owner writes that already in early July 2026 verified Googlebot requests were getting 403 with a note in Security Events: "Blocked by 'Block AI training crawlers'".

And then an important date. From 15 September 2026 Cloudflare changes the defaults: for new domains, bots in the Training and Agent categories are blocked on pages where ads are shown, Search stays allowed. So a new domain now arrives with blocking already on, and you have to deal with it before launch, not after a complaint.

  1. 1 July 2026Cloudflare splits bots into Search, Agent, Training
  2. 4 August 2026report of 403s for Googlebot and Bingbot on sitemap
  3. 5 August 2026warning: Googlebot, Bingbot, Applebot count as mixed
  4. 31 August 2026AI surface reports in all Search Console accounts
  5. 15 September 2026new blocking defaults for new domains

Why this stopped being a rare story

A year ago crawler blocking was a topic for a dozen large publishers. Now it's the background. According to Cloudflare, the share of 4xx responses across all crawlers in their network was 35.79% in July 2026 against 14.04% in July 2025, a rise of 21.75 points. Weekly series confirm it: from 11.70% to 16.77% in 2025 and from 34.85% to 37.10% in 2026.

The reason is clear. Cloudflare's June 2026 stats show training crawlers made up 50.6% of model bot traffic, search crawlers only 10.7%, and more than half of the crawls hit pages that hadn't changed since the previous visit. In June 2026 Cloudflare CEO Matthew Prince said bot traffic had exceeded human traffic for the first time. When logs show that half the load is re-downloading unchanged pages into somebody's dataset, the "deny" button presses itself. The only question is what else gets caught.

One more detail that changes habits: bot management has moved from the robots.txt level to the infrastructure level. Cloudflare categories, Content Signals with the use field (immediate, reference, full), cryptographic identification via Web Bot Auth. Since 7 May 2026 Shopify applies stricter limits to bots that don't sign requests through Web Bot Auth. Edits in robots.txt aren't going anywhere, they're just no longer the only place where your site tells a bot "no".

Who should always stay on the allowlist

The list is short, and it's not about AI. These are the bots without which search, ads and product feeds break.

For AdsBot there is an official position, and it settles the argument. The Google Ads help page on "Destination not accessible" names 404 and 403 codes during an AdsBot crawl, a disallow for AdsBot in robots.txt, and server configurations that block access as reasons for disapproval. The recommendation, same page: allowlist the AdsBot-Google and AdsBot-Google-Mobile user agents and make the site accessible from all countries. That last part hurts anyone who closes off half the world with a geo filter at the CDN level: the bot doesn't come from where you expect it.

Google InspectionTool is what powers URL inspection in Search Console. Storebot-Google walks product pages. If you don't let them in, you don't get a ranking drop, you get weirdness: reports are empty, and support says "everything looks fine on our side".

Why a green check in Search Console proves nothing

The same 7 September 2026 case study has one line worth the whole read: a successful URL Inspection check confirms access for the search bot, but doesn't prove access for AdsBot. These are different user agents, and a CDN rule can let one through while cutting the other.

So the order of checks is: first logs by user agent for the last 30 days, then security events in the CDN dashboard, and only then Google's tools. In logs you see response codes, not opinions. If Googlebot comes in and gets a 403, the log says 403, and no green check in the interface overrides that. On my own sites I run this check once a month and after every touch to the infrastructure, mine or someone else's [need a figure: how many times a year I found a stray rule].

What a 41 day experiment showed

The second half of the losses is technical too, but of a completely different kind. Vinicius Stanule, Associate Director SEO at LOCOMOTIVE, ran an experiment for 41 days and wrote it up on Search Engine Land. The site got 11 sections linked with plain HTML and 10 sections where links were inserted through JavaScript. The sitemap returned 404, breadcrumbs and hierarchy panels were removed, so the bot was left with nothing but links.

The result: Googlebot reached 2% of the pages available through JS links. And GPTBot, ClaudeBot, Bingbot, Meta's crawlers and Amazonbot found essentially zero of those pages.

pages behind JS links, Googlebot reached2
same pages, GPTBot, ClaudeBot, Bingbot, Meta, Amazonbot0

This lines up with an earlier measurement by Vercel and MERJ from December 2024: no major AI crawler was recorded executing JavaScript. ChatGPT's crawlers downloaded JS files in 11.50% of requests, Claude's crawlers in 23.84%, but didn't execute them. Downloading and executing are not the same thing, and that difference is exactly where catalogs, filters, click-loaded blocks and infinite feeds disappear.

The practice from here is simple. Open a page, turn JavaScript off in the browser, look at what's left. If an empty frame remains instead of products and text, then for AI bots the page doesn't exist. Links have to be an a tag with an href attribute, not a click handler, and key pages have to be in a sitemap that returns 200.

How to notice you've dropped out of AI answers

The separate trouble is that this drop is invisible in the usual reports. Classic rankings can hold, organic traffic flows along smoothly, and the site is gone from AI answers. A standard monthly report simply has no column for that.

Google already hands over something. On 3 June 2026 Search Console launched Search Generative AI performance reports, and on 31 August 2026 they rolled out to every site in the world. There are impressions, pages, countries, devices and dates. There are no clicks. On that same 31 August 2026 a toggle went live globally that lets you keep content from appearing in Google's AI surfaces, including AI Overviews, AI Mode and generative features in Discover. The reports also gained views for third party platforms, among them Instagram, YouTube and TikTok.

And a limitation worth knowing in advance: generative data lives only in the interface. A check on 11 August 2026 showed that neither the Search Analytics API nor the BigQuery export returns those numbers. So an automated AI visibility dashboard can't be built yet, you'll have to log in by hand and take weekly screenshots. Impressions without clicks, no export, gaps in history: the data exists, but you can't lean on it the way you lean on a normal query report.

Why it's too early to blame a drop on the algorithm

When rankings fall, the first theory is always the same: an update. Let's check the calendar. The last confirmed ranking update is the spam update: launched on 18 August 2026 around 12:30 US Eastern time, completed on 21 August 2026 at 4:50, the rollout took about 2.5 days, global and across all languages, the third spam update of 2026. Everything else that sank in September is more likely technical: bot access, rendering, broken landing pages.

Another September story fits here. Google Data Manager is changing the data collection scheme on the site, a breakdown of the risks came out on 8 September 2026. The meaning is the same: while everyone discusses text quality, the layer nobody looks at breaks, precisely because it always worked.

What to do this week

  1. Open 30 days of logs and find every 403 and 404 response by user agent: Googlebot, AdsBot-Google, AdsBot-Google-Mobile, Google InspectionTool, Storebot-Google.
  2. In the CDN dashboard, check whether training crawler blocking or a bot fighting mode is enabled that catches mixed crawlers. For new domains do this before 15 September 2026, not after.
  3. Turn JavaScript off in the browser and walk through three typical pages: home, category, product or article. Whatever isn't visible without scripts doesn't exist for AI bots.
  4. Make sure the sitemap returns 200 and that important links are a tags with href.
  5. Open the AI surface reports in Search Console, record current impressions by section and set yourself a reminder to record them once a week. There's no export, the history won't build itself.

Blocking training crawlers is a perfectly fine decision, by the way. What isn't fine is learning who you blocked from an email about paused ads. If you want, I'll look at your case through logs and settings, write to support.

Author: Aleksandr DolgopolovSeptember 10, 2026
Aleksandr Dolgopolov
Aleksandr Dolgopolov founder of mrpopular

mrpopular has been running since 2014, and promotion has been in front of my eyes all that time: social networks, search engines, ads, suppliers, orders, disputes, statistics.

A marketing blog without fairy tales. What works, what stopped working, what it costs and why.

About the author →

Read also

TikTok turned comments into a separate channel

TikTok turned comments into a separate channel

On September 3, 2026 TikTok switched on voice comments, polls, carousels of up to nine photos and live photos. Here is what a creator and a brand should do with all this while the format is still empty.

September 9, 20268 min read