Simplytics

AI crawlers vs. AI referrals: what your analytics actually sees

On July 1, 2026, Cloudflare announced that from September 15 it will block AI "training" and "agent" bots by default on ad-supported pages for new domains. If that made you open your analytics to check how much of your traffic is AI, here's the thing worth knowing first: the AI activity in the headlines and the AI numbers in your dashboard are usually not the same thing. One is bots reading your pages. The other is people arriving from a chatbot's answer. Your analytics tool measures the second and mostly can't see the first — and confusing the two is how people end up with wildly wrong ideas about their traffic.

The short version: AI crawlers fetch your content; AI referrals send you visitors. A crawler is a bot pulling your pages to train a model or answer a question inside a chatbot. A referral is a real person who clicked a link in ChatGPT, Perplexity, or Google's AI answers and landed on your site. Most analytics tools — Simplytics included — are built to count the second and to reject the first. So the "AI is eating my traffic" story and the "AI is sending me traffic" story live in different systems, and you need both lenses to read the moment correctly.

The two layers, side by side

AI crawlers AI referrals
What it is Bots fetching your pages (ClaudeBot, GPTBot, PerplexityBot, Googlebot, CCBot…) A person clicking through from an AI answer to your site
Purpose Training models, indexing, or answering a question in-chat A real visit — reading, buying, signing up
Where it shows up Your server logs, CDN, and Cloudflare's bot reports Your analytics dashboard, as a referral
Does your analytics count it? No — declared bots are filtered out before they're recorded Yes — if the visit carries a referrer
What Cloudflare's July 2026 move affects This layer — who's allowed to crawl Not directly; people can still click through

Keep that table in mind and most of the confusion dissolves. Cloudflare's Content Independence Day is about the crawler layer — who gets to read your content and on what terms. It splits automated AI traffic into three buckets — Search (collecting or indexing content to answer questions later), Agent (acting in real time on a person's behalf), and Training (crawling to train or fine-tune a model) — and from September 15, 2026, new domains on Cloudflare will block Training and Agent bots by default on ad-monetized pages, while leaving Search allowed. (Existing sites can opt out before then in their security settings.) None of that changes what lands in your analytics dashboard, because your dashboard was never counting those crawlers in the first place.

Why the crawler numbers look enormous — and why they're not "visits"

The reason this distinction matters so much in 2026 is the sheer size of the gap between the two layers. Cloudflare publishes a "crawl-to-refer" ratio — how many pages a company's bots crawl for every one visitor it sends back — and the numbers are lopsided in a way that surprises people the first time they see them:

AI company Pages crawled per visitor referred back
Anthropic ~38,000 : 1
OpenAI ~1,090 : 1
Perplexity ~195 : 1
Google ~5.4 : 1

Source: Cloudflare's crawl-to-click gap analysis, July 2025 data. Cloudflare's own real-time AI Insights on Radar show the gap has persisted through 2026; exact ratios move week to week. Google's far lower ratio reflects that its crawling is still mostly traditional search indexing — which sends real clicks back — not AI training.

Cloudflare's data also found that, across the past year, roughly 80% of AI crawling was for training, about 18% for search, and only ~2% for real-time user actions. So the overwhelming majority of "AI traffic" in the infrastructure sense is bots ingesting content — not humans on their way to your site. If you mistook the crawler volume for visits, you'd think AI was flooding you with audience. It isn't. It's reading, at industrial scale, and sending a trickle of people back.

That trickle is the part your analytics can actually count — and it's the part worth caring about, because those are real readers and customers.

Why so many of the real AI visits still hide in "Direct"

Here's the frustrating twist, and it's the same on every analytics tool: even the genuine human clicks from AI assistants frequently show up unattributed. When someone follows a link out of an AI chat, the referrer header that would say "this came from ChatGPT" is often missing. Cloudflare noted this directly — visits referred by Claude's native app don't include a Referer: header at all — which is one reason the crawl-to-refer ratios overstate the gap, and exactly why so much real AI traffic quietly piles up in Direct.

This is not a flaw you can configure away, and it hits GA4 and privacy-first tools alike: no referrer on the request means no tool can honestly say where the visit came from. It's the single biggest reason your "AI referrals" number understates reality. We covered the attribution side of this — including how GA4 files Perplexity under Referral and buries stripped-referrer AI clicks in Direct — in what GA4's AI Assistant channel still hides. The visits Simplytics can attribute show up in a dedicated AI channel (ChatGPT, Perplexity, Claude, Gemini and others); the ones that arrive bare fall to Direct, same as everywhere. We're honest about that limit rather than papering over it. The mechanics of that AI channel are in seeing ChatGPT, Perplexity and Gemini traffic in your analytics.

What Simplytics does with the crawler layer

So if the crawlers aren't visits, do they pollute your numbers? For a well-behaved analytics tool, no — and it's worth being precise about how, because "we block bots" is a claim every vendor makes and most explain badly.

Simplytics never lets a declared crawler become a recorded visit. Every hit enters through one edge endpoint we control, and before anything is stored, the request's user-agent is checked against a bot list. The AI crawlers making all those headlines announce themselves in that user-agent — ClaudeBot, GPTBot, PerplexityBot, Googlebot, CCBot, Bytespider — and any user-agent containing bot, crawler, or spider is rejected with a 400 and recorded nowhere. So the training and indexing traffic that dominates Cloudflare's charts simply doesn't appear in your Simplytics dashboard. Your visitor count stays a count of people.

Now the honest caveat, because this blog exists to make them: this is user-agent matching, and user-agent matching is not a wall. A crawler that disguised itself with an ordinary browser user-agent — or a brand-new one that doesn't put bot in its name — could slip past. No client-side analytics tool can perfectly separate a determined, well-disguised bot from a person; anyone claiming otherwise is overselling. What we can say plainly is that the declared AI crawlers, which are the ones in the news and the vast majority of the volume, are filtered out at the door on infrastructure we control, before they can touch your numbers. (This is the same single-endpoint, check-before-record design that makes fabricated hits harder to inject, which we dug into in why you can't filter GA4 ghost traffic.)

So what should you actually do?

The bottom line: AI crawlers and AI referrals are two different layers, and July 2026's Cloudflare change is about the crawler one. Your analytics tool should keep the crawlers out of your visitor counts entirely — Simplytics rejects declared bots at a single edge endpoint before recording anything — while showing you the human clicks AI genuinely sends, in a dedicated AI channel, and being upfront that referrer-less AI visits land in Direct for everyone. You get that honest, cookie-less picture — visitors, channels, countries, events — for a lot less money than the rest: $1/month versus the $9–15/month competitors.

← All posts · Compare Simplytics with alternatives