Skip to content

AI Crawlers: How Much They Take for Every Visit They Send Back

The crawl-to-refer ratio shows what you actually get for the AI crawler load. How to find your own number, how to read it, and when it matters.

~800 words 6 common questions ~4 min read Updated: 2026-07-19
AI Crawlers: How Much They Take for Every Visit They Send Back
Quick answer

The crawl-to-refer ratio tells you how many pages an AI platform crawls on your site for every visitor it sends back. In Cloudflare network data for July 2025 it ranged from 5.4 to 38,065.7 depending on the operator. Roughly 80 percent of that crawling goes to model training.

The crawl-to-refer ratio tells you how many pages an AI platform crawls on your site for every visitor it sends back. In Cloudflare network data for July 2025 it ranged from 5.4 to 38,065.7 depending on the operator. Roughly 80 percent of that crawling goes to model training.

AI crawler load usually gets discussed in terms of volume: how many requests, how much data transferred. This article is not about server performance or hosting costs. It’s about a metric that’s more useful for making decisions: what you get back for that load.

What the crawl-to-refer ratio measures

Cloudflare defines it as the number of pages a platform crawls compared with how often it sends a person to your site. The higher the number, the more content leaves for every visit that comes back.

Aspekt Ratio Change since January 2025
Google 5.4 +43%
Microsoft 40.7 +5.7%
Perplexity 194.8 +256.7%
OpenAI 1,091.4 −10.4%
Anthropic 38,065.7 −86.7%

The gap between the extremes is roughly seven thousand fold. But the main lesson from the table isn’t in the specific numbers — it’s that “AI crawlers” aren’t one thing with one behavior.

Why you can’t rely on someone else’s numbers

Look at the right-hand column. In just six months, Anthropic fell from 286,930 to 38,065.7 while Perplexity jumped from 54 to 195. Anyone who built a strategy in January around the values at the time had it backwards on both platforms by July.

This part holds up better than the specific values. Cloudflare reports that model training accounts for roughly 80 percent of AI crawler activity, up from 72 percent a year earlier. Search takes about 18 percent, and user-triggered actions 2 percent.

So the vast majority of what crawlers take away doesn’t come back as a visit — and that isn’t a failure or bad faith, it’s the purpose of that crawling. It’s exactly why Cloudflare’s AI crawler settings split training, search, and agents into three separate controls.

How to find your own number

Since July 2026, Cloudflare offers a report showing the crawl-to-refer ratio for your site and broken down by operator, over the past 24 hours, 7 days, or 30 days.

  1. Find the report in Cloudflare

  2. Pick 30 days, not 24 hours

  3. Break the number down by operator

  4. Compare against the previous period

  5. No Cloudflare? Go to the logs

When the number is worth acting on

The size of the ratio doesn’t decide anything on its own. A framework for thinking it through:

  • When your content is easy to rewrite into an answer (how-tos, descriptions, comparisons), a high ratio on a given platform means real value walking out the door. This is where deciding deliberately pays off.
  • When the content is supporting material and the sale happens elsewhere (referrals, resellers, a physical location), the same number may not sting at all. Load isn’t the same as loss.
  • When the number shifts suddenly, don’t block right away. Verify the 30-day trend and check whether the change traces back to something on your side.
  • When you don’t know what the AI channel brings you at all, that’s a signal to fix measurement first, not to block. The article on the limits of measurement works through it.

What this article doesn’t cover

It doesn’t recommend blocking or allowing. The framework for that decision is in the article on opting out of AI answers, and the CDN-level implementation is in the article on Cloudflare.

It doesn’t cover hosting costs and isn’t a guide for server admins. Caching and rate limiting are a separate topic.

What to take away

More useful than “how much are AI crawlers costing me” is the question of how much content each operator takes for every visit it sends back. We’ve found it helps to picture it as an exchange rate — Cloudflare doesn’t describe it that way, but it captures nicely that the figure varies by orders of magnitude between operators and changes fast.

Take two things from this. Most crawling goes to training, so a large share of the load will never come back by design — and that’s intended. And for the first time, instead of someone else’s numbers you can see your own, broken down by operator and over a sensible window.

If you see differences between platforms in those numbers but don’t know what to allow, what to restrict, and how to reflect it in measurement, that’s exactly when it’s worth bringing someone in.


Not sure how to read your own numbers? This site is run by Sniper Design — an AI SEO audit covers measurement, visibility, and technical setup, and names what you should actually be deciding from.

Sources: the Cloudflare blog (“The crawl-to-click gap,” July 2025 data) and the Cloudflare changelog of July 1, 2026. Current as of July 19, 2026 — these values change fast, and we’ll keep the article updated.

Sniper Design
Help with implementation

Don't want to handle it in-house? We'll build it for you.

At Sniper Design we do full‑service AI SEO — strategy, audit, implementation, and content. E‑commerce specialists since 2016, 600+ e‑shops delivered. We build AI search in from the ground up — into homepage designs, content structures, and client site audits.

  • E‑commerce since 2016
  • 600+ e‑shops
  • Our own e‑shop
FAQ · 6 questions

Common questions on this topic

01 What is the crawl-to-refer ratio?
Cloudflare defines it as the number of pages a platform crawls relative to how often it sends a visitor to your site. A high ratio means a lot of crawling and few visits back. It's a more practical metric than request volume on its own, because it shows directly what you're getting for the load.
02 What were the numbers for the major platforms?
In Cloudflare data for July 2025, Google came in at 5.4, Microsoft at 40.7, Perplexity at 194.8, OpenAI at 1,091.4, and Anthropic at 38,065.7. These are figures from the Cloudflare network for that period, not from any one site — your own numbers may look very different.
03 How do I find my own number?
Since July 2026, Cloudflare offers a report that shows the crawl-to-refer ratio for your whole site and broken down by operator, over the past 24 hours, 7 days, or 30 days. Choose 30 days; shorter windows are vulnerable to one-off spikes. If you aren't on Cloudflare, you'll need to go into your server logs.
04 Do these numbers change fast?
Very. Between January and July 2025, Anthropic's ratio dropped by 86.7 percent while Perplexity's climbed by 256.7 percent. Treat any specific number as a snapshot on a date, not a permanent trait of the platform.
05 Why does most crawling go to training?
Cloudflare reports that training accounts for roughly 80 percent of AI crawler activity, up from 72 percent a year earlier. Search takes about 18 percent and user-triggered actions 2 percent. Content pulled for training doesn't come back as a visit — that isn't a failure, it's the purpose of that crawling.
06 Does a high ratio mean I should block the crawler?
Not automatically. It's an input to the decision, not the decision itself. A high ratio tells you how much content leaves for every visit that comes back — but whether that trade bothers you depends on how you make money and what you expect from the AI channel.
Keep reading

Related articles

All blog articles Back to home