The crawl-to-refer ratio tells you how many pages an AI platform crawls on your site for every visitor it sends back. In Cloudflare network data for July 2025 it ranged from 5.4 to 38,065.7 depending on the operator. Roughly 80 percent of that crawling goes to model training.
AI crawler load usually gets discussed in terms of volume: how many requests, how much data transferred. This article is not about server performance or hosting costs. It’s about a metric that’s more useful for making decisions: what you get back for that load.
What the crawl-to-refer ratio measures
Cloudflare defines it as the number of pages a platform crawls compared with how often it sends a person to your site. The higher the number, the more content leaves for every visit that comes back.
| Aspekt | Ratio | Change since January 2025 |
|---|---|---|
| 5.4 | +43% | |
| Microsoft | 40.7 | +5.7% |
| Perplexity | 194.8 | +256.7% |
| OpenAI | 1,091.4 | −10.4% |
| Anthropic | 38,065.7 | −86.7% |
The gap between the extremes is roughly seven thousand fold. But the main lesson from the table isn’t in the specific numbers — it’s that “AI crawlers” aren’t one thing with one behavior.
Why you can’t rely on someone else’s numbers
Look at the right-hand column. In just six months, Anthropic fell from 286,930 to 38,065.7 while Perplexity jumped from 54 to 195. Anyone who built a strategy in January around the values at the time had it backwards on both platforms by July.
Most crawling goes to training, not search
This part holds up better than the specific values. Cloudflare reports that model training accounts for roughly 80 percent of AI crawler activity, up from 72 percent a year earlier. Search takes about 18 percent, and user-triggered actions 2 percent.
So the vast majority of what crawlers take away doesn’t come back as a visit — and that isn’t a failure or bad faith, it’s the purpose of that crawling. It’s exactly why Cloudflare’s AI crawler settings split training, search, and agents into three separate controls.
How to find your own number
Since July 2026, Cloudflare offers a report showing the crawl-to-refer ratio for your site and broken down by operator, over the past 24 hours, 7 days, or 30 days.
-
Find the report in Cloudflare
-
Pick 30 days, not 24 hours
-
Break the number down by operator
-
Compare against the previous period
-
No Cloudflare? Go to the logs
When the number is worth acting on
The size of the ratio doesn’t decide anything on its own. A framework for thinking it through:
- When your content is easy to rewrite into an answer (how-tos, descriptions, comparisons), a high ratio on a given platform means real value walking out the door. This is where deciding deliberately pays off.
- When the content is supporting material and the sale happens elsewhere (referrals, resellers, a physical location), the same number may not sting at all. Load isn’t the same as loss.
- When the number shifts suddenly, don’t block right away. Verify the 30-day trend and check whether the change traces back to something on your side.
- When you don’t know what the AI channel brings you at all, that’s a signal to fix measurement first, not to block. The article on the limits of measurement works through it.
What this article doesn’t cover
It doesn’t recommend blocking or allowing. The framework for that decision is in the article on opting out of AI answers, and the CDN-level implementation is in the article on Cloudflare.
It doesn’t cover hosting costs and isn’t a guide for server admins. Caching and rate limiting are a separate topic.
What to take away
More useful than “how much are AI crawlers costing me” is the question of how much content each operator takes for every visit it sends back. We’ve found it helps to picture it as an exchange rate — Cloudflare doesn’t describe it that way, but it captures nicely that the figure varies by orders of magnitude between operators and changes fast.
Take two things from this. Most crawling goes to training, so a large share of the load will never come back by design — and that’s intended. And for the first time, instead of someone else’s numbers you can see your own, broken down by operator and over a sensible window.
If you see differences between platforms in those numbers but don’t know what to allow, what to restrict, and how to reflect it in measurement, that’s exactly when it’s worth bringing someone in.
Not sure how to read your own numbers? This site is run by Sniper Design — an AI SEO audit covers measurement, visibility, and technical setup, and names what you should actually be deciding from.
Sources: the Cloudflare blog (“The crawl-to-click gap,” July 2025 data) and the Cloudflare changelog of July 1, 2026. Current as of July 19, 2026 — these values change fast, and we’ll keep the article updated.