Your biggest audience is now a robot: 8 changes to make before the traffic goes

Your biggest audience is now a robot: 8 changes to make before the traffic goes
Business & Technology

Your biggest audience is now a robot: 8 changes to make before the traffic goes

Oct 04, 2026
In July 2025, Cloudflare measured how many pages each AI company crawled for every visitor it sent back. Anthropic's ratio was 38,065 to one. OpenAI's was 1,091 to one. Google's was 5.4 to one.

Those figures come from Cloudflare's crawl-to-click analysis, and they move fast. Anthropic's own ratio fell from 286,930 in January to 38,065 by July, which tells you how unsettled this is. The direction is the point rather than the decimal.

What the numbers describe is the quiet end of an arrangement that held for twenty years. You published something useful, a search engine indexed it, and it sent you people. The reading and the visit arrived together. They have now come apart, and most websites are still built for the version where they did not.



What the network data shows

  • More than half of crawler requests are now for AI training.
    Cloudflare's one year review of the agentic internet puts it at 52% as of June 2026, up from 22% in spring 2025.

  • Pure search crawling is now a small and declining share of all crawler activity.
    Mixed-use crawlers, which blend search, agent use and training in the same visit, account for over 36%. The tidy distinction between a search bot and a training bot is disappearing.

  • Some heavily crawled categories lost up to 40% of their human traffic in under a year.
    Same source. Not every category, and not evenly, but the pattern is visible at network scale rather than anecdotal.

  • People spend about fifteen minutes on the open web for every hour they spend looking for information.
    The rest happens inside an interface that reads your page on their behalf.

  • Training grew from 72% of AI crawling in July 2024 to 79% a year later, while search fell from 26% to 17%.
    The machines reading your site are increasingly not the ones that send anybody to it.

Worth saying plainly: this is Cloudflare network data, so it reflects the sites behind Cloudflare rather than the entire web, and the per-company ratios are volatile enough that quoting one as a fixed fact would be wrong by next quarter. Use them to understand the shape of the change, not to build a forecast.



1. Find out what your own split actually is

Before changing anything, look at your server logs. Not your analytics, which mostly cannot see bots, but the raw request log.

Count requests by user agent. Separate the obvious search crawlers from the AI crawlers, and compare the crawl volume against the referral traffic each platform sends you. Your ratios will not match the network average, and the gap between them is the only number that should drive your decisions. A documentation site and a recipe site are having completely different experiences right now.

2. Decide what you are giving away, and to whom

There is a real choice here, and it is not binary. Cloudflare's managed robots.txt, launched in July 2025 and available on free plans, blocks training-specific crawlers while leaving search visibility intact. There is also an option to block only on pages carrying advertising.

Many publishers have settled on a middle position: block the crawlers that train models, allow the ones that answer questions and cite sources. That is a defensible stance, and it is worth knowing that the distinction is eroding as mixed-use crawlers grow. Make the decision deliberately rather than inheriting a default you never chose.

3. Serve your content without requiring JavaScript

This is the single most common technical reason a site gets read badly by machines. Many crawlers and assistant fetchers execute little or no JavaScript, so a page that assembles its content client-side can be effectively blank to them.

Server rendering or static generation fixes it. The test is simple and takes a minute: disable JavaScript in your browser and load your most important page. Whatever remains is roughly what a machine sees. If that is a loading spinner, nothing else in this list matters yet.

4. Put the facts where they can be lifted

A model summarising your page needs to extract specifics: what you do, where you operate, what it costs, who you serve, what the answer to the question is.

That means structured data that matches what the page actually says, clear headings that state the thing rather than tease it, prices and specifications in text rather than baked into images, and a page title that works as an answer out of context. Schema markup is not a ranking trick here. It is a way of writing down the facts so they survive being read by something that does not look at your layout.

5. Answer first, elaborate second

The journalistic inverted pyramid has become a technical requirement. If the answer to the page's question is in the final paragraph, a model summarising the first few hundred words will miss it, and a human skimming will too.

Lead with the direct answer in plain language, then supply the reasoning, the caveats and the detail. This is also simply better writing, which is why it is one of the few changes in this list with no downside if the whole landscape shifts again.

6. Give the citation somewhere worth going

When an assistant does cite you, the person arriving has already received the answer. They are not clicking for information. They are clicking to check you are real, to see what you charge, or to find out whether you serve their city.

Make that visitor's landing experience match that intention. Credentials, examples of work, pricing or at least a pricing range, and a clear way to start a conversation. A cited page that answers the question and then offers nothing beyond it earns you a view and nothing else. This is where most of the commercial value of AI visibility is actually won or lost.

7. Build the channels that nobody else controls

Email lists, direct bookmarks, communities, a product people sign into. These are unglamorous, and they are the only distribution that cannot be restructured by a platform decision you were not consulted about.

Every content strategy that depended entirely on being discovered by an intermediary is now carrying a risk it did not price. The organisations weathering this best are the ones that spent the last few years converting borrowed attention into a direct relationship.

8. Raise conversion, because the visits are getting scarcer

If fewer people arrive, the arithmetic only works if more of them do something. This is the least fashionable item on the list and probably the highest return.

Fix the forms, speed up the pages, clarify the offer, remove the steps nobody needs. A site converting at 1% that loses a third of its traffic needs to reach 1.5% to stand still, and reaching 1.5% is usually more achievable than replacing the traffic. The work is unexciting and the maths is unforgiving.



Nobody knows how this settles

It would be easy to write a confident conclusion here, and every confident conclusion currently on offer is a guess. Licensing frameworks, pay-per-crawl arrangements and new standards for declaring terms are all being built in public, and which of them becomes normal is genuinely unresolved. Anyone claiming certainty about the web in 2028 is selling something.

Blocking everything is also a bet, not a safe harbour. It protects your content from being used without compensation, and it removes you from the interfaces where a growing share of people now start looking. Both outcomes are real, and the right answer depends on whether your business runs on attention or on transactions. A local builder and a news publisher should not make the same choice, and most of the advice circulating treats them as if they should.

What is not a guess is the asymmetry. The reading and the visiting have separated, and they are unlikely to recombine. Whatever the eventual settlement looks like, a site that is readable by machines, that answers questions directly, that converts the visitors it does get, and that owns a direct line to its customers is better positioned under every scenario. That is the useful part of this: the hedges and the sensible moves are the same moves.

Which means the decision is less dramatic than the headlines suggest. Do the durable things now and stay flexible about the rest.




Three tests to run this week

  • Turn off JavaScript and load your three most important pages.
    Whatever is still visible is approximately what a machine reads. Thirty seconds per page, and the result often explains a lot.

  • Grep your server logs for AI crawler user agents.
    Compare the request counts against the referrals those platforms sent you last month. That ratio is your actual situation rather than the industry average.

  • Ask an assistant a question your business should be the answer to.
    See whether you appear, who does instead, and what it says about you. Then read the page it would have cited and ask whether that page would convert a stranger.

None of that requires a budget or a strategy document. It produces the three facts you would need before writing one.

Want your site readable by machines and persuasive to the people they send? Talk to Spark about building for both audiences properly.
Found this article useful? Share it with your network.