When bots bend the rules: why AI crawlers became a governance question
A public dispute between Cloudflare and Perplexity shows that AI crawlers are no longer a neutral infrastructure detail. For brands and operators that raises a new question: how do I keep control over what agents do with my website, and over what they see at all?
For a long time web crawlers were technical background noise. A bot announced itself by name, read the rules in robots.txt and either honoured them or did not. That was a question for administrators, not for management.
That has changed. In the summer of 2025 the infrastructure provider Cloudflare publicly accused the AI company Perplexity of using covert, undeclared crawlers to get around the no-crawl rules set by websites. Perplexity firmly disputed that account. Regardless of who turns out to be right, the incident marks a turning point. It shows that the question of which bot may read which content stopped being a purely technical matter a while ago. It touches the relationship between publishers, infrastructure providers and AI companies. Bots have become political.
What Cloudflare actually criticised
To understand the dispute, it helps to look at the ground rules that have held the open web together so far.
The robots.txt file is a simple document in which an operator defines which parts of a site may be fetched automatically and which may not. It is not a technical barrier but a convention. Reputable providers honour it. Alongside it sits the concept of verified bots: a crawler identifies itself through its user agent, and its behaviour can be checked against known address ranges. That way an operator knows who is reading.
Cloudflare now describes behaviour that sidesteps those conventions. According to its account, a crawler that had been blocked under its official identifier changed identity: a generic user agent that looks like an ordinary browser, plus rotating IP addresses from different networks. Cloudflare says it tested this with purpose-built pages that nobody could have known about, so-called honeytraps, explicitly disallowed in robots.txt. The content still turned up in answers.
For site operators the consequence is uncomfortable. Anyone relying on a robots.txt block to work can no longer be certain that it does. Content deliberately held back can still be fetched and processed underneath the surface. The control you believed you had turns into an assumption.
What this means for brands and operators
At the business level the incident has two sides, and they pull in opposite directions.
One side is a governance problem. When bots fetch content against your own rules, a brand loses authority over what happens to its texts, prices and statements. That is not only a copyright question, it is also about control over how you are represented. If you do not know which content an agent actually read, you also do not know what its answers rest on.
The other side is a visibility problem, and it is often the bigger one. While much of the debate is about unwanted crawling, the opposite risk is more real for most brands: bots that do not see content although they should. A page that is technically hard to read gets passed over. The brand is then missing not because it was blocked, but because it stayed invisible.
Both sides lead to the same four questions every brand should be able to answer today. Which bots read my site? What do they see of my brand? Which answers do they build from my content? And how do I notice when that changes? Anyone who cannot answer these leaves their own presence in AI answers to chance.
Two different things: readability and behaviour
This is where a distinction is worth making, one that often blurs in the argument about stealth crawlers. There are two separate layers, and a brand needs clarity on both.
The first layer is technical readability, often called AI readiness. It answers the question of whether an agent can capture the content cleanly at all. That includes server-side rendering, so content does not appear only after JavaScript runs, structured data that makes statements machine-readable, clean feeds and clear access rules, including newer conventions such as an llms.txt that explains to AI systems how to treat a site. This part is the operator’s responsibility.
The second layer is the behaviour of the agents. It answers the question of which bots arrive when, how often they read and whether they honour the rules. This layer sits outside a brand’s direct control, but it can be observed. The Cloudflare incident is an example of why that observation matters: optimising your own readability without ever checking which bots actually turn up shows you only half the picture.
The craft lies in thinking about both layers together. A brand needs to know whether its website is technically readable for AI, and at the same time understand which agents are helping themselves to it and which ones are in disguise.
Klariton as infrastructure for control and visibility
This is exactly where an infrastructure layer of the kind Klariton provides comes in. Not as a weapon against individual crawlers, because nobody wins the arms race against bots in disguise for long, but as a foundation that makes both layers visible and steerable. Five building blocks show what that looks like.
Technical AI Audit. The first step checks whether AI crawlers see usable answers server-side at all, rather than an empty JavaScript skeleton that only fills up in the browser. Many sites deliver their actual content only after scripts have run, and most bots do not wait for that. The audit exposes the gap before it costs visibility.
Agentic Reach. Instead of relying on the assumption that a rule works, Agentic Reach measures which agents actually read the content, broken down by provider, category and page. That produces a real picture of who comes by, and deviations become visible, including ones you would otherwise miss.
LLM Discovery. This layer shows where a brand appears in the answers of assistants and in what role: as a recommendation, as a cited source, or not at all. A vague sense of somehow being present turns into a solid basis for decisions.
Trust Intelligence and Safe Guard. Both watch which published answers are weak or outdated, and stop exactly those from becoming the base material for agents. Because what an agent picks up once, it repeats.
Compliance Center. It governs the rules under which answers may go out into the world at all: which sources are permitted, which claims are allowed, which policies apply. Control over how you are represented stays in one place instead of getting lost across individual texts.
The thinking behind it is sober. You cannot force every bot to behave. But you can make sure your own brand is technically readable, that you see who is reading, and that the answers built from it are correct.
Practical guidance
For people responsible for marketing, e-commerce, digital and technology, the incident has concrete consequences.
Treat bots as a category of their own. Keep a separate section for AI bots in your monitoring, apart from classic traffic. Only what is visible can be steered. Agentic Reach delivers that view by provider and page.
Treat robots.txt and policies as AI governance. Access rules are no longer just an SEO detail, they are part of the question of who may use which content in your brand’s name. A compliance centre turns those rules into a deliberate decision.
Check technical AI readability. Server-side rendering, structured data, clean feeds and an llms.txt decide whether an agent can capture your content. An audit shows the gaps before they cost revenue.
Make visible who uses you and who only uses your content. Distinguish between bot reads and real AI referral clicks. One shows who reads your content, the other who actually sends you customers as a result. Both belong on the same dashboard.
Build answers that are AI-ready. Verified, quotable units beat loose scraps of text. Agents are more likely to recommend what is unambiguous and backed by evidence. Buying answers and a guided advisor deliver exactly those units.
Schedule regular safe checks. Drift, outdated content and new guidelines quietly change what agents make of your material. Safe Guard keeps checking against the current source and reports before a customer sees the wrong answer.
The common thread is the same throughout. Trying to lock out individual crawlers is a race without a finish line. The more durable answer is a layer that keeps readability, behaviour and answers visible and steerable over time.
From defence to control
The dispute between Cloudflare and Perplexity will not be the last of its kind. The more important AI assistants become for purchase decisions and research, the harder the fight over access to content will get. For brands, though, the decisive insight is not which side is right in this particular case.
The insight is that control over your own presence in AI answers no longer comes from blocking alone, but from knowing. Anyone who knows whether their site is readable, who reads it and what comes of it can act. Anyone who does not know depends on everybody honouring rules that this incident shows are not a given. Control shifts from defence to observation, and from assumption to measurement.
Ask your question about Klariton.
Grounded in Klariton’s own knowledge, cited rather than invented.
How visible is your brand to AI?
The free AI visibility check shows you in under a minute how AI assistants see your shop today.