Insights

How AI Collects Data on Your Site, and What Traffic You Really See

Five kinds of AI traffic, the crawl-to-click ratio, and why last-click attribution makes AI’s contribution invisible. From the first line of markup to the attributed purchase.

10 min read·
Crawler lesen eine Seite; der meiste AI-Traffic bleibt unsichtbar, nur ein Klick wird sichtbar

While you read this, AI bots are visiting your site. Not once, but continuously, and most of them never send a click back. They read, learn, cite and enrich. The catch: in Google Analytics you see almost none of it, because AI crawlers do not run JavaScript. What reaches the server is the full truth, and it is surprising.

Five ways the AI visits your site

AI traffic is not one thing. It is at least five different intents, and each one means something different for you:

Class

Example bots

What happens

What it means

Standard crawl (training)

GPTBot, ClaudeBot, CCBot, Google-Extended

Ingests content for the model, broad and periodic

You enter the model’s knowledge, but without a live link

On-demand research

ChatGPT-User, OAI-SearchBot, Perplexity-User

Fetches live because someone is asking right now

Your page is checked for a specific answer

Deep Research

Burst of the same on-demand bots

Many pages at once for a report

Someone is researching you deeply and in stages

Citation

The subset with a link in the answer

Shows up visibly in the answer, rare click

The actual referral, the rest is silent consumption

Enrichment / agents

Amazonbot (Rufus), Applebot-Extended (Apple Intelligence, Siri)

Enriches product and company data for assistants and voice

Your data flows into buying agents and voice answers

A training crawl is something completely different from someone actively investigating you in a Deep Research. Lumping both together means not understanding your AI traffic.

Not all of it is AI, separate the noise

Alongside the AI bots comes a lot of traffic that has nothing to do with AI, and that skews your numbers if you count it:

  • SEO tools: AhrefsBot, SemrushBot and the like crawl for backlink and keyword indexes, not for AI answers.

  • Competitor monitoring: tools others use to watch you.

  • Scrapers and spam: often from datacenter ranges, grabbing content without ever bringing you anything.

And to be honest: our own Klariton scans that happen while working with you (checks, verification) are not counted in your reach numbers. Otherwise you would be skewing your own measurement. Clean AI reach means: AI bots yes, SEO tools, spam and our own tooling out.

How to spot the patterns in your logs

These signals live in your server logs, not in Google Analytics (which loads via JavaScript the bot never starts). The key is the user agent of every request. In a log, an AI visit looks roughly like this:

203.0.113.7 [21/Jul/2026] "GET /product/xy" 200
"Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)"

The token in the user agent, here ChatGPT-User, reveals sender and intent. The key tokens, as of July 2026 (they change, check them regularly):

User agent contains

Who and class

What it means

GPTBot

OpenAI · training

For the model, broad and periodic

OAI-SearchBot, ChatGPT-User

OpenAI · on-demand

Fetches live for a specific answer

PerplexityBot

Perplexity · training

For the model

Perplexity-User

Perplexity · on-demand

Live for an answer

ClaudeBot

Anthropic · training

For the model

Claude-User

Anthropic · on-demand

Live for an answer

Google-Extended

Google · training

For Gemini training

Amazonbot

Amazon (Rufus, Alexa) · enrichment

Enriches product data

Applebot-Extended

Apple (Apple Intelligence, Siri) · enrichment

For assistants and voice

AhrefsBot, SemrushBot

SEO tool · NOT AI

Filter out, or it skews your AI numbers

Single lines become a pattern once you add time: several on-demand tokens (ChatGPT-User, Perplexity-User) within a few minutes across different pages means Deep Research, someone is investigating you deeply right now. Empty or generic user agents (python-requests, curl) from pure datacenter ranges are scraper noise and do not count as AI. That separation is exactly what the Agentic Reach module does for you, so you never have to read logs yourself.

Read a lot, clicked rarely

Cloudflare put numbers on it (the crawl-to-click gap, July 2025): the ratio of crawls to clicks sent back is dramatic. Anthropic hit about 38,000 crawls per single visitor, OpenAI around 1,100, Perplexity just under 200. Your content is consumed massively, but the AI keeps the user inside its own answer instead of sending them to you.

Your brand can be read in AI a hundred times over, long before it sees a click. Whoever measures only clicks mistakes that for silence.

That is not bad news, it is invisible news. AI visibility happens whether you measure it or not. The only question is whether you can see it.

Visibility is not the goal, the purchase is

Being read is the start of the chain, not the end. The full chain runs: crawl → citation → click → purchase. Every stage loses volume, and every stage is its own measurement. Only the last one pays your bills.

  • Crawl: how often and by which bot is your page read?

  • Citation: do you make it into the answer, or only into silent consumption?

  • Click: does the citation lead to a real visit?

  • Purchase: does the visit turn into a sale, in the shop?

Most tools measure the first stage and call it success. But a crawl is not revenue. Only when you follow the chain to the purchase do you know what AI visibility actually gets you.

Why last-click makes AI invisible

This is the most expensive measurement error. AI is almost always the first contact, the discovery in the upper funnel. The customer asks the AI, finds you, but comes back on the second or third visit via Google or direct, and buys then. Measure by last-click and AI gets zero credit, the purchase looks like it came from Google.

That is how practically everyone underestimates AI’s contribution today. The right model is fractional, multi-touch attribution: first visit via AI referral, second via Google, third direct, and AI gets its fair share of the chain instead of vanishing from the stats.

AI is the first touch. Last-click erases it. Fractional attribution brings it back.

What you can measure of all this

Because AI crawlers do not run JavaScript, the measurement has to happen server-side, at the edge, before any script runs. That is exactly where Klariton’s Agentic Reach module sits:

  • Which bot read which page, in real time, cleanly split by intent (training, on-demand, Deep Research).

  • Active research: a burst of on-demand reads means someone is investigating you right now, and the topic of the pages reveals the intent.

  • All the way to revenue: first-touch capture and session grouping over time, so an AI first contact is not lost when the purchase later arrives via another channel.

The free Klariton check shows you the entry: do ChatGPT, Perplexity, Gemini and Claude even see and cite you? The rest of the chain, from research to attributed purchase, is exactly what Agentic Reach makes visible, while others stop at counting crawls.

Ask Klariton

Ask your question about Klariton.

Grounded in Klariton’s own knowledge, cited rather than invented.

Or ask your own question:
Next step

How visible is your brand to AI?

The free AI visibility check shows you in under a minute how AI assistants see your shop today.