How-to
The Bots & Agents Page
Verified crawlers, fake Googlebots, AI agents, and everything else that declares itself a bot, split apart so you can set policy per group.
Jump to section
What the page is for
"Bot traffic" is not one thing. Googlebot indexing your catalogue, a scraper pretending to be Googlebot, an AI agent citing your product page in an answer, and a declared monitoring bot all need different treatment. The Bots & agents page (under Monitor) splits them apart so you can decide policy per group instead of guessing.
The page reads top to bottom in priority order: impersonation first, but only while there is any (a clean window gets a one-line all-clear instead of an empty threat panel), then what AI is citing, the AI agents themselves, search crawlers, the unclassified group, and finally the banked history. The four tiles at the top jump straight to their sections.
The four groups
- Verified crawlers — traffic that claims to be a known crawler (Googlebot, Bingbot, etc.) and passes verification. Usually welcome; blocking it hurts your SEO.
- Impersonation — traffic that claims to be a known crawler but fails verification. There is no innocent explanation for a fake Googlebot: real crawlers always verify. This group feeds the Verified Bot Impersonation detector, which opens an alert when fake-crawler traffic crosses its floor.
- AI agent requests — traffic from declared AI agents. Every agent carries an intent label, because an AI fetch means one of three very different things: training crawl (GPTBot, ClaudeBot, Bytespider and friends reading your catalogue in bulk), AI search (fetches that power AI search indexes, such as OAI-SearchBot and ExaSearchBot), and live retrieval (ChatGPT-User, Claude-User, Perplexity-User fetching a page while composing an answer for a real person). The AI Agent Traffic detector alerts specifically when agents attempt transactions (POST requests: add to cart, checkout, login), because reading a product page and attempting a purchase deserve different policies.
- Other bots — declared bots that aren't verified crawlers or known AI agents: monitors, SEO tools, unclassified automation.
Each group drills into the networks behind it, so "impersonation traffic is up" turns into "and it comes from these two hosting providers" in one click. The impersonation table also names the network next to its AS number, because the identity is the story: a fake Googlebot from a rented cloud server is a different situation from one on a residential line.
What AI is citing
The live-retrieval agents check a page at the moment they cite it, so each of their fetches is your content appearing in a generated answer: the closest server-side signal there is to an AI impression. The What AI is citing panel lists the pages those agents fetched in your selected window, with the split per agent, so "is AI recommending us?" becomes a list of actual product and landing pages with numbers on them.
For these agents, and only these, Edge records the exact page — /fr/chemise-sans-col---blanc/CSR0806WHT.html, not a :slug placeholder — because for a citation the specific product is the whole point. Query strings are never kept, so the privacy posture is unchanged. Bulk crawls and AI-search index fetches deliberately stay grouped by URL pattern instead: for an index, the pattern is the right level of detail, and their volume makes per-URL capture counterproductive.
Treat it as a good proxy rather than a measurement: agents cache and sometimes cite without re-fetching. Claude and Perplexity retrieval agents are identified separately, and exact pages are recorded, from 19 August 2026 onward, so their history builds from that date. Edge also banks a weekly per-page summary of these fetches indefinitely, so the trend outlives the raw log window.
The impersonation summary
When there is impersonation in the window, the panel opens with a short plain-English summary: what the traffic claims to be, what network it actually comes from, whether any of it touched checkout, auth, or cart paths, and where that leaves you. It is written by AI from the same figures the panel shows; the numbers and any drafted rule come from the deterministic pipeline, not the AI, and nothing is applied without your approval. When there is no impersonation, there is no summary, because the panel's own "none detected" line already says everything worth saying.
Sensitive paths touched
Every agent, crawler, and impersonator row carries a sensitive-paths flag: a red marker with checkout / auth / cart chips when that bot's traffic hit an SFCC checkout, account, or cart controller. It is computed independently of volume, so a bot that is millions of image requests but pokes a checkout controller thirty times still lights up. This is the signal a raw total hides: "139k fake Googlebot" reads as image scraping until the flag shows a handful of those requests landed on your payment pages.
Treat the flag as a prompt to look, not a verdict. Real search crawlers do render some pages, so a verified crawler touching a cart endpoint can be legitimate. The point is that you see it and decide.
Drilling into what a bot scanned
Click any row to expand it. Edge shows:
- Where it came from — the networks behind the bot's traffic, each with its share and its verified/unverified split. For a crawler like Googlebot this is the fastest honesty check on the page: the real crawler runs from its own network, fully verified, while every impersonator shows up as another network with an unverified share. Click a network to focus the whole drawer on it; when a network's traffic is mostly unverified, the drawer offers to draft a challenge rule for it.
- When it hit — an activity timeline across the selected window, so a steady crawl and a sudden burst stop looking the same. Intervals that are mostly errors are tinted red: a burst that is all errors is probing, not crawling. The timeline is also the drawer's time control: click a bar, or drag across several, and every other section recomputes for just that slice, so a spike becomes "these pages, from these networks, in that hour". Time and network focus combine (select the burst, then click the network behind it), and Show whole window clears the slice.
- Top paths the bot hit, each with its request count and, where relevant, its POST volume and error (4xx/5xx) count. POST volume on a crawler or agent means it tried to transact, not just read; errors mean it is probing endpoints that reject it.
- Sensitive paths touched — the exact checkout / auth / cart controllers behind the row flag, listed even when their volume is far too low to appear in the top paths.
- What it fetched — the bot's traffic split by content type: pages (html), data (json), images, and static assets. Pages and data mean the bot is reading your catalogue; heavy image and asset weight is the signature of media harvesting. This split appears when your Logpush job ships the
EdgeResponseContentTypefield (see the required-fields article for how to add it). - Where your logs ship Cloudflare's
BotDetectionIDsfield, the drawer also lists the detection signature IDs Cloudflare stamped on this traffic, for correlating with Cloudflare's own tooling.
For an impersonator row, the drawer scopes to the fake (unverified) half only, from that row's network only, so you see what each spoofing source specifically went after, separated from any genuine crawler on the same name and from other networks running the same disguise. Each impersonation row also carries a Draft rule shortcut that opens the network's detail page with a challenge rule already drafted — generated disabled, applied only if you decide.
The drawer shows the busiest paths first. Where per-path detail ends, the drawer says how many further requests it holds no detail for, rather than hiding them inside a bucket. On windows that reach back before 18 August 2026, bot traffic ran through an older rollup that kept less per-bot path detail, so the long tail on those windows shows as a single collapsed row.
When you need every URL a suspicious bot touched, exactly as sent, use the Capture this bot's exact URLs link in the drawer. It opens a short forensic capture window scoped to that one bot: exact paths, query strings if you ask for them, and the IPs behind the traffic. The capture runs on live traffic from the moment you start it.
Bots over time
Everything above shows the window you have selected; Bots over time shows the history. Every completed week, Edge banks a per-bot summary (requests, verified share, POSTs, errors, traffic that reached your origin) and keeps it indefinitely — long after the raw log window has rolled off. The panel charts the busiest agents week by week and tables the latest week with its change against the week before, so "GPTBot has doubled since July" is a fact you can read off, not a feeling.
The banked rows are deliberately free of personal data: a bot name and counters, no IPs, no addresses, no URLs, no query strings. That is what makes keeping them forever straightforward.
Download CSV exports the same table — one row per bot per week — for your own reporting tools.
The weekly report
Settings offers an optional weekly bot report, delivered each Monday with the previous completed week. It reads the way this page does, in the same priority order: crawler impersonation first — with a plain all-clear line when the week was clean — then what AI cited from your store, the AI agents with their intent labels, the search crawlers, and the other declared bots. The numbers come from the same per-bot data the page shows, so the report never disagrees with the screen, and like everything banked weekly it contains no personal data.
The report has its own recipients, separate from alert routing. Under the Weekly bot report switch in Settings you can set one or more report email addresses (comma separated; leave blank to keep using your default notification address) and, optionally, a report webhook. Paste a Teams workflow or Slack webhook URL there and the report arrives in the channel as a card; any other HTTPS endpoint receives it as structured JSON — one object with the impersonation, citations, agent, crawler, and other-bot sections — for your own tooling. Week-over-week comparisons appear only when both weeks are directly comparable; across a dataset upgrade the report pauses them for one week and says so, rather than showing changes the traffic never made.
You don't have to wait for Monday: Send report now in the same Settings block sends the report on demand, to the same recipients, whether or not the weekly schedule is switched on. Pick a date and it sends the report covering the last completed week as of that date, so you can pull up any banked week from the history.
The report also carries a Marketing channels section once channel history is banked: landing views per acquisition channel with Cloudflare's verdict, money channels first, opening with a short generated summary composed only from the week's own numbers. See the Channels page article for the full picture behind those numbers.
What to do with each group
- Verified crawlers: leave alone unless crawl volume is causing origin load.
- Impersonation: treat as hostile. The alert carries a drafted rule; challenge-first is safe because a real crawler never solves a challenge from a fake source anyway.
- AI agents: a policy decision, not automatically hostile. Some merchants welcome agent traffic (it converts), others challenge transactions from agents. Edge tells you the volume and lets you decide.
- Other bots: audit occasionally; ignore known-good monitors via the usual pin/ignore controls on Offenders. The Other declared bots panel at the bottom of the page lists what is inside this group by the name each bot gives itself, up to 40 names so low-volume tools are not buried under the heavy crawlers. Each name carries an identity label saying who operates it and what it does — SEO crawlers (Ahrefs, Semrush, Majestic), link-preview bots (Facebook, Pinterest, TikTok), ad checkers, uptime monitors, archives — with a fuller sentence on hover. Where nobody credibly claims a name, the label says so explicitly: an unidentified crawler with real volume deserves more attention than a known SEO tool, and the panel makes that difference visible at a glance. Performance testers such as GTmetrix and Lighthouse, which fetch pages with a real browser identity and used to be invisible here, are classified and listed by name too. Every name expands into the same drill-down as the named bots — networks, activity, paths, content types — so an unfamiliar name is one click from being understood. Being unverified is normal in this group (most tools sit outside Cloudflare's verified-bot programme), so the drawer says so plainly rather than treating it as suspicious. A recurring heavy name is worth flagging to support so it can be classified properly.
Edge also watches this population for you: the New bot detected and Bot volume spike alerts fire when a never-seen name arrives with real volume, or a known bot suddenly runs at a multiple of its own usual rate. Both link straight back to the bot's row on this page.
Requirements
Crawler verification depends on the BotScoreSrc field in your Logpush job. If it's missing, the page shows a banner and impersonation detection is inactive until you add the field. See HTTP Methods and Required Logpush Fields.