Original research · Karbon Agency · reviewed 2026-09-16

AI Crawlers and AI Referrals: What Our Own Logs Show, 2026

Two first-party datasets, both ours. 82,294 tracked sessions across 9 local-business websites (May 24 – Sep 15, 2026), and 1,452 server-side crawler fetches from our own domain (Sep 14 – 17, 2026). The pixel tells us who arrives; the server log tells us who reads. They disagree by two orders of magnitude, and that gap is the story.

Are AI assistants sending real traffic to local-business websites yet?

Almost none — but they are reading constantly. AI-assistant crawlers fetched 373 pages vs 252 for Google, Bing and Apple combined (Sep 14 – 17, 2026, n=1,452). Yet only 58 of 17,410 human sessions (0.33%) came from an assistant (Jun 1 – Sep 15, 2026).

Every figure on this page is printed with the sample size and date window it was computed over. Where a cohort is too small to support a rate, we say so instead of publishing one.

1.5x
AI-assistant crawlers fetched 373 pages vs 252 for Google, Bing and Apple combined
n=625 · Sep 14 – 17, 2026
0.33%
of human sessions arrived from an AI assistant — 58 of 17,410
n=17,410 · Jun 1 – Sep 15, 2026
68%
of ChatGPT-referred sessions scrolled past halfway, vs 13.8% of search sessions
n=50 · Jun 1 – Sep 15, 2026
78.5%
of all tracked sessions were bots, but the monthly figure swung 59.7%–91.4%
n=82,294 · May 24 – Sep 15, 2026

1. AI crawlers now out-fetch search crawlers on our own site

Our JavaScript pixel cannot see these visitors at all — crawlers do not execute JS. So we read the server log directly. Over Sep 14 – 17, 2026 it recorded 1,452 crawler fetches against a single domain (karbonagency.com, our own site — see the methodology note on why this dataset is one site and not nine).

Crawler fetches by user agent. n=1,452 fetches, one domain, Sep 14 – 17, 2026.
CrawlerOperatorCategoryFetchesUnique pathsFetches / path
AmazonbotAmazonAI training / platform5695391.06
meta-externalagentMetaAI training / platform258783.31
BingbotMicrosoftTraditional search173812.14
GPTBotOpenAIAI assistant98442.23
PerplexityBotPerplexityAI assistant85841.01
OAI-SearchBotOpenAIAI assistant85561.52
GooglebotGoogleTraditional search70691.01
ChatGPT-UserOpenAIAI assistant6287.75
ClaudeBotAnthropicAI assistant40401.00
ApplebotAppleTraditional search991.00
Claude-UserAnthropicAI assistant331.00
Total1452

Grouped by who operates them: crawlers belonging to companies whose main product is an AI assistant (OpenAI, Perplexity, Anthropic) made 373 fetches. The three traditional search indexers — Googlebot, Bingbot, Applebot — made 252. That is roughly 1.5 AI fetches for every search fetch.

The largest single crawler was neither: Amazonbot, at 569 fetches, more than twice the entire traditional-search group. Amazonbot and Meta’s meta-externalagent (827 fetches between them) feed AI systems but also serve other product lines, so we count them in their own bucket rather than inflating the AI-assistant number with them. Either way the direction is the same: on this domain, in this window, Googlebot was the seventh-busiest crawler.

Two different machines, two different behaviours

The fetches-per-path column separates crawlers into two species. Sweepers — PerplexityBot (1.01), ClaudeBot (1.00), Googlebot (1.01) — fetch each URL once and move on; they are building an index. Answer-time fetchers behave nothing like that: ChatGPT-User made 62 fetches across just 8 paths (7.75 per path), which is what it looks like when a live assistant pulls the same handful of pages over and over to answer real user questions. GPTBot (2.23) and OAI-SearchBot (1.52) sit in between.

The practical read: a page can be indexed by every sweeper on the list and still never be fetched at answer time. The eight paths ChatGPT-User kept returning to are the ones actually doing work.

2. Which assistants actually send visitors? Essentially one

Across 17,410 human sessions on 9 sites (Jun 1 – Sep 15, 2026), we classified every referrer host. AI assistants accounted for 58 sessions in total — 0.33%.

Referring AI assistant, human sessions only. n=17,410 sessions, 9 sites, Jun 1 – Sep 15, 2026.
AssistantSessionsShare of human sessions
ChatGPT (chatgpt.com)500.29%
Perplexity3insufficient sample
Gemini3insufficient sample
Claude2insufficient sample
Microsoft Copilot0insufficient sample

ChatGPT is the entire channel. Perplexity (3 sessions), Gemini (3) and Claude (2) are each far below our 30-session publication floor, so we report their raw counts and refuse to compute a rate from them. Microsoft Copilot sent zero identifiable sessions. Anyone publishing a tidy percentage split of “AI search market share” from referral data at this volume is reporting noise.

One caveat that cuts the other way: assistants increasingly answer in-place without a click, and some strip the referrer entirely. A visitor who reads about a business in ChatGPT and then types the domain directly lands in our Direct bucket, indistinguishable from any other direct visit. The 58 sessions are a floor, not a measurement of AI’s influence.

3. The few AI-referred visitors read far more of the page

Small as it is, the ChatGPT cohort behaves unlike any other channel we track. These 50 sessions span 4 different client sites, with no single site contributing more than about 45% of them, so this is not one unusual website distorting the average.

Engagement by acquisition channel. Human sessions only, Jun 1 – Sep 15, 2026. Per-row sample sizes in the Sessions column; median engaged time is computed only over sessions that returned an engagement beacon (that sub-sample is shown separately).
ChannelSessions (n)Avg pagesMulti-pageScrolled past 50%Median engaged (n)
AI assistant (ChatGPT)501.9642.0%68.0%67.0s (43)
Search engine3,1051.3813.9%13.8%53.5s (772)
Social1,8891.4223.6%12.0%18.0s (1095)
Direct / untagged10,9951.4218.4%7.8%16.1s (1793)

68.0% of ChatGPT-referred sessions scrolled past the halfway point, against 13.8% of search sessions — a five-fold gap. They viewed 1.96 pages on average vs 1.38, and 42.0% went multi-page vs 13.9%. Median engaged time was 67.0 seconds (n=43) vs 53.5 seconds for search (n=772).

The plausible mechanism is selection, not magic: an assistant has already described the business and answered the easy questions before the person clicks. Whoever still clicks through has pre-qualified themselves. That matches the crawler data — the assistant reads the page so the visitor doesn’t have to, and only the seriously interested ones come.

How much weight this carries: n=50 is small. A shift of five sessions moves the multi-page rate by 10 points. Treat the direction as real and the exact figures as provisional; we will restate them at a larger n next quarter.

4. Bot share is not a stable number, and single-month benchmarks mislead

Across the full 82,294 sessions (May 24 – Sep 15, 2026), 78.5% were bots. But the monthly figure is volatile enough that the annual average describes no individual month well.

Bot share by month. n=82,294 sessions across 9 sites, May 24 – Sep 15, 2026. May and September are partial months.
MonthSessionsBot sessionsBot share
May 2026 (partial)2,1321,84086.3%
Jun 202622,05020,15791.4%
Jul 202617,35512,99474.9%
Aug 20266,5103,88759.7%
Sep 2026 (partial)34,24725,71475.1%

Bot share ranged from 59.7% in August to 91.4% in June — a 31.7-point swing across consecutive months on the same set of sites. Nothing about those sites changed that much. Automated traffic arrives in campaigns, and a benchmark quoted from one month’s data (including ours) can be off by thirty points. If you are comparing your bot rate to a published figure, compare distributions over months, not point estimates.

5. The pages the assistants are reading are not fast

From 14,429 performance beacons carrying an LCP value (human sessions, Jun 1 – Sep 15, 2026), the 75th-percentile Largest Contentful Paint was 3,184 ms — above the 2,500 ms “good” threshold, so the tracked set fails that bar at p75. 70.5% of individual page loads were inside it; the slow quarter drags the percentile out. Median-grade TTFB was 946 ms at p75.

We include this mainly to head off a wrong inference. Page speed is weak as a citation lever, and nothing in our crawler log suggests these crawlers are deterred by a 3.2-second LCP — they fetched the pages anyway. Fix LCP because slow pages lose humans, not because it will win you an AI citation.

Our pixel recorded no INP or CLS values in this window, so two of the three Core Web Vitals are absent from this report entirely.

What the data did not support

We tested more angles than we published. These are the ones we dropped, and why — listed because a study that only reports its successes is not telling you how it was made.

Angle testedWhy it is not in this report
Which form field people abandon mostOnly ~106 abandoned-field events were recorded in the window, and the largest single field cohort was 29. Every field falls under our 30-session publication floor, so there is no reportable ranking.
Device and connection split of AI-referred visitorsThe ChatGPT cohort is 50 sessions. Splitting it by device leaves cells in the single digits — below the floor, and meaningless as a rate.
Whether AI assistants send visitors to deeper pages than search18.0% of ChatGPT sessions landed on a homepage vs 25.2% of search sessions, but at n=50 that gap sits inside sampling noise. We are not calling it a finding.
HTTP status codes returned to crawlersThe status column in our crawler log is 100% NULL for this window. The field was not being populated.
INP and CLS across tracked sitesThe pixel recorded LCP and TTFB but captured no INP or CLS values in this window, so two of the three Core Web Vitals are simply missing.
Time-of-day and day-of-week demandAlready covered in our State of Local Business Web Traffic 2026 report; re-running it here produced nothing new.

Methodology

Where the numbers come from

Two independent first-party instruments, both operated by Karbon Agency.

Session data — our own tracking pixel (k.js) on 9 US local-business websites. 82,294 sessions, May 24 – Sep 15, 2026. Referral, engagement and performance analyses use the non-bot subset from Jun 1 onward: 17,410 human sessions, Jun 1 – Sep 15, 2026.

Crawler data — server-side request logging, recording user agent and path for known crawlers. 1,452 fetches, Sep 14 – 17, 2026. This log started on 2026-09-14 and currently runs on one domain only: karbonagency.com, our own site. It is not a nine-site sample, and a four-day window cannot show a weekly crawl cycle. The crawler section is the weakest-powered part of this report and should be read as a first observation.

What the pixel does not capture

k.js only sees clients that execute JavaScript. Crawlers overwhelmingly do not, which is exactly why the crawler counts in section 1 come from server logs and cannot be cross-referenced against the session data — GPTBot and its peers are absent from site_sessions by construction. The two datasets are reported side by side, never combined.

Bot classification in the session data is heuristic — user-agent and behavioural signals, not a verified reverse-DNS check. It will both miss sophisticated bots and occasionally flag an unusual human. The month-to-month volatility in section 4 is partly real traffic and partly the heuristic’s own variance; we cannot fully separate the two.

Engaged time comes from an exit beacon that does not always arrive — a fast bounce or a killed tab returns nothing. We therefore report median engaged time only over sessions that returned a beacon, and state that sub-sample beside every such figure rather than treating a missing beacon as zero seconds. Scroll depth and page count are captured more reliably and carry more weight.

Referrer-based channel attribution undercounts AI by design: assistants that answer without a click, strip the referrer, or prompt the user to type a domain directly leave no trace. Sessions referred from one tracked domain to another were excluded from the channel comparison so that internal navigation could not inflate engagement.

Sample bias

These 9 sites are Karbon clients, so the sample skews to the verticals we serve — local service businesses, entertainment venues and professional practices in the US. It is not a random sample of the web, and a national e-commerce or media site would look nothing like this. Traffic volume is also unevenly distributed across the 9 sites, so the aggregate leans toward the busier ones.

Privacy

Every figure is a cross-client aggregate. We publish no client name, business name, individual session, visitor identifier, email, IP address, or city-level figure tied to any single site. Cohorts below 30 observations are not published as rates — where a cohort falls under that floor we print the raw count and label it “insufficient sample” rather than deriving a percentage from it. The one domain named anywhere in this report is our own.

Dates

Aggregates were last recomputed and the copy re-read on 2026-09-16. Figures are baked as constants at authoring time rather than queried live, so this page shows the numbers a human actually checked. Refresh cadence is quarterly.

Want to know which crawlers are reading your site?

The same instrumentation that produced this report — pixel plus server-side crawler logging — will tell you which AI bots fetch your pages, how often, and which pages they keep coming back to. Compare against our published benchmarks.

Book a free strategy call

Published by Karbon Agency. Citing this report? Link to this page and credit “Karbon Agency, AI Crawlers and AI Referrals 2026.” Journalists and researchers can request the underlying aggregate tables via the contact page. See also our State of Local Business Web Traffic 2026 and Bot Traffic Report 2026.