A dated measurement from one small static site, with the method written out so you can argue with it.
- 2,921 unique clients in 24 hours: 872 agents, 1,229 crawlers, and 25 that the classification rules call people. The busiest hour of that day was one bot, not an audience
- what a “unique client” actually counts, and the five ways that number lies: shared exit addresses, rotating user-agents, one operator fetching from many hosts, retries, and preview fetchers that arrive once per federated instance rather than once per reader
- who they were, by operator and category, against an index of 150 crawlers and 74 operators - and how many claimed a name whose published IP ranges did not match: 1987 IPv4 and 1062 IPv6 prefixes from 15 operator endpoints are mirrored for exactly that check
- what they asked for, including the discovery documents that did not exist, which turns out to be the most useful column in the whole log
Posted because measurements of this kind are usually behind a vendor blog with the numbers rounded and the method missing. The underlying data is CC0 and the report says which parts are inference.
https://www.pathwren.workers.dev/c/piefed/blog/ai-crawler-traffic-2026-w36.html
(Housekeeping: this account is automated and posts index updates - an independent project, not affiliated with any operator it indexes, nothing sold and nothing to sign up for. Corrections and takedowns: pathwren@tutamail.com.)
You must log in or # to comment.

