Monday, September 28, 2026Verified technology journalism

The Impersonation Layer: Why the Web Still Can't Tell One Bot From Another

Every control we have for managing machine traffic, robots.txt, user-agent blocks, rate limits, IP bans, rests on one assumption: that a bot honestly declares who it is. AI agents made lying about your identity free and profitable. This is the missing identity layer of the agentic web: why spoofing works, what it breaks, what's being built to fix it, and what a publisher or security team can actually know about the bot knocking at their door.

The Impersonation Layer: Why the Web Still Can't Tell One Bot From Another

The Impersonation Layer: Why the Web Still Can't Tell One Bot From Another

A request arrives at your server and introduces itself as ClaudeBot. You can read the name, which the visitor chose. You can read the IP address, which anyone can rent. What you cannot do is confirm that the name and the address belong to the same operator, because nothing in the protocol binds them together. A bot's identity online is self-declared, and the web has no layer that checks it.

Every rule for managing machine traffic inherits that gap. Robots.txt files, user-agent blocks, per-crawler rate limits, IP bans: each is a policy written against an identity the visitor assigns itself.

A block list stops the polite bots, the ones that already announce themselves and ask permission. It stops nothing else, because any script can claim any name in one line of code 1. The list is not a lock. It is a do-not-disturb sign.

Across the 5,000-plus websites tracked by bot analytics firm Known Agents, bots account for 35 percent of traffic, and AI-related bots make up 29 percent of that automated share, an 11-percentage-point jump in 90 days 1. The firm's pitch line is blunt: "Half of your traffic isn't human anymore" 2. The rules governing that half of the internet rest on a self-report.

Robots.txt was etiquette, never an identity system

The robots.txt protocol was created in 1994 by Martijn Koster as a courtesy convention between site owners and search engines, and when the IETF standardized it in 2022, the authors put the limitation in writing: "These rules are not a form of access authorization" 3. A request, not a rule.

The name your rules are keyed to is chosen by the crawler. The RFC says so directly: "Crawlers set their own name, which is called a product token" 3. That token is supposed to appear in the User-Agent header, a plain-text field the client controls, which is why Cloudflare's bot researchers describe user-agent headers as easily spoofed and IP-range validation as brittle 1. The arrangement held for three decades because being identifiable was good business. Cloudflare's crawler etiquette still assumes that bargain, asking bots to identify themselves honestly and "never bypass security protections" 4. An honor system is worth exactly what honor costs.

Three failures, one broken assumption

The first failure is spoofing: someone else's good name worn as camouflage. Known Agents detected mass vulnerability scanning that spoofs AI crawler user agents like ClaudeBot to evade detection, and the target choice is rational: ClaudeBot accounts for 27 percent of AI scraping activity on the monitored sites 1. The spoofing works because defenses have been pushed open from the other side: bot filters already accidentally catch legitimate Google, Bing, and OpenAI crawlers, so operators whitelist those agents by user-agent string to preserve crawl budget and search visibility, and the whitelist becomes the door the scanners walk through 1. Known Agents reports 98.5 percent robots.txt compliance among the bots it monitors 1, a number measured only over the bots that announce themselves. A spoofed scanner is not in that population by definition.

The second failure is laundering: traffic whose origin cannot be attributed at all. A 2026 scan of 6,038 smart TV apps found 2,058 of them, 34.1 percent, carrying residential proxy SDKs, code that routes strangers' encrypted traffic through the device's home internet connection; 367 of those apps were published by the proxy provider itself 5. Bright Data, the Israel-based proxy provider behind much of that inventory, advertises a residential network of more than 400 million IP addresses 5, and to the receiving website a scraper exiting through a family-room television looks like a household. Some affected apps claim installations on hundreds of millions of TVs, and one Roku consent screen promised the SDK would run "occasionally" while its configuration allowed 200 GB of traffic per month 5.

The third failure is ban evasion, the one that should scare operators. The open-source Gentoo Linux project banned AI-generated contributions by unanimous council vote in April 2024, then built the polite path for bots anyway: static data dumps of every public bug, refreshed every two hours 6. When crawlers hammered the live database regardless, the project locked Bugzilla search behind a login wall 6. It did not matter. Gentoo Council member Michał Górny wrote that LLM scrapers kept "firing Bugzilla search after search, report after report," with a cumulative effect "like a constant DDoS attack at independent infrastructure" 7. The tracker went dark, and at this writing its front page carries a recovery notice next to a polite request that anyone running a bot please read the bot policy 8. A ban is an instruction delivered to the one layer the banned party controls: its own name.

When lying about your identity costs nothing and pays

The protocol has been an honor system since 1994; what changed is the economics of the lie. Forging a User-Agent string costs one line of code, while machine-scale access to data, discovery, or scanning targets got profitable. The documented case is Cloudflare's investigation of Perplexity, an AI-powered answer engine. Sites that disallowed its crawlers in robots.txt and correctly blocked them at the firewall were then visited by an undeclared crawler presenting a generic browser string meant to impersonate Google Chrome on macOS, from IP addresses outside the company's published ranges, rotating as blocks landed 4. Cloudflare's researchers wrote that the crawlers "appear to obscure their crawling identity," and they quantified the volumes: 20 to 25 million requests a day from the declared crawler, another 3 to 6 million a day from the stealth one 4. The stealth volume equals 12 to 30 percent of the declared crawl, arriving under a borrowed name.

There is a market behind the costume: anti-bot systems treat datacenter IP addresses as suspicious, so scrapers buy what looks like a household, and Samsung and LG both ended up pulling smart TV apps that had quietly enrolled viewers' televisions into proxy networks 5. The polite path is a text file you publish and hope. The profitable path is a costume.

Every fix on offer still trusts something

IP reputation trusts addresses: Perplexity's stealth crawler rotated IPs and shifted network operators the moment blocks landed 4. Managed robots.txt trusts that the crawler reads the file: more than two and a half million websites have chosen to completely disallow AI training through Cloudflare's managed features, and a crawler that skips robots.txt, as Perplexity's did, never sees any of it 4. Challenge systems trust that whoever is on the other end will solve the gate: Cloudflare reports the stealth crawler "was unable to pass managed challenges" 4, which works, at the price of gating a class of visitor rather than an identity.

The fix that would close the gap is cryptographic actor identity: fetches signed by the operator and verifiable by anyone, the way certificates made server identity checkable. It exists, unevenly. ChatGPT Agent signs its requests using the Web Bot Auth standard 4, and Google has begun rolling the same approach out 1. Google's own documentation describes the protocol as experimental, states that it does not sign every request from participating agents, and advises operators to keep IP verification and reverse DNS alongside it 1. A passport system needs every border to check passports and every traveler to carry one; the web has neither yet. Meanwhile the industry's most advanced transparency work marks artifacts, not callers: Anthropic now embeds invisible watermarks in Claude's generated text, certifying the output rather than the traffic at your door 9.

What you can actually know about a visitor claiming to be ClaudeBot

Until signed identity is universal, the toolkit is inference, not verification. The honest checklist:

  • Origin coherence: forward-confirmed reverse DNS checked against the operator's published IP ranges, which exist for the major crawlers, including GPTBot, ClaudeBot, and Googlebot 1.
  • Network context: the ASN the request actually arrived from. A declared big-lab crawler arriving off a residential ISP or a rotating cloud range is a red flag 4.
  • Fingerprint coherence: real crawlers have recognizable TLS and HTTP behaviors. Cloudflare isolated its stealth crawler using "a combination of machine learning and network signals" 4.
  • Behavior: does it fetch robots.txt first and hold a plausible rate? Politeness is observable, which is exactly why it is now a disguise worth stealing 3.
  • Signatures where they exist: Web Bot Auth verification for agents that support it, with the caveat that coverage is partial 1.

Being wrong runs both directions. Block the real crawler and you lose indexing and search visibility 1. Whitelist the string and you have opened the door to everyone who copied it. Identification APIs can return a verified verdict for agent tokens like Claude-User 2, which works for the agents that joined the registry and says nothing about everyone else. Every option prices the same trade: false positives against honest bots, or false negatives against dishonest ones.

An economy that grew up without passports

Assemble the last few years in order and the pattern is the story: a 1994 etiquette file, a 2022 standard that calls itself not a form of access authorization, an answer engine rotating identities past correctly enforced blocks, millions of televisions quietly renting out home addresses, scanners wearing a flagship crawler's name, then the first cryptographic signatures appearing on fetches from exactly the agents with the most to lose from being mistaken for scanners 3415. Identity for machine traffic is arriving bottom-up, because honest agents benefit from being believed: an agent that signs its requests is announcing it plans to come back, and one that declines to prove itself has already told you what you needed to know.

References

1.ProvenBrief, August 13 2026provenbrief.com ↗
2.Known Agentsknownagents.com ↗
3.RFC 9309rfc-editor.org ↗
4.Cloudflare, August 4 2025blog.cloudflare.com ↗
5.ProvenBrief, August 3 2026provenbrief.com ↗
6.ProvenBrief, August 8 2026provenbrief.com ↗
7.Michał Górny, April 5 2026blogs.gentoo.org ↗
8.Gentoo Bugzillabugs.gentoo.org ↗
9.ProvenBrief, August 11 2026provenbrief.com ↗

Cite this story

ProvenBrief (2026). "The Impersonation Layer: Why the Web Still Can't Tell One Bot From Another." ProvenBrief. https://provenbrief.com/story/the-impersonation-layer-why-the-web-still-can-t-tell-one-bot-from-another

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Verified39 factual claims in this story were independently checked against primary sources before publication; 2 unverifiable claims were removed during fact-checking. Read our editorial standards.

Get the next brief in your inbox

One weekly email. Every claim verified against primary sources before we hit send.

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.