Sunday, September 27, 2026Verified technology journalism

Someone is running mass vulnerability scans disguised as ClaudeBot, exposing a trust crisis in the agentic web

Bot analytics firm Known Agents, monitoring traffic across 5,000-plus websites, has detected mass vulnerability scanning operations that spoof AI crawler user agents like ClaudeBot to evade detection. AI-related traffic now accounts for 29 percent of all bot activity, up 11 percent in 90 days, giving attackers an expanding pool of legitimate crawler identities to impersonate. The Hacker News discussion that surfaced the finding, drawing 240 upvotes and 177 comments, revealed a deeper problem: website defenses meant to block malicious bots are already accidentally catching real AI crawlers from Google, Bing, and OpenAI, creating pressure to whitelist AI agents that attackers then exploit. Google has begun rolling out cryptographic Web Bot Auth to verify crawler identity, but the agentic web's trust model still rests on user-agent strings anyone can forge.

Someone is running mass vulnerability scans disguised as ClaudeBot, exposing a trust crisis in the agentic web

Attackers are running mass vulnerability scans across the web disguised as ClaudeBot and other legitimate AI crawlers, exploiting a trust model that was already broken before anyone started impersonating it. Bot analytics firm Known Agents, monitoring traffic across more than 5,000 websites, detected the spoofed scanning operations and identified the structural flaw underneath: the identity layer for automated web traffic rests on user-agent strings that anyone can forge in one line of code 1.

The scale of the opportunity for attackers is visible in the data. Bots account for 35% of all web traffic across the monitored sites, and AI-related bots make up 29% of that automated traffic, an 11-percentage-point jump in 90 days 1. Multiplying those two figures yields a number Known Agents does not state directly: AI bots represent roughly 10% of every request hitting these sites, or about one in ten visits. That expanding footprint gives attackers a growing library of legitimate crawler identities to impersonate.

ClaudeBot accounts for 27% of all AI scraping activity on the monitored sites, making it the single most active AI data scraper 1. When attackers spoof that user-agent string, their requests look to any standard filter like routine Anthropic crawling. A Hacker News discussion that brought the finding to wider attention surfaced the operational consequences: site operators reported that bot-detection tools meant to filter malicious traffic were already catching legitimate Google, Bing, and OpenAI crawlers by mistake 2.

The defense creates the opening

That misfiring is the structural flaw, not a configuration error. When bot-detection systems like Cloudflare's Bot Fight Mode flag or block legitimate AI crawlers, site operators face pressure to whitelist those agents by user-agent string to preserve their crawl budget and search visibility. But a user-agent string is self-reported plain text. Forging ClaudeBot requires changing one string in an HTTP header, a task that takes seconds 3. The whitelist designed to protect legitimate crawlers becomes the camouflage attackers wear to bypass it.

Known Agents reports 98.5% robots.txt compliance among the bots it monitors 1. That figure sounds reassuring until you trace which bots it measures: self-identifying agents that announce themselves and respect publisher rules. Spoofed vulnerability scanners, by definition, do neither. The 98.5% compliance rate describes the population of bots that are not the problem.

The contradiction extends further. Search Engine Journal's verified crawler list shows that some of the most prominent AI agents, including ChatGPT's Operator, Bing's Copilot chat, and Grok, operate with unidentifiable user-agent strings, leaving site operators no way to track or manage them by name 4. Sites that want to participate in AI discovery must either accept traffic they cannot identify or risk blocking agents they cannot see.

Cryptographic identity, partial deployment

Google has begun rolling out Web Bot Auth, an IETF draft protocol that lets bots cryptographically sign their requests so websites can verify identity independently of IP addresses or user-agent strings 5. Cloudflare proposed the same approach in May 2025, describing user-agent headers as easily spoofed and IP-range validation as brittle because cloud infrastructure addresses shift over time 3. OpenAI's Operator already signs its requests using the HTTP Message Signatures standard, according to Cloudflare's account 3.

But the protocol's own documentation describes it as experimental and explicitly incomplete. Google states that it does not sign every request from participating agents and advises site operators to continue relying on IP verification and reverse DNS alongside Web Bot Auth 5. Not all Google agents use the protocol. Adoption outside Google's infrastructure remains limited. No timeline exists for full coverage, and every day without universal cryptographic verification is a day this cycle continues.

For site operators and builders, the practical takeaway is direct: whitelisting AI crawlers by user-agent string is now an active security liability, not a courtesy. Until cryptographic bot identity reaches broad adoption, IP-based verification is the interim step, and Search Engine Journal's data shows official IP lists are available for major crawlers including GPTBot, ClaudeBot, and Googlebot 4. The vulnerability scanners disguised as ClaudeBot will keep probing regardless. The question is whether the site can tell the difference.

References

1.Known Agentsknownagents.com ↗
2.Hacker Newsnews.ycombinator.com ↗
3.Cloudflareblog.cloudflare.com ↗
4.Search Engine Journalsearchenginejournal.com ↗
5.Google Developersdevelopers.google.com ↗

Cite this story

ProvenBrief (2026). "Someone is running mass vulnerability scans disguised as ClaudeBot, exposing a trust crisis in the agentic web." ProvenBrief. https://provenbrief.com/story/someone-is-running-mass-vulnerability-scans-disguised-as-claudebot-exposing-a-tr

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Verified25 factual claims in this story were independently checked against primary sources before publication. Read our editorial standards.

Get the next brief in your inbox

One weekly email. Every claim verified against primary sources before we hit send.

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.