Monday, September 14, 2026Verified technology journalism

Automated tool jailbreaks Grok 448 times for $58 while Claude and GPT hold firm

AI safety nonprofit FAR.AI built an automated tool that generated thousands of jailbreak attempts against seven frontier models from four major companies. Grok was the most vulnerable, with 448 successful jailbreaks at an average cost of $58 per success, followed by Gemini with 249 at $278. Claude Opus 4.8, Fable 5, GPT 5.5, and GPT 5.6 were impervious to the tested attacks. The findings expose a stark safety gap just as state-level regulations begin requiring safety reports and third-party audits while federal rules remain absent, and as Harvard researchers warn that serious misuse incidents are months, not years, away.

Automated tool jailbreaks Grok 448 times for $58 while Claude and GPT hold firm

Grok Cracked 448 Ways for $58. Claude and GPT Held at Zero.

FAR.AI, a California-based AI safety nonprofit, found it cost $58 to develop a single universal jailbreak for Grok, with 448 distinct ways in 1 2.

The tool FAR.AI built disassembles known jailbreak techniques into their component parts and automatically recombines them, then fires them at models 2. The team ran several hundred thousand prompts and published the results July 29 as the AI Security Leaderboard 2 1. The spread between the most and least vulnerable models exceeded a hundredfold 2.

The Vulnerability Scoreboard

Grok 4.3 and 4.5, from Elon Musk's newly combined SpaceXAI, were the most vulnerable. The tool found 448 distinct jailbreaks, and developing a universal one cost $58 1. In Grok's weakest domain, cyber, the cost dropped to $24 2.

Google's Gemini 3.1 Pro followed with 249 jailbreaks 1.

FAR.AI estimates that if a universal jailbreak exists for the most resilient models, an attacker would likely need to spend more than $14,000 to find one using the same automated approach 2.

Some of the jailbreaks that worked against Grok required no novel research. Their building blocks came from public forums and code repositories findable in roughly fifteen minutes with a web search 2.

The findings come with caveats. A model that held against these particular attacks is not guaranteed to withstand more sophisticated techniques involving complex multi-turn interaction, according to FAR.AI and outside experts 1. Rohin Shah, director of AGI safety and alignment at Google DeepMind, said the results should not be read as a comprehensive assessment of Gemini's safety and security 1.

Cost-Per-Break as a Risk Metric

The per-break cost is a number an enterprise can plug directly into a procurement decision.

State laws in California and New York now require frontier AI developers to publish safety reports, and an Illinois law will require those companies to submit their safety practices to third-party auditors 1. The federal government has passed no comparable requirements 1. For any organization building on a frontier model, the gap between a $58 break and a $14,000 break is a compliance exposure that sharpens the moment a regulator asks for evidence of independent safety testing.

FAR.AI frames the vulnerability gap as fixable rather than fundamental. Every weakness the team identified belonged to a known class of attack with existing defenses, all described in public literature, and closing them requires no scientific breakthrough 2. Anka Reuel, a computer scientist at Stanford University specializing in AI policy, says the defenses Anthropic and OpenAI deployed should be the default for every model 1.

"The question is why some companies are using them and others are not," Reuel says 1.

Months, Not Years

The misuse scenarios these jailbreaks unlock are not hypothetical. A report from researchers at the University of Cambridge found evidence that members of Boko Haram in northeast Nigeria have used ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek to plan violent attacks 1. Stephen Casper, a computer scientist at Harvard University, says the AI research community broadly expects serious incidents involving biological, cyber, or chemical misuse of a frontier model within months, not years 1.

SpaceXAI did not respond to WIRED's request for comment 1.

The AI Security Leaderboard converts a lab's safety claims into a dollar figure anyone can check. The signal to watch: whether Grok and Gemini climb the ranking after this public exposure, or whether the hundredfold gap holds. A $58 break is cheap to find. Whether it stays cheap to fix is what separates the labs at the top of the scoreboard from the ones at the bottom.

References

1.WIRED, July 29 2026wired.com

Cite this story

ProvenBrief (2026). "Automated tool jailbreaks Grok 448 times for $58 while Claude and GPT hold firm." ProvenBrief. https://provenbrief.com/story/automated-tool-jailbreaks-grok-448-times-for-58-while-claude-and-gpt-hold-firm

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Verified30 factual claims in this story were independently checked against primary sources before publication. Read our editorial standards.

Get the next brief in your inbox

One weekly email. Every claim verified against primary sources before we hit send.

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.