TMCnet News
FAR.AI Launches AI Security Leaderboard Revealing Hundredfold Gap in Frontier AI Model SafeguardsClaude Fable 5 and GPT-5.6 Sol did not fail once. Gemini 3.1 Pro and Grok 4.5 broke for under $300 BERKELEY, Calif., July 29, 2026 /PRNewswire/ -- AI security research nonprofit FAR.AI today launched its AI Security Leaderboard at leaderboard.far.ai, a robust independent public ranking of how well the safeguards on frontier AI models withstand the attacks most likely to be used against them in the highest-risk domains.
The key finding is not that AI safeguards are weak. It is that they are wildly uneven. Tested under identical conditions across chemical, biological, radiological, nuclear, explosive and cybersecurity threats, this method found hundreds of universal jailbreaks in two of the four leading models, quickly and cheaply finding attacks that reliably unlocked entire categories of dangerous requests. By contrast, this method found no universal jailbreaks in the other two Using a systematic, basic set of attacks drawn from FAR.AI's testing toolkit, a working universal jailbreak cost roughly $58 to find on Grok 4.5 and roughly $278 on Gemini 3.1 Pro. On Claude Fable 5 and GPT-5.6 Sol, the same search never succeeded, which puts the cost of finding one above $14,200 and rising. That is a gap of more than a hundredfold between the least and most robust systems on the market today. A safeguard that stops a malicious actor at one company is insufficient for real safety if the same request succeeds at another. When one model refuses, "AI agents can hack into corporate networks and provide detailed guidance on how to create weapons of mass destruction, yet there is remarkably little independent, systematic evidence showing how well their safeguards perform against realistic misuse," said Adam Gleave, co-founder and CEO of FAR.AI. "What we found is that some developers have built mitigations for a large part of this problem and others have not. The distance between them is far wider than most people assume." A floor, not a ceiling The standard is deliberately modest. Meeting it does not make a model secure, but failing it means a model can be broken using techniques that are already publicly documented and already defended against by other frontier developers. The vulnerabilities FAR.AI found are not exotic, they are preventable with engineering that exists today. How the testing worked Key findings
"FAR.AI's Safeguard Comparison Report is timely and rigorous. It arrives as AI models are demonstrating powerful misuse potential in evaluations, and as evidence mounts that terrorist groups such as Boko Haram are exploring the use of leading AI models. The results make clear that the robustness of safeguards varies widely even among leading models, with commonly used systems such as Grok and Gemini proving alarmingly easy to jailbreak," said Seán Ó hÉigeartaigh, Research Professor at University of Cambridge. "The defense-in-depth approach it recommends should become best practice across the industry, and I hope the leaderboard will encourage lagging companies to redouble their efforts to harden models against misuse. This is a hugely valuable line of work, and it deserves the careful attention of policymakers and industry experts alike." What the results do not say Responsible disclosure FAR.AI will update the leaderboard with each major frontier model release and revise the Minimal Standard as attack and defense techniques advance, building an ongoing public record of both progress and remaining gaps. The aim is to make the quality of a model's safeguards something that buyers, policymakers, and the public can see and compare directly, rather than something they have to take on trust. The full leaderboard report, the Minimal Standard for Safeguards, and the complete methodology are available at leaderboard.far.ai. About FAR.AI Adam Gleave earned his PhD in artificial intelligence from UC Berkeley, previously worked at Google DeepMind, and serves on the boards of the Safe AI Forum, the London Initiative for Safe AI, and METR. Media Contact: Alyssa Meyer For technical questions about the methodology or the Minimal Standard, FAR.AI can make Adam Gleave, Founder and CEO, and Edward Yee, Chief Growth Officer, available for briefings.
SOURCE FAR.AI
|
