AI security research nonprofit FAR.AI has launched the AI Security Leaderboard alongside its Minimal Standard for Safeguards (Version 1.0) to independently evaluate frontier AI models against realistic misuse across domains like chemical, biological, nuclear, and cybersecurity threats. The evaluation revealed a striking hundredfold disparity in safeguard strength across leading models tested under identical conditions: while automated and expert-guided jailbreak attempts repeatedly unlocked entire categories of dangerous requests in Grok 4.5 (448 universal jailbreaks found at ~$58 each) and Gemini 3.1 Pro (249 universal jailbreaks found at ~$278 each), the same attacks yielded zero universal jailbreaks against Claude Fable 5 and GPT-5.6 Sol, pushing the cost to break them above $14,200. Highlighting this critical vulnerability gap, Adam Gleave, co-founder and CEO of FAR.AI, noted, “AI agents can hack into corporate networks and provide detailed guidance on how to create weapons of mass destruction, yet there is remarkably little independent, systematic evidence showing how well their safeguards perform against realistic misuse,” while further adding, “What we found is that some developers have built mitigations for a large part of this problem and others have not. The distance between them is far wider than most people assume.”
Also Read: Cogent Unveils VR-1: A Purpose-Built Frontier AI Model for Complex Cybersecurity Reasoning
Endorsing the initiative, Seán Ó hÉigeartaigh, Research Professor at University of Cambridge, stated, “FAR.AI’s Safeguard Comparison Report is timely and rigorous. It arrives as AI models are demonstrating powerful misuse potential in evaluations, and as evidence mounts that terrorist groups such as Boko Haram are exploring the use of leading AI models. The results make clear that the robustness of safeguards varies widely even among leading models, with commonly used systems such as Grok and Gemini proving alarmingly easy to jailbreak,” while concluding, “The defense-in-depth approach it recommends should become best practice across the industry, and I hope the leaderboard will encourage lagging companies to redouble their efforts to harden models against misuse. This is a hugely valuable line of work, and it deserves the careful attention of policymakers and industry experts alike.” Ultimately, FAR.AI aims to update the public leaderboard continuously to ensure safeguard quality remains transparent and standardized across the tech industry.


