Jailbreak Anthropic's new AI safety system for a $15,000 reward

2025-02-04T19:35:13+00:00 View Original

Full Report

In testing, the technique helped Claude block 95% of jailbreak attempts. But the process still needs more 'real-world' red-teaming.

Analysis Summary