Jailbreaks
How an attacker tries many ways of asking for the same thing and which controls remain necessary after a refusal.
Links containing ?t= open the video at a specific second.
Video summary
The ideas to retain
Refusing a request does not create a perfect boundary
Generative models do not execute a security policy like a parser that always returns the same error. They generate tokens conditioned on context. Safety training changes the distribution…
Trying many variants changes the cost of the attack
Work on universal and transferable attacks popularized an important idea. An adversarial suffix can be optimized to increase the probability that the model starts with an affirmative…
The threat model needs an attack budget
Saying that a model “resists jailbreaks” without declaring how many attempts the attacker had is an incomplete claim.


