All videosAI security0:48

Jailbreaks

How an attacker tries many ways of asking for the same thing and which controls remain necessary after a refusal.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

Refusing a request does not create a perfect boundary

Generative models do not execute a security policy like a parser that always returns the same error. They generate tokens conditioned on context. Safety training changes the distribution…

02

Trying many variants changes the cost of the attack

Work on universal and transferable attacks popularized an important idea. An adversarial suffix can be optimized to increase the probability that the model starts with an affirmative…

03

The threat model needs an attack budget

Saying that a model “resists jailbreaks” without declaring how many attempts the attacker had is an incomplete claim.