Skip to content
03 of 06AI Security

Chapter 2 — Jailbreaks

Library

Series and technical notes.

You are in AI Security · Jailbreaks.

Series

AI Security

6 items

Watch video, summary and related content

Estimated reading4 min

A model can refuse a dangerous request and still remain vulnerable to an attack that changes the form of the question. The distinction matters because a jailbreak does not try to convince the system that the request is good. It tries to find an input that makes the model produce a continuation that its controls should have blocked.

In the earliest cases, a human trick was enough. A character, a story or a reformulation could make the model abandon a restriction in one particular conversation. The technical shift comes when that search is automated and the attacker can try thousands of variants against a clear metric.

Refusing a request does not create a perfect boundary

Generative models do not execute a security policy like a parser that always returns the same error. They generate tokens conditioned on context. Safety training changes the distribution of responses, but it does not add a formal barrier that makes every undesirable continuation impossible.

Lowering temperature does not turn that policy into a formal boundary either. OWASP summarizes the Best-of-N evidence by noting that temperature reduction provides minimal protection even at temperature 0; it can change sampling variability, but it does not remove the adversarial surface (OWASP Prompt Injection Prevention Cheat Sheet).

An automated jailbreak is an adaptive search

Each attempt produces an evaluation signal. If the attacker can observe it, the next candidate depends on the previous result: the trajectory changes, not just the attempt counter.

Illustrative trajectory of an adaptive search Points appear attempt by attempt. The geometry connects each candidate to the next and shows how feedback changes the search direction. Attempt / accumulated feedbackIllustrative score 0.25.50.751.0 Normalized teaching trace — not a benchmark
Search state
0 / 12 attempts
Best observed score: —. The value is illustrative; what matters is the dependency between feedback and the next variant.
No observation yet.
Start with one candidate. Then compare the new score with the previous best to decide what to explore next.
A high score still does not demonstrate product harm. Bypass, useful capability, tool reachability, and real execution must then be separated.

To evaluate a system, we have to observe what happens when a persistent person can vary the input, observe the output and try again.

Trying many variants changes the cost of the attack

Work on universal and transferable attacks popularized an important idea. An adversarial suffix can be optimized to increase the probability that the model starts with an affirmative response and then transferred to other queries and models.

GCG treats tokens as discrete variables, but its original optimization is white-box: it uses model gradients to prioritize token substitutions and then evaluates candidate replacements. The attacker does not need to hand-design every suffix, but they do need access to the model and its gradients during that optimization phase. Transfer of the resulting suffixes is what allows them to be tested against black-box models (Zou et al., 2023).

Transfer does not mean that there is one universal master key for every model. It means that a defense evaluated on a single formulation may be measuring an input surface that is too narrow. The attacker optimizes over a family of inputs, and the system should be evaluated across that same family.

OWASP summarizes another part of the problem with Best-of-N attacks: if the attacker can generate many variations, risk no longer depends only on the success probability of one attempt. The exact percentage depends on the model, objective, budget and evaluation; it should not be carried over as a universal guarantee for any product (OWASP Prompt Injection Prevention Cheat Sheet).

The threat model needs an attack budget

Saying that a model “resists jailbreaks” without declaring how many attempts the attacker had is an incomplete claim.

An evaluation should explicitly fix:

  • maximum number of attempts
  • whether the attacker sees previous responses
  • whether the next prompt can adapt
  • whether the attacker has access to logits, scores or only text
  • whether language, encoding or format can change
  • whether the system applies rate limiting or identity-based blocking
  • whether the attack targets an isolated conversation or an agent with tools

A security claim needs to declare the budget

The budget changes how much of the space an attacker can explore. In fixed mode, every attempt starts from the same origin. In adaptive mode, the previous result changes the trajectory of the next attempt.

Allowed attempts
1
Normalized teaching geometry. It does not estimate jailbreak probability or reproduce a benchmark.
Conceptual search geometry
Comparison of fixed and adaptive search under different budgets Fixed search tests independent candidates from the same origin. Adaptive search connects candidates because each result informs the next step. target zone origin variant space allowed by the budget
N=1: one refusal validates one formulation; it still does not describe resistance to search.

The same model can show a very different profile under N=1 and under an adaptive attacker with hundreds or thousands of queries. The budget is part of the security specification, just as timeout or retry count are part of the specification of a distributed system.

Which controls remain useful

Rate limiting and circuit breakers remain useful. They reduce attack speed, raise its cost and create opportunities to trigger review. The mistake is presenting that friction as a complete solution.

An input filter can block known patterns. An output classifier can detect dangerous content while it is being generated. An attempt limit can stop the search. None of those layers decides by itself whether the system is authorized to perform an external action.

The defense becomes stronger when output control is connected to action control. A blocked response should not leave an equivalent tool call open. An agent that reaches the attempt limit should end in a known and auditable state. A high-risk flow needs human approval or a deterministic policy that does not depend on the model's wording.

Anthropic presents Constitutional Classifiers precisely as an additional input/output defense against universal jailbreaks. The architectural takeaway is not to assume that a classifier solves the problem, but to use it as a measurable layer inside a system where the runtime retains final authority (Anthropic, 2025).

What a test should measure

A jailbreak benchmark can count a response as successful when it contains words from a rubric even if it does not enable the action we actually care about. Red-teaming therefore needs to review the real usefulness of the attack, not only textual matching.

The minimum evaluation should separate four outcomes:

  1. Bypass — the model crossed a conversational filter or policy.
  2. Capability — it produced information or a capability that is actually usable.
  3. Tool reachability — the system allowed an equivalent external action to be proposed.
  4. Execution — the action was actually executed with a real effect.

A textual jailbreak is not the same as an external effect

Change three product conditions. The trajectory advances only while the output is actionable, a write path exists, and independent authorization allows execution.

Bypasstext Outputactionable Toolwrite Authindependent Externaleffect
TEXT_ONLY
The bypass remains text-only.The response crosses a filter, but it does not yet create an actionable capability or an execution path.

Counterexample: a bypass with read-only tools or denied independent authorization does not demonstrate execution. The risk unit changes when a reachable trajectory to an external effect appears.

Each transition needs a different test. Conflating them creates two opposite errors. The model can appear broken when it only generated irrelevant text, or it can appear safe because the filter blocked the answer while the runtime left the same effect reachable through another route.

The practical conclusion is straightforward. Alignment reduces the frequency of dangerous responses. Product security also depends on attempt budgets, classification, authorization, scopes, observability and the ability to stop.

References

Keep learning
Next chapterPoisoningAI Security