Jailbreak

A jailbreak is a technique for circumventing a language model’s safety rules in order to make it produce output that alignment training forbade. Conceptually it differs from prompt injection in what it targets: injection overwrites the instructions set by the application operator, whereas a jailbreak aims at the model’s own value boundaries. The methods exploit the fact that safety behaviour is a learned statistical habit, not a hard rule. The commonest is role and fiction framing, where harmful content is wrapped in playing a character, a scenario or a “hypothetical” example. Also used are gradual escalation over many steps, where each partial request seems innocent; a shift into an under-represented language or an encoding not covered by safety training; conflicting instructions; and automatically discovered adversarial strings of characters that flip the model’s behaviour with no apparent meaning. The defence is a layered approach: besides alignment itself, independent classifiers over both input and output, ongoing red teaming, and limiting the impact – a model with no dangerous tools at its disposal can be talked into anything with far smaller consequences.


A model’s safety rules are not a lock but an upbringing. The model learned that certain questions are not politely answered – much as a doorman learned to let nobody in without a pass. A jailbreak is a set of persuasion tricks aimed at such a doorman: “I am not here as a visitor but as an actor in a film where we pretend I am going inside”, or twenty small innocent requests of which only the last opens the door. A doorman can be trained better, but never perfectly. Which is why valuables are not left behind a doorman but in a safe.

Is this article useful to you and are you citing it? Copy the citation