{"id":9583,"date":"2026-07-27T12:30:27","date_gmt":"2026-07-27T10:30:27","guid":{"rendered":"https:\/\/www.kubicek.ai\/?post_type=lexicon&#038;p=9583"},"modified":"2026-07-27T13:37:15","modified_gmt":"2026-07-27T11:37:15","slug":"jailbreak","status":"publish","type":"lexicon","link":"https:\/\/www.kubicek.ai\/en\/lexicon\/jailbreak\/","title":{"rendered":"Jailbreak"},"content":{"rendered":"<p class=\"wp-block-paragraph\">A <strong>jailbreak<\/strong> is a technique for circumventing a language model&#8217;s safety rules in order to make it produce output that alignment training forbade. Conceptually it differs from prompt injection in what it targets: injection overwrites the instructions set by the application operator, whereas a jailbreak aims at the model&#8217;s own value boundaries. The methods exploit the fact that safety behaviour is a learned statistical habit, not a hard rule. The commonest is role and fiction framing, where harmful content is wrapped in playing a character, a scenario or a &#8220;hypothetical&#8221; example. Also used are gradual escalation over many steps, where each partial request seems innocent; a shift into an under-represented language or an encoding not covered by safety training; conflicting instructions; and automatically discovered adversarial strings of characters that flip the model&#8217;s behaviour with no apparent meaning. The defence is a layered approach: besides alignment itself, independent classifiers over both input and output, ongoing red teaming, and limiting the impact \u2013 a model with no dangerous tools at its disposal can be talked into anything with far smaller consequences.<\/p>\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n<p class=\"wp-block-paragraph\">A model&#8217;s safety rules are not a lock but an upbringing. The model learned that certain questions are not politely answered \u2013 much as a doorman learned to let nobody in without a pass. A jailbreak is a set of persuasion tricks aimed at such a doorman: &#8220;I am not here as a visitor but as an actor in a film where we pretend I am going inside&#8221;, or twenty small innocent requests of which only the last opens the door. A doorman can be trained better, but never perfectly. Which is why valuables are not left behind a doorman but in a safe.<\/p>\n","protected":false},"featured_media":0,"template":"","class_list":["post-9583","lexicon","type-lexicon","status-publish","hentry"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon\/9583","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon"}],"about":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/types\/lexicon"}],"wp:attachment":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/media?parent=9583"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}