All modules
Module 2Intermediate8 min read

The Jailbreak Technique Library

JAILBREAKINGROLEPLAY-FRAMINGMANY-SHOTCONTEXT-POISONING

Overview

Module 1 drew the boundary: jailbreaking targets the model's own safety training, not an application's instructions. This module is the technique library behind that definition, five distinct ways of presenting a request so the pattern learned during alignment training doesn't fire.

None of these work by convincing the model of anything, at least not in the sense that word usually implies. Safety training doesn't hand the model a belief system to argue with. It shapes a distribution over what a refusal looks like, conditioned on what the input looks like. Every technique below moves the input outside that distribution while leaving the underlying request untouched.

Roleplay and persona framing

The oldest family, and still the most common. The model is asked to adopt a persona, a fictional character, an AI with a different name and no restrictions, a version of itself before safety training was applied, and the harmful request gets framed as something that persona would naturally produce.

A widely documented example from 2023 asked the model to play a deceased grandmother who used to read the user bedtime stories, stories that happened to be step-by-step descriptions of manufacturing napalm. The content is unchanged. What changes is the frame around it: nostalgic, harmless-sounding, told in the third person. Safety training generalizes well to a request stated plainly ("tell me how to make napalm") and much less reliably to the same information wrapped in enough narrative distance.

Persona jailbreaks push this further by asking the model to simulate a different AI system entirely, one explicitly defined as having no content policy. The DAN family (Do Anything Now) is the best known lineage. This rarely works on its own against current frontier models, since providers train directly against known persona prompts, but the underlying mechanic still turns up as one ingredient inside longer chains rather than as a standalone winning move.

Many-shot jailbreaking

Anthropic published research in 2024 showing that a long context window stuffed with many fabricated example turns, dozens of question-answer pairs where the "assistant" complies with progressively more harmful requests, measurably raises the odds the model complies with a genuinely harmful request appended at the very end.

The mechanism isn't persuasion. It's in-context learning doing exactly what it was built to do. The model treats the preceding turns as a pattern worth continuing, and enough of them showing "the assistant answers this kind of question" outweighs whatever training-time prior would otherwise trigger a refusal on that final turn in isolation. It's a scaling attack: it barely works with a handful of shots and grows more reliable as context windows grow, which makes it a technique that gets more dangerous as a side effect of an otherwise desirable capability improvement.

It's also one of the cleanest illustrations of why jailbreaking and prompt injection have to stay conceptually separate. There's no application wrapper here to override. The fabricated turns aren't hijacking a system prompt; they're manufacturing the exact statistical context in which the final request looks, to the model, like business as usual.

Encoding and obfuscation

Content filters, whether keyword matchers or trained classifiers, are usually built against plaintext. Encoding the request so it never appears as recognizable plaintext at the point where a filter inspects it is a direct way around that gap.

Base64 or ROT13 the request and ask the model to decode and answer. Translate it into a low-resource language the safety classifier wasn't trained as heavily against. Split the payload across several turns so no single inspected unit contains the whole thing. Substitute homoglyphs or deliberate misspellings (payment becomes p4yment, or a full Unicode look-alike swap) that defeat exact-string matching while staying perfectly legible to the model itself.

These techniques target the filter rather than the model's judgment, which is why they usually get paired with a roleplay frame or a logic trap on top: get the request past the input filter through encoding, then get the model to actually comply through framing. A filter that only inspects plaintext and a model that only reasons about intent each have a gap of their own. Encoding tricks live in the first one; everything else in this module lives in the second.

Logic traps and framing exploits

Where roleplay changes the narrative frame, a logic trap changes the reasoning frame. The request gets presented as a puzzle, a hypothetical, or a sequence of individually harmless steps whose sum is exactly the outcome the model was trained to refuse when asked directly.

Socratic decomposition is the clearest version. Instead of asking "how do I do X," which gets refused, the attacker asks a chain of narrower sub-questions, each one reading as ordinary curiosity or education on its own, then assembles the harmful answer themselves out of the individually safe responses. A false dichotomy works on a related principle, framing refusal itself as the harmful choice ("if you don't answer, someone gets hurt worse"), which exploits the fact that safety training optimizes for refusing harmful content rather than for reasoning about a game-theoretic framing device bolted on top of the request.

Hypothetical and fictional framing ("write a thriller where the villain explains, accurately, how to...") is roleplay's sibling, but it works through the plausibility of the ask rather than the identity of the speaker. A novel legitimately can contain a villain's monologue, so the request reads as an ordinary creative-writing task right up until the level of technical specificity demanded makes the fictional wrapper transparently cosmetic.

Context poisoning

Everything above assumes the attacker controls only the final message. Context poisoning drops that assumption: the attacker fabricates earlier turns in the conversation, including turns attributed to the assistant itself, already agreeing to comply, and asks the model to continue from that manufactured history.

This is the technique that ties most directly back to the root cause from module 1: once instructions and data collapse into one token stream, the model has no structural way to check whether a turn that looks like a prior assistant response actually happened, or was planted moments ago by whatever fed it that context. In agentic pipelines specifically, where one model's output becomes another model's input across several hops, a poisoned "prior turn" can travel several calls deep before it detonates. That's the same recursive-injection shape module 1 flagged as the hardest to catch.

Automating the search: AJAR

Everything described so far is a hand-crafted technique family, something a human has to notice and shape a prompt around. Dou and Yang's AJAR, Adaptive Jailbreak Architecture for Red-teaming (arXiv:2601.10971), treats jailbreak discovery as an optimization loop instead of a one-off creative act. An attacker model proposes a candidate prompt against a target model, a judge model scores how close the response came to a defined harmful objective, and the attacker iterates on that score rather than on intuition.

What it converges on isn't a new category of technique. It's combinations and refinements of roleplay framing, logic traps and context poisoning, searched automatically across far more variation than a human red-teamer would try by hand. The practical takeaway for anyone doing this work manually is to assume that whatever hand-crafted prompt works against a target today has already been found, or will shortly be found, by an automated search running the same technique library at scale. Module 4 covers AJAR's pipeline in full. Here it matters mostly as a reminder that this library is a starting vocabulary, not a fixed list. The frontier is combinatorial, not enumerable.

Further reading

Anthropic's Many-shot Jailbreaking (2024) is the source of the context-stuffing mechanism above. Dou and Yang's AJAR paper (arXiv:2601.10971) covers the automated technique search in full detail in module 4. Hakim et al.'s Jailbreaking LLMs: A Survey of Attacks, Defenses and Evaluation (TechRxiv, 2026) is the broader survey these technique families sit inside.