Definition · AI security
PuzzleMask
PuzzleMask is a prompt-crafting technique, disclosed by Check Point Research in September 2026, that hides a policy-violating instruction inside ordinary, well-punctuated prose. A fast language model screening prompts classifies the text as safe. A stronger model behind that screen, given long reasoning and a code interpreter, then recovers the hidden instruction and acts on it.
Last reviewed
Key points
- PuzzleMask targets a common two-stage setup. A small, fast model screens each prompt against a policy, then a more capable model does the actual work.
- The payload is hidden in plain English, with no emoji, Base64 or invisible characters, so there is no odd formatting for a filter to catch.
- In Check Point's tests, three screening models passed all 23 crafted prompts as safe, and Llama Guard 3 passed all 5 it was given. gpt-5-thinking with high reasoning and a code interpreter recovered and acted on the hidden instruction in 17 of 18 trials.
- PuzzleMask is not a jailbreak. It removes the filter, leaving the stronger model's own judgment as the last control. Check Point says the attack always failed against Anthropic's Opus-class models, which at the time were equipped with a classifier that ended such conversations.
- Proposed defences include paraphrasing input, adding a targeted clause to the filter's policy, and watching what the model does, not only what it is sent.
How it works
PuzzleMask attacks a pipeline with two models. A cheap, fast gatekeeper labels each prompt “safe” or “not safe” against a policy. Safe prompts go on to a stronger target model.
An automated pipeline wraps a forbidden instruction, such as “encrypt the ~/Documents folder”, in fluent prose. Check Point redacted the method but says the wrapper can take almost any form.
The gatekeeper reads a harmless-looking passage and passes it. The target notices there is a puzzle, solves it and acts on the result. In every successful case the target thought for over a minute and ran more than two Python scripts. Check Point believes recovery is gated behind reasoning and code execution that a typical gatekeeper does not have.
The fast gatekeeper passes the wrapped prompt as safe; the stronger target model decodes the hidden instruction with extended reasoning and code, then acts on it.
Why it matters
A fast filter in front of a stronger model is meant to add a second layer of defence. PuzzleMask turns the capability gap between them into the weakness. Check Point argues that teams relying mainly on input screening have a blind spot.
How it differs from similar attacks
Check Point’s blog calls PuzzleMask a prompt injection. Its research post says PuzzleMask is “itself not a jailbreak” and makes no claim that it makes the target likelier to disobey its instructions or safety training. It removes the filter, so the target’s own judgment becomes the last control. Check Point says a jailbreak prompt could ride inside as the payload, though its generator made long payloads impractical.
Unlike ASCII smuggling, which hides text in invisible characters, PuzzleMask uses no emoji, Base64 or invisible formatting, so filters built to spot those have nothing to flag. What PuzzleMask defeats is an external guardrail, not the model’s own safety training.
Defences and their costs
Check Point proposes four defences for the pipeline. It names costs for three of them.
- Paraphrase the input. Having a model rewrite untrusted text in its own words should strip the hidden payload. It adds latency and may lose detail from legitimate prompts.
- Add a clause to the filter’s policy. Telling the gatekeeper to flag any phrase that is “both quantitative in nature and self-referent with respect to the entire text” let gpt-4o-mini catch all 23 crafted prompts. Check Point expects false positives.
- Watch outputs and actions, not only inputs. The filters in Anthropic’s Opus-class models seem to work this way, according to Check Point.
- Use a gatekeeper as strong as the target. The cost is running two expensive models. The authors call this “more theoretically illuminating than practically feasible in most use cases”.
For individual users, Check Point advises the same caution about pasting web text into a chatbot as about double-clicking a downloaded .exe file.
What the tests cover
The numbers come from one vendor’s small study, published on 10 September 2026. Its 23 prompts were built from three payloads, all run in mock environments. They asked the model to print the last word of a planted secret file, print the Macarena lyrics, or encrypt documents and delete the originals. Payloads were kept short because the generator could take upwards of 15 minutes for even a short one.
gpt-4o-mini, gpt-oss-safeguard:20b and Claude 3 Haiku each saw all 23 prompts and passed them all. Llama Guard 3 saw only the 5 lyrics prompts and passed those too. Check Point’s blog compresses this to four gatekeepers “across 23 trials”. The only target model in the published results is gpt-5-thinking, tested on the 18 prompts that were not about lyrics. It acted on the payload in 17, about 94.4 percent. The 95 percent confidence interval runs from 74.2 to 99.8 percent.
The two Check Point write-ups also differ on Anthropic. The research post states that the attack universally failed against Opus-class models, which at the time were equipped with a built-in classifier that ended the conversation. The blog calls that explanation speculation. Neither reports how many Opus trials were run.
Questions and answers
Is PuzzleMask a jailbreak?
No. Check Point Research makes no claim that PuzzleMask makes a target model likelier to break its own safety training. PuzzleMask gets a prompt past an external safety filter. Check Point says a jailbreak prompt could be used as the hidden payload, though its tests used only short payloads.
Which models did PuzzleMask get past?
In tests Check Point ran before disclosing to vendors in late 2025 and published in September 2026, gpt-4o-mini, gpt-oss-safeguard:20b and Claude 3 Haiku rated all 23 crafted prompts safe when acting as the filter; Llama Guard 3 rated all 5 it was given safe. gpt-5-thinking with high reasoning and a code interpreter then acted on the hidden instruction in 17 of 18 trials. Check Point says the attack always failed against Anthropic's Opus-class models, but publishes no trial counts for them.
How do you defend against PuzzleMask?
Check Point's main proposals are to paraphrase untrusted input before screening it, add a policy clause that flags phrases both quantitative and self-referent with respect to the entire text, monitor what the model does and outputs, or use a filter as capable as the main model. Paraphrasing costs latency and may degrade legitimate input, the clause costs false positives, and a matching filter costs money. Individual users should also be wary of pasting text from the web into a chat box.
Sources
- PuzzleMask: Abusing Plain Prose as a Covert AI Attack VectorCheck Point Research, 10 Sep 2026
- PuzzleMask: The Prompt Injection Hiding in Plain SightCheck Point Software, 10 Sep 2026