Definition · AI security
ASCII smuggling
ASCII smuggling hides text inside invisible Unicode tag characters from the Tags block (U+E0000–U+E007F), a shadow copy of the printable ASCII set that most fonts and interfaces do not render. A human reading the text sees nothing unusual, while a language model that ingests the raw bytes decodes the tag characters as ASCII and reads the hidden message.
Last reviewed
Key points
- ASCII smuggling lives in the Unicode Tags block (U+E0000–U+E007F), a shadow copy of the printable ASCII set; most of those code points are invisible to humans yet present to anything that processes the raw text.
- In prompt injection the attacker hides instructions from the operator while a model that decodes the tags follows them, defeating human-in-the-loop reviews and guardrails.
- MITRE ATLAS files the class as AML.T0068, LLM Prompt Obfuscation: instructions hidden in text, images or metadata to evade humans, guardrails and detectors.
- In 2026 Microsoft documented the same tag characters crossing into high-volume phishing — a single invisible tag-space spliced inside words like 'funding' broke string matches and tokenizers at peaks over two million messages a day.
ASCII smuggling got its name in the AI security research of early 2024, as the clearest gap between what a human sees and what a model reads. A model ingesting the raw bytes of a page decodes the hidden text; the person looking at the page never knew it was there.
How it works
The code points come from the Unicode Tags block, U+E0000 to U+E007F, built for language tagging and now largely deprecated. They mirror the printable ASCII set — U+E0041 stands for “A”, U+E0061 for “a” — so letters, digits and spaces all have shadow versions, but most fonts and interfaces never render them: a string that looks clean to a person still carries a hidden message in the bytes.
A model reads all of it. It ingests the raw text, tokenizes tag characters as letters and follows an instruction the operator never saw. It can also write that way: a reply wrapped in tag characters is invisible on screen, turning the response channel into a second smuggling lane.
Why it matters
In January 2024 a proof of concept hid an instruction in pasted text that made ChatGPT call DALL-E. Weeks later the technique had a name and tooling — the ASCII Smuggler could encode and decode tag payloads — and it reaches beyond chatbots: anywhere a person forwards or approves text they cannot see, hidden instructions ride along and the human-in-the-loop check fails.
The mechanism is MITRE ATLAS technique AML.T0068, LLM Prompt Obfuscation, rated Demonstrated: hiding prompt injections from “humans, large language model guardrails, or other detection mechanisms.” A 2026 benchmark, Reverse CAPTCHA, measured the gap across 8,308 outputs: tool use sharply raised how often invisible instructions were followed, and which encoding a model decoded varied by provider.
In September 2026 the same characters crossed into mass phishing; that crossover and the definitional question it raises are examined in depth below.
What the 2026 crossover changes
Microsoft’s research ran a hunt for hidden prompt-injection content in email and instead surfaced a finance-theme phishing campaign. The mail did not carry smuggled instructions for an AI assistant. It carried a single invisible tag character — U+E0020, TAG SPACE — spliced inside high-signal words: “funding” became fun⟨tag⟩ding. The byte sequence no longer contained the contiguous keyword, so literal matches and tokenizers failed, while the word still read as “funding” to a recipient.
The intent is inverted rather than the mechanism. In 2024 the attacker hid text from a human so the model would read it; in this campaign the attacker hid text from a detector so the recipient would read it normally. Microsoft is explicit that strictly speaking this is “invisible-character insertion using a code point from the ASCII-smuggling tag block, rather than full message smuggling.”
The scale is what is new. Signature hits jumped from about 21,000 messages on 8 February 2026 to more than 1.3 million the next day and peaked over 2.3 million on 11 February 2026, sent from roughly 150 recycled finance-vocabulary domains with a weekday-on, weekend-off cadence. Microsoft’s own layered detection flagged over 99 percent of the messages without relying on the tag characters, and it argues the real defence is “normalize before you match”: strip tag-block characters before keyword and signature logic runs. Their presence is itself a strong anomaly signal — legitimate mail almost never carries them, with the notable exception of England, Scotland and Wales flag emojis, which are encoded from tag characters.
Where definitions disagree
The term has drifted since Rehberger coined it. His original ASCII smuggling means full message smuggling: encoding a whole hidden message into tag characters, and decoding one — the name of the tool is symmetrical. Microsoft uses the term for any invisible tag character from the block, and notes that a single splicing character is not smuggling a message at all. Read the specific campaign before mapping a control to “ASCII smuggling”.
ASCII smuggling also sits inside a family it does not own. Zero-width spaces, zero-width non-joiners, no-break spaces, soft hyphens and homoglyph substitutions have been fracturing words to defeat naive string matchers for years. What changed in 2026 was the character choice — a block that became famous through AI security research — and the discipline and scale of the campaign. Accounts that present invisible-character filter evasion as new are mistaking the character for the idea.
Questions and answers
Is ASCII smuggling the same as prompt injection?
No. ASCII smuggling is a way of hiding text, not an attack of its own. Loaded with an instruction it becomes a delivery mechanism for prompt injection, which is why MITRE ATLAS files it under AML.T0068, LLM Prompt Obfuscation. The same mechanism is reused with other intents: hiding data in a model's reply, and since 2026 hiding phishing keywords from filters. The technique carries no intent; the intent decides what the hidden text does.
Why is the hidden text invisible to humans but readable by AI?
The Tags block code points (U+E0000–U+E007F) mirror the printable ASCII set, but typical fonts and user interfaces do not render them, so they occupy space in the byte stream without showing on screen. A language model processes the raw bytes, tokenizes the tag characters as letters and reads the hidden message, which a person scanning the same text can never see. That asymmetry is the point of the attack.
Does ASCII smuggling need an AI system to work?
No. The mechanism is characters in a text stream. Rehberger noted in 2024 that the technique implies "smuggling of data in plain sight" beyond LLM applications, and Microsoft documented the proof in 2026: a phishing campaign used a single invisible tag-space inside words such as 'funding' to defeat email filters, which do not reason about the hidden characters the way a whitelisted recipient does.
Sources
- ASCII Smuggler Tool: Crafting Invisible Text and Decoding Hidden Codes, Embrace The RedEmbrace The Red (Johann Rehberger), 14 Jan 2024
- Video: ASCII Smuggling and Hidden Prompt Instructions, Embrace The RedEmbrace The Red (Johann Rehberger), 12 Feb 2024
- ASCII smuggling crosses over from AI prompt injection to phishing evasion, Microsoft Security BlogMicrosoft Security Research, 3 Sep 2026
- MITRE ATLAS, AML.T0068 LLM Prompt Obfuscation (collection 2026.08)MITRE
- Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction InjectionarXiv (Marcus Graves), 26 Feb 2026
- Unicode Technical Standard #51, Unicode Emoji, section 1.4.5 (ED-14a)Unicode Consortium, 4 Sep 2025