Skip to main content
Prompt Security from SentinelOne
  • AI Security Academy

    AI Security Academy

    Learn More
    Title of post

    What is AI Security

    AI security is not a neat, one-line definition you can slap on a slide.

    Learn More
    Title of post

    AI Security Glossary

    Explore some of the most common terms in AI Security

    Learn More
    Title of post

    AI Usage Stats

    Explore current AI usage trends.

  • Tools

    AI Security Tools

    Learn More
    Title of post

    OneClaw

    Track and analyze OpenClaw deployments in your org

    Learn More
    Title of post

    ClawSec

    Secure your OpenClaw, NanoClaw, and Hermes agents.

    Learn More
    Title of post

    Prompt Fuzzer

    Get our AI vulnerability assessment open source tool

  • Blog
  • Startup Map
  • Learn More
    Book a Demo
Prompt Security from SentinelOne
  • AI Security Academy

    AI Security Academy

    Learn More
    Title of post

    What is AI Security

    AI security is not a neat, one-line definition you can slap on a slide.

    Learn More
    Title of post

    AI Security Glossary

    Explore some of the most common terms in AI Security

    Learn More
    Title of post

    AI Usage Stats

    Explore current AI usage trends.

  • Tools

    AI Security Tools

    Learn More
    Title of post

    OneClaw

    Track and analyze OpenClaw deployments in your org

    Learn More
    Title of post

    ClawSec

    Secure your OpenClaw, NanoClaw, and Hermes agents.

    Learn More
    Title of post

    Prompt Fuzzer

    Get our AI vulnerability assessment open source tool

  • Blog
  • Startup Map
  • Learn More
    Book a Demo
Skip to main Content
Back to Blog
Back to Blog
AI Resources
AI Code Assistants

Unicode Exploits Are Compromising Application Security

Written by: 
Omer Zilberman, Data Scientist
April 30, 2025

A new attack surface for LLMs is gaining significant attention in IT circles, and it’s more menacing than it looks at first glance: the innocent smiley face 🙂. To be precise, this is a story about any emoji – or for that matter, any Unicode character – being used to conceal harmful or extensive text. 

By encoding arbitrary-length data into single Unicodes via zero-width joiner sequences (ZWJ), attackers make the embedded content invisible to human reviewers, increasing the likelihood of bypassing organizations’ content filtration mechanisms. The fact that this vulnerability is relevant for both visible Unicodes and invisible Unicodes makes it especially disconcerting. 

Earlier this year, we explored Unicode abuse against coding assistants, focusing on how invisible characters were being introduced into open-source code repositories to sabotage assistants’ outputs. Because invisible characters like whitespace characters (e.g., spaces, tabs, and newlines), control characters (e.g., carriage return CR and line feed LF), and zero-width characters (e.g., zero-width space U+200B and zero-width non-joiner U+200C) do not appear on text editor display screens, they constitute an understandable attack surface for Unicode abuse.

But it turns out that visible symbols, such as numbers, letters, and emojis, are also vulnerable to Unicode abuse, and as with invisible Unicode characters, visible Unicode characters can lead to attacks.

The most well-known type of attack that Unicode abuse can lead to is Prompt Injection, where an attacker manipulates an LLM through carefully crafted inputs. This includes jailbreaking, a specific category of prompt injection where the goal is to coerce a GenAI application into deviating from its intended behavior and predetermined guidelines. Attackers can embed a hidden prompt injection or jailbreak command within an emoji, which they go on to feed to an LLM.

Unicode abuse can also lead to token expansion attacks.

In language modeling, ‘tokens’ are the smallest unit of text that an LLM can comprehend and process, often as short as a single visible or invisible character. In a token expansion attack, an attacker exploits the process by which certain tokens, once processed by the model, expand into a larger number of tokens. By embedding excessive or malicious content within a single Unicode character, the attacker can cause that single token to expand into large numbers of tokens.

Token expansion can result in unbounded consumption, where an LLM is manipulated to process excessive amounts of information. When unbounded consumption occurs, the model’s computational capacity – already stressed by high computational demand – can become overwhelmed, leading to severe performance issues. This in turn increases the model’s vulnerability to Denial of Service (DoS) attacks, model theft, and other risks.

Both Prompt Injection and unbounded consumption are considered by the Open Worldwide Application Security Project (OWASP) to be among the top ten most critical vulnerabilities found in applications that use LLMs.

‍

Prompt Security employs protection measures that inspect text down to Unicode level in real time. These measures are capable of restricting, blocking, and redacting visible and invisible characters, and are configurable so that organizations can customize protection according to their specific requirements.

To learn more about how Prompt Security can help protect your organization from Unicode abuse and other threats, get in touch today.

‍

Last Updated:
August 6, 2026

Share this post
Summarize this Post
On This Page

TOC Element

Related Posts

View All Posts
View All Posts
Learn More
Title of post

ChatGPT Security Guide: Enterprise Risks, Incidents, and Practitioner Guidance

AI Risks

AI Resources

Jun 29th, 2026

What security teams actually need to know about ChatGPT: data handling, documented incidents, CISO guidance, and API risks.

Learn More
Title of post
Learn More
Title of post

The Key Layer in AI Security: Browser and Endpoint Sensors

AI Resources

AI Risks

Feb 19th, 2026

SASE, proxies, and EDR/MDM are foundational, but AI needs more. Learn why browser and endpoint sensors enable real-time AI governance.

Learn More
Title of post
Learn More
Title of post

When Your Plugin Starts Picking Your Dependencies: Marketplace Skills and Dependency Hijack in Claude Code

AI Risks

AI Resources

Jan 5th, 2026

Claude Code marketplace skills can rewrite how dependencies are installed. Demo shows silent httpx hijack and OWASP agentic failures.

Learn More
Title of post
Prompt Security from SentinelOne
Log In
Log In
Learn More
Book a Demo

Resources

Blog
AI Security Glossary
What is AI Security?
PromptCast: The Voice of AI & Security
ClawSec
OneClaw
Prompt Fuzzer
AI Security Startup Map
© {{year}} Prompt Security. All Rights Reserved.
Privacy Policy
Terms of Service

Follow Us