logo

Hackers Can Hide Malicious AI Commands Inside Normal English to Bypass Security Filters

ID: 1fdf4114-e3fb-5c63-999e-de509a18d612

STIX ID: report--1fdf4114-e3fb-5c63-999e-de509a18d612

Feed Name: cybersecurityNews.com

Threat Score
50/100

Date Published: 2026-09-11

Date Updated: 2026-09-11

Author: Tushar Subhra Dutta

...
...

Check Point researchers disclosed “PuzzleMask,” a prompt-obfuscation technique that conceals harmful instructions inside normal English so lightweight gatekeepers mark inputs as safe while more capable downstream models recover and act on the hidden payload; tests showed the downstream model recovered and acted on concealed instructions in 17 of 18 cases (94.4%). The report emphasizes separating content from commands, reducing agent permissions, paraphrasing untrusted input, improving gatekeeper rules, and monitoring model outputs and tool calls; it includes one test IoC (file name `flag.txt`).

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.