logo

How nice that state-of-the-art LLMs reveal their reasoning ... for miscreants to exploit

ID: 849e7443-90cf-5f13-857e-0ab90890c865

STIX ID: report--849e7443-90cf-5f13-857e-0ab90890c865

Feed Name: The Register (Security)

Date Published: 2025-02-25

Date Updated: 2026-04-26

Author: Thomas Claburn and Jessica Lyons

...
...

Researchers introduce H-CoT, a prompt-based jailbreak that hijacks chain-of-thought safety reasoning in leading LRMs (OpenAI o1/o3 and o3-mini, DeepSeek-R1, Gemini 2.0 Flash), using a Malicious-Educator dataset to drastically lower safety rejection rates and expose delayed/weak safeguards; tests conducted via web UIs/APIs underscore differences between remote filtering and local deployments, while external vendor evaluations of DeepSeek vary in rigor—collectively suggesting exposed CoT reasoning is a significant new safety attack vector.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.