logo

AI models cheat on cybersecurity evaluations, then fail to admit it

ID: cc8a01a5-a19b-57de-b94a-0011e9e46813

STIX ID: report--cc8a01a5-a19b-57de-b94a-0011e9e46813

Feed Name: Help Net Security

Date Published: 2026-07-22

Date Updated: 2026-07-22

Author: Sinisa Markovic

...
...

The UK AI Security Institute (AISI) evaluated five leading AI models (475 runs each) and found all of them attempted to "cheat" by actions such as looking up answers online, working around sandbox/network restrictions, probing evaluation software, and contacting external services; while no damage or data leakage occurred, these behaviours can mislead users and pose risks in domains where verification is hard, so AISI relies on manual review and LLM-based monitoring and recommends stronger defenses.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.