Roblox chat moderation gets bypassed by leet speak and code words
ID: 4d75925e-82ff-573a-9dc8-7e72f862b7fe
STIX ID: report--4d75925e-82ff-573a-9dc8-7e72f862b7fe
Feed Name: Help Net Security
An independent audit examined ~2 million Roblox chat messages from public servers and found the platform's automated moderation frequently fails to block grooming of minors, sexual solicitation, violent threats, self-harm references, and harassment. Researchers documented common bypass methods (splitting phrases, spelling variants, code words, leetspeak) and recommend combining pattern-based detection with language models, evaluating entire conversations rather than single messages, tracking repeat offenders across games/sessions, and improving user feedback; private servers were not included and reported figures are a lower bound.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
