Bypassing Amazon Bedrock Guardrails Content Filters
This article documents three prompt-obfuscation techniques that disguised harmful intent as harmless string-processing tasks and bypassed Amazon Bedrock Guardrails content filters configured with High filter strength on the Standard tier. To the best of my knowledge, this is the second published report of an Amazon Bedrock Guardrails bypass, following earlier work by NR Labs.
Remediation status: AWS has addressed all three bypasses described below. They are no longer reproducible; this article preserves the original research and disclosure timeline for reference.