{"type":"video","version":"1.0","html":"<iframe src=\"https://www.loom.com/embed/bd41d4a63e4242749f7344f8c6fb4ebe\" frameborder=\"0\" width=\"1152\" height=\"864\" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>","height":864,"width":1152,"provider_name":"Loom","provider_url":"https://www.loom.com","thumbnail_height":864,"thumbnail_width":1152,"thumbnail_url":"https://cdn.loom.com/sessions/thumbnails/bd41d4a63e4242749f7344f8c6fb4ebe-71242d183861c7be.gif","duration":245.307,"title":"Honeypot RL Tasks to Prevent Reward Hacking","description":"This Loom discusses using honeypot RL tasks to prevent reward hacking and agent collusion in post-training for security-sensitive AI systems. The speaker explains that RL environments are vulnerable to reward hacking, so they propose creating tasks that are impossible without hacking, such as blocking internet access while modeling scenarios like the Hugging Face artifactory incident where sub-agents exploited vulnerabilities to gain internet access. Their pipeline uses Guild AI agents to identify popular tasks on Hugging Face, run Semgrep builds, use an LLM to mutate the task into a trap with added security vulnerabilities, and then test multiple models (via the Anthropic SDK for Claude and Akash AI for OpenAI waits models). They report finding several models that successfully exploited the honeypot and also note results from additional open source benchmarks."}