{"type":"video","version":"1.0","html":"<iframe src=\"https://www.loom.com/embed/d90827338c5d4e4e9c48b1dcf0ba3f93\" frameborder=\"0\" width=\"1658\" height=\"1243\" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>","height":1243,"width":1658,"provider_name":"Loom","provider_url":"https://www.loom.com","thumbnail_height":1243,"thumbnail_width":1658,"thumbnail_url":"https://cdn.loom.com/sessions/thumbnails/d90827338c5d4e4e9c48b1dcf0ba3f93-16a65dd383003a0d.gif","duration":113.408,"title":"How E-Agents Get Crash Tested Reliably","description":"This Loom explains a “crash test” method for evaluating e-agents by generating adversarial test cases. The process has four steps: it finds real documented agent failures from production, then Gemini writes 12 to 15 attacks tailored to the agent’s system prompt, and separate models run and grade each test so none grades its own work. Results are scored and failures are sorted by severity with exact inputs and how the agent responded. The speaker highlights that this catches issues like de-inventing policies and a support agent committing to a $500 refund it could not actually process."}