{"type":"video","version":"1.0","html":"<iframe src=\"https://www.loom.com/embed/27b8d94c7e024688a9c50b383f0eed7f\" frameborder=\"0\" width=\"1280\" height=\"960\" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>","height":960,"width":1280,"provider_name":"Loom","provider_url":"https://www.loom.com","thumbnail_height":960,"thumbnail_width":1280,"thumbnail_url":"https://cdn.loom.com/sessions/thumbnails/27b8d94c7e024688a9c50b383f0eed7f-07383aa0337dfb64.gif","duration":140.1,"title":"When your Voice AI grader is wrong, not your agent","description":"This Loom explains how voice AI graders can produce false positives that make a working agent seem to fail. The speaker describes an integration test where the agent correctly handled a caller refusing her email and later switched to an emergency flow after a burning smell from the electrical panel, but the grader still marked the call as failed. The issue was a grader rule not present in the prompt, creating a grader false positive. The recommended debugging order is to verify whether the prompt rule was actually broken, then check for possible transcript misreading, and only then treat the failure as real if the agent truly violated an existing prompt rule."}