<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/168a5cac9ed64678ae1badf359bf57dc&quot; frameborder=&quot;0&quot; width=&quot;1816&quot; height=&quot;1362&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1362</height><width>1816</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1362</thumbnail_height><thumbnail_width>1816</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/168a5cac9ed64678ae1badf359bf57dc-9b409362065311c5.gif</thumbnail_url><duration>109.7077</duration><title>Evaluating AI Agents with Realistic Simulated Interactions</title><description>In this video, I discuss how Scorecard is addressing the challenge of evaluating AI agents by creating realistic simulation users for interaction. We leverage pre-existing data to train new personas, using the chat2BT5 response API to enhance our multi-turn interactions. I demonstrate this with a specific example involving Alaska Airlines, showcasing how our system can now provide detailed responses and visual feedback on agent performance. I encourage viewers to explore the Scorecard platform to see the results of these simulations and understand how our agents are performing. Your feedback on this process will be invaluable as we continue to refine our approach.</description></oembed>