<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/19563ead3964477694e80be4e0a1f97c&quot; frameborder=&quot;0&quot; width=&quot;1862&quot; height=&quot;1396&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1396</height><width>1862</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1396</thumbnail_height><thumbnail_width>1862</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/19563ead3964477694e80be4e0a1f97c-6b24d3dc13e85d5c.gif</thumbnail_url><duration>152.863</duration><title>LLM Observatory - Demo</title><description>This Loom explains a tool for gaining real time visibility into AI production costs and latency before spending becomes a problem. It shows a main dashboard with total cost, total requests, average latency, and an audit log where every request and response is recorded with tokens, cost, and per request latency as it happens. The presenter sends requests to Haiku and Opus to demonstrate the trade off, noting that Opus increases time to first token and cost compared with the lightweight model. The cost is also broken down by model so teams can see which models drive spend and decide routing without re architecture, migration, or changing how the application is built beyond a single client side configuration change.</description></oembed>