<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/6a1f47f8ff0b40568744e3bd66685143&quot; frameborder=&quot;0&quot; width=&quot;1920&quot; height=&quot;1440&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1440</height><width>1920</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1440</thumbnail_height><thumbnail_width>1920</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/6a1f47f8ff0b40568744e3bd66685143-1886ca066e38b4d1.gif</thumbnail_url><duration>286.555</duration><title>Add Statistical Significance to Eval Deltas</title><description>This Loom presents a statistical significance layer for the BrainTrust platform to determine whether prompt changes that raise eval scores reflect real improvements or noise. The speaker shows that dashboards may display small gains like 62% to 64%, which can be misleading, especially in paired factual Q&amp;A tests where a “clear win” improved by about 16% across 20 cases. They then demonstrate a deceptive case where scores rose by roughly 2%, but a paired bootstrap significance check yields a confidence interval indicating the result is not significant. The script pulls per-question scores via the BrainTrust SDK and computes confidence bounds from paired bootstrapping, concluding that raw score deltas can mislead shipping decisions without this gate.</description></oembed>