<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/79151d2fc76b4eb38802ae6a457f5680&quot; frameborder=&quot;0&quot; width=&quot;1662&quot; height=&quot;1246&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1246</height><width>1662</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1246</thumbnail_height><thumbnail_width>1662</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/79151d2fc76b4eb38802ae6a457f5680-7612db520be92e4f.gif</thumbnail_url><duration>177.174</duration><title>Exploring Impossible Moments: A Benchmark for Future Models 🚀</title><description>In this video, I introduce my project, Impossible Moments, which serves as a benchmark for future models, focusing on reasoning capabilities across various disciplines such as physics and philosophy. I developed this benchmark using five different agents during a single Cloud Code session, resulting in 12 scenario categories and five top tiers: Spark, Fracture, Rapture, Singular, and Impossible. The Impossible tier presents scenarios that are currently unsolvable, like making a submarine invisible. I encourage researchers to explore these open-source prompts and frameworks to enhance their own benchmarks. My goal is to foster collaboration and innovation in the development of large language models.</description></oembed>