<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/77e2d77299e3472d8f85febbad78484b&quot; frameborder=&quot;0&quot; width=&quot;1900&quot; height=&quot;1425&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1425</height><width>1900</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1425</thumbnail_height><thumbnail_width>1900</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/77e2d77299e3472d8f85febbad78484b-ac3e67d37d5da760.gif</thumbnail_url><duration>68.864</duration><title>How Zork Learning Harnesses Replay Memories</title><description>This Loom explains how a reinforcement-style harness helps a Zork-playing agent learn from past runs instead of repeating the same mistakes. It describes Zorkinator rebuilding a small prompt each move from its working knowledge, map, items, and prior attempts in the current room, then predicting outcomes and logging every move as evidence. After each game, fixed code carries the map and action outcomes forward, and a reflector turns deaths and surprises into new memories that cite the proving moves. Those updates are published into MongoDB Atlas as a new fixed version, and each subsequent game loads exactly that version, allowing rules to block commands only after replaying past games.</description></oembed>