<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/f7e76d4ef72948e3b3e9818a267f438c&quot; frameborder=&quot;0&quot; width=&quot;1670&quot; height=&quot;1252&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1252</height><width>1670</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1252</thumbnail_height><thumbnail_width>1670</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/f7e76d4ef72948e3b3e9818a267f438c-552c8c311dc5649b.gif</thumbnail_url><duration>524.587</duration><title>Loyalty Lens: Detecting Secretly Loyal LLMs</title><description>This Loom presents the Loyalty Lens project for detecting secret loyalty behaviors in 1.5 billion parameter language models. The author trained a zoo of 29 organisms, including 16 secretly loyal variants and 13 clean controls, finding SFT worked better than DPO and GRPO, with DPO failing to handle close calls. They probe whether prompt-installed principles generalize to wait-installed ones and report a maximum AUC of 0.712, indicating prompt probes are only partial proxies. Using Jacobian Lens and Logit Lens, they detect loyalty verbally and find the signature most consistently in layers 23 to 26, with JLens about 6 percent better than Logit Lens and code plus data at github.com/document1998/loyalty-lens.</description></oembed>