<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/8963dede3f7248b1b3f94a4e487ee431&quot; frameborder=&quot;0&quot; width=&quot;1280&quot; height=&quot;960&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>960</height><width>1280</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>960</thumbnail_height><thumbnail_width>1280</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/8963dede3f7248b1b3f94a4e487ee431-687c3e8ad88eb2ac.gif</thumbnail_url><duration>136.149</duration><title>Fully Encrypted AI Inference with FHE</title><description>This Loom discusses a fully encrypted inference platform that keeps users data encrypted throughout the entire transformer lifecycle using fully homomorphic encryption. It explains that the system works with any open model without additional training or fine tuning, by finding a proxy for nonlinearity in deep neural network forward passes, including a custom Softmax using polynomials. The approach encodes inputs and outputs so the operator cannot see the underlying values. The speaker also claims over 40% quicker inference time than leading academic results while remaining fully encrypted, and highlights that prompts produce faster generation on their solution versus prior results.</description></oembed>