<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/b8d14f283a494682b6917a64475e0dbd&quot; frameborder=&quot;0&quot; width=&quot;1756&quot; height=&quot;1317&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1317</height><width>1756</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1317</thumbnail_height><thumbnail_width>1756</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/b8d14f283a494682b6917a64475e0dbd-7e94d93c36d1397e.gif</thumbnail_url><duration>297.152</duration><title>VelloxCon KV Caching, 43 Methods, 98% Reduction</title><description>This Loom introduces VelloxCon, a unified API for combining KV caching compression methods to reduce memory use on MacBook while enabling longer conversations. The creator built it over four to five months to address out-of-memory errors when downloading open source models, combining around 43 compression algorithms, with peak memory reduction of about 98% and up to 16x improvement versus a normal fp16 baseline. They also mention custom Metal kernels that provide about 13x speedup for a single conversation, plus multi-model support and classifier categories like quantizers, token eviction, and cross-layer merging. The Loom demonstrates experimenting in a playground to compare compression settings and reports that, for TurboCont, the context fitting increases from about 61,4400 words to around 460,000 watts.</description></oembed>