{"type":"video","version":"1.0","html":"<iframe src=\"https://www.loom.com/embed/b8d14f283a494682b6917a64475e0dbd\" frameborder=\"0\" width=\"1756\" height=\"1317\" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>","height":1317,"width":1756,"provider_name":"Loom","provider_url":"https://www.loom.com","thumbnail_height":1317,"thumbnail_width":1756,"thumbnail_url":"https://cdn.loom.com/sessions/thumbnails/b8d14f283a494682b6917a64475e0dbd-7e94d93c36d1397e.gif","duration":297.152,"title":"VelloxCon KV Caching, 43 Methods, 98% Reduction","description":"This Loom introduces VelloxCon, a unified API for combining KV caching compression methods to reduce memory use on MacBook while enabling longer conversations. The creator built it over four to five months to address out-of-memory errors when downloading open source models, combining around 43 compression algorithms, with peak memory reduction of about 98% and up to 16x improvement versus a normal fp16 baseline. They also mention custom Metal kernels that provide about 13x speedup for a single conversation, plus multi-model support and classifier categories like quantizers, token eviction, and cross-layer merging. The Loom demonstrates experimenting in a playground to compare compression settings and reports that, for TurboCont, the context fitting increases from about 61,4400 words to around 460,000 watts."}