<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/23f92f47f790497b8e026b49f86dcaef&quot; frameborder=&quot;0&quot; width=&quot;1150&quot; height=&quot;862&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>862</height><width>1150</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>862</thumbnail_height><thumbnail_width>1150</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/23f92f47f790497b8e026b49f86dcaef-3da9bb8481d4c50b.gif</thumbnail_url><duration>205.236</duration><title>Mintoken Cuts AI Token Costs with Caching</title><description>This Loom explains how Mintoken reduces AI token cost by using a self-evolving cache and routing to the cheapest capable model. It checks cached responses for a proper answer first, and if none exists it queries the lowest-cost model such as open source or cloud options like ChatGPT or Gemini, then caches the result. The presenter claims users see up to 91% cost savings and that the system improves as more people use it. It also describes how tools like Pioneer, Senso, Guild, Ban, and Replay are used to route questions, enforce a 90% quality floor, promote only policies that preserve quality while saving cost, announce promotions, and validate reused answers.</description></oembed>