<?xml version="1.0" encoding="UTF-8"?><oembed><type>video</type><version>1.0</version><html>&lt;iframe src=&quot;https://www.loom.com/embed/1ee2441f7bc742508c747be567f4ddd1&quot; frameborder=&quot;0&quot; width=&quot;1920&quot; height=&quot;1440&quot; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;</html><height>1440</height><width>1920</width><provider_name>Loom</provider_name><provider_url>https://www.loom.com</provider_url><thumbnail_height>1440</thumbnail_height><thumbnail_width>1920</thumbnail_width><thumbnail_url>https://cdn.loom.com/sessions/thumbnails/1ee2441f7bc742508c747be567f4ddd1-0fc6e4a9395cf6bb.gif</thumbnail_url><duration>355.947</duration><title>Benchmarking Speech to Text Models for Interviews</title><description>This Loom benchmarks multiple automatic speech recognition models for an interview speech to text system, focusing on accuracy, word error rate, entity recall, and latency. The presenter tests 32 audio samples from different speakers using phone mics in noisy conditions, evaluating ground truth entities like place names (for example Electronic City, Jaya Nagar) and multi-entity sentences. Sarvam performs best overall, with particularly strong results on audio clips under 30 seconds, while audio above 30 seconds is not processed properly for Sarvam. The Loom also explains using RealTimeFactor to compare processing speed and notes that larger models like Google CSSD and OpenAI Whisper can produce semantically correct but differently mapped outputs that score worse in WER.</description></oembed>