{"schemaVersion":"1.0","type":"Article","types":["Article"],"slug":"whistle-speech-to-text-in-16-9-mb-subc8","url":"https://zyvop.com/whistle-speech-to-text-in-16-9-mb-subc8","title":"Whistle: Speech to Text in 16.9 MB","subtitle":"A 16.9 MB speech-to-text model that runs on the CPU, handles seven languages and reaches its first token in about 11 ms, built for wearables, phones and tiny devices.","tldr":"Cactus Compute’s Whistle packs on-device speech recognition into one 16.9 MB file: no GPU, no dependencies, seven languages, word timestamps and keyword biasing. Here’s how it works, how it compares to Whisper base, and where it falls short.","keywords":["Cactus Compute","speech recognition","Whisper","on-device AI","edge computing","News"],"entities":["Anshu Pathak","Cactus Compute","speech recognition","Whisper","on-device AI","edge computing","News","ZyVOP"],"keyTakeaways":["A speech recognition model small enough to live inside a watch, and quick enough that you stop noticing it's there.","Speech-to-text has a size problem.","The popular open models are happy on a laptop or a phone with room to spare, but the moment you move to a budget handset, a wearable, a smart speaker or a microcontroller-class board, a 145 MB model stops being a footnote and becomes the whole conversation."],"headings":["What Whistle actually does","A look under the hood","How it compares","The honest fine print","From speech straight to action","Try it in a few lines","Where it fits","The takeaway"],"outboundLinks":[],"contentText":"A speech recognition model small enough to live inside a watch, and quick enough that you stop noticing it's there. Speech-to-text has a size problem. The popular open models are happy on a laptop or a phone with room to spare, but the moment you move to a budget handset, a wearable, a smart speaker or a microcontroller-class board, a 145 MB model stops being a footnote and becomes the whole conversation. On October 2, 2026, Cactus Compute released Whistle, a speech recognition model that attacks that problem head-on: the entire thing is one 16.9 MB file. It runs on the CPU, needs no dependencies and no GPU, and loads into the same C++ engine as the company's Needle model. What Whistle actually does Whistle takes 16 kHz mono audio, up to 30 seconds per pass, and handles three jobs entirely on the device: Transcription in English, German, French, Spanish, Italian, Dutch and Polish, with the language auto-detected unless you name it. Word timestamps, with each word's start, end and confidence, aligned using the decoder's own attention. That's enough to highlight words as they're spoken, seek to a phrase, or cut audio on a word boundary. Speech embeddings: the encoder's output, one row per 80 ms frame, which you can use for matching and retrieval without ever producing a transcript. There is also keyword biasing. Hand Whistle a list of names or terms and it nudges the beam search toward them. This is the fix for the classic failure where a recognizer has never heard \"Siobhan\" or \"Krzysztof\" and quietly substitutes something ordinary. A look under the hood Whistle isn't a shrunk-down Whisper. It's a purpose-built design with roughly 55 million parameters (about 36 million active), quantised to 2 bits to reach that 16.9 MB footprint. The pipeline goes like this: Front end. The audio becomes 80-bin log-mel features, which a small convolutional stem shortens to 375 frames for a 30-second clip, one frame per 80 ms. Encoder. Eight Simple Attention blocks, the same family of blocks Needle uses, attend over the whole clip at once. Decoder. Eight more blocks generate the transcript, with one speech-specific addition: a gated cross-attention in every layer that reads the encoder output. Search. A five-beam search picks the transcript, with keyword biasing applied along the way. Two details are worth noting. The decoder is \"laddered\", meaning every depth from two layers upward was trained as a deployable model in its own right, so you can pick a shallower decoder at load time to trade accuracy for speed on weaker hardware. And the engine checks loudness before decoding, so a silent clip returns an empty transcript without ever entering the search. How it compares The launch benchmarks pit Whistle against Whisper base and Moonshine tiny v2, two of the obvious small-model references. Whistle Whisper base Moonshine tiny v2 File size 16.9 MB 145.3 MB 41.9 MB Time to first token (10 s clip) 11.1 ms 73.2 ms 22.8 ms Decode speed 1,319 tokens/s 266 tokens/s 262 tokens/s That works out to Whistle being roughly 8.6x smaller than Whisper base and about 6.6x quicker to the first token. These timings come from a 10-second clip on an Apple M4 Pro CPU, with each model on its official runtime. On accuracy, measured as word error rate (lower is better), Whistle edges out Whisper base on the sets both report: Benchmark Whistle Whisper base LibriSpeech test-clean 4.31 4.9 LibriSpeech test-other 10.49 11.0 FLEURS (average) 21.4 24.5 Whistle also posts 7.65 on SPGISpeech and 19.01 on Earnings-22, two benchmarks Whisper's authors never published figures for. Its latency also scales with clip length rather than staying flat: about 5.9 ms to first token at 5 seconds, 11.1 ms at 10 seconds and 36.3 ms at 30 seconds, because it doesn't pad every input out to a full 30 seconds the way Whisper does. The honest fine print A small model makes trade-offs, and the launch results say so plainly. Whisper base still wins in places. It's ahead on TED-LIUM, on AMI and on the MLS average. Whistle looks strongest on read speech, business speech and the seven languages it targets. For long-form lecture audio, the bigger decoder in Whisper base buys real accuracy. Seven languages only. If you need broad multilingual coverage, this isn't that model. 30 seconds per call. Longer audio has to be split, or handled through the streaming API, which commits words once two consecutive passes agree so the text you've already received never changes. Benchmarks come from a high-end chip. The speed figures were measured on an M4 Pro, not on the budget phones and wearables the model is aimed at, so test on your own target hardware. Competitor numbers are published figures. Whisper's and Moonshine's error rates are what their authors reported, while Whistle's were measured over 86,174 utterances, so it's not a perfectly like-for-like rerun. It's new. Treat it as an early-stage model and validate it on your own audio before betting a product on it. From speech straight to action The more interesting idea is what happens when Whistle sits next to Needle. Because both load into the same engine from the same container format, one binary can take an audio clip, transcribe it, run the transcript against your tool definitions, and hand back a single JSON object containing both the function calls and the speech fields. Say \"turn off the kitchen lights\" and you get back a set_lights call with the right arguments, plus the transcribed text and detected language. The caller never has to touch a transcript. For voice-controlled gadgets, that collapses what is usually a two-model, two-runtime pipeline into one small, offline package. Try it in a few lines Install the package: pip install cactus-needleThen transcribe a clip: import needle result = needle.transcribe(\"clip.wav\") print(result[\"text\"], result[\"language\"])Every call also reports time to first token and decode speed. A few options are worth knowing: # Word-level timing needle.transcribe(\"clip.wav\", word_timestamps=True) # Bias toward names your users actually say needle.transcribe(\"clip.wav\", keywords=[\"Aoife\", \"Wojciech\"]) # Skip detection and force a language needle.transcribe(\"clip.wav\", language=\"de\")If you'd rather poke at it first, the project ships a command-line playground that transcribes from your microphone, plus a compare mode that runs the same clip through Whistle, Whisper and Moonshine side by side. The model also has an in-browser demo that runs locally, so your audio never leaves the device. Where it fits The engine ships prebuilt for seventeen targets, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. That makes Whistle a natural fit for: Wearables and smart-home devices where storage and battery are tight. Privacy-sensitive apps where audio shouldn't leave the device. Offline-first products that have to work without a connection. Servers on a budget. The community has already wrapped Whistle in an OpenAI-compatible transcription API that idles around 80 MB of RAM, small enough for free-tier containers that a multi-gigabyte Whisper image would never fit. The takeaway Whistle isn't trying to be the most accurate speech model in the world, and its creators say as much. The bet is different: compress capable speech recognition until it fits places that have been left out of the on-device AI conversation. At 16.9 MB, with competitive accuracy on its target languages and a first token in about 11 milliseconds, it makes that bet credibly. If you build for small devices, it's worth ten minutes with the playground. Just remember the limits: seven languages, thirty seconds a pass, and benchmarks that deserve to be rechecked on your own hardware and your own audio. Source material: the Cactus Compute launch post, the model card on Hugging Face and the project's usage guide.","contentHash":"sha256:02b54f739ba9c131adb1a5f68d63868f97819c5a0352fd36c09c58608731d896","authorName":"Anshu Pathak","authorUrl":"https://zyvop.com/author/anshu","authorSameAs":[],"category":"News","tags":["Cactus Compute","speech recognition","Whisper","on-device AI","edge computing"],"audience":"Software engineers and developers building applications with News","tone":"Practical and evidence-based engineering guidance","readingTimeMinutes":6,"wordCount":1316,"faqs":null,"primaryTopic":"News","publishedAt":"2026-10-09T04:50:12.410Z","updatedAt":"2026-10-09T04:50:12.410Z","canonicalUrl":"https://zyvop.com/whistle-speech-to-text-in-16-9-mb-subc8"}