Cactus Whistle (16.9 MB on-device speech-to-text)
Aimed at phones, wearables, robots, smart home, cars and microcontrollers. Compared with Whisper base (a small 2022 model), not with large ASR models; vendor-run numbers. Launch post ~167k views; Nobara/GE-Proton maintainer GloriousEggroll reported it made his home assistant respond 'SO damn fast'.
- Input
- audio
- Output
- text
- License
- apache-2.0
- Verified
- 2026-10-03
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | Cactus-Compute/whistle | huggingface.co/Cactus-Compute/whistle | — |
| GitHub (engine) | — | github.com/Cactus-Compute/needle3 | — |
| Browser demo | — | cactuscompute.com/blog/whistle | — |
Notable capabilities (1)
- Speech recognition in a 16.9 MB file, on CPU: One 16.9 MB file with no dependencies, running in the same C++ engine as Cactus's Needle model; ~11 ms to first token; transcribes up to 30 s per pass in English, German, French, Spanish, Italian, Dutch and Polish with language detection, word timestamps and speech embeddings. Vendor benchmarks: ahead of Whisper base on most sets at ~9× smaller and ~6× faster. source
Launch: https://x.com/cactuscompute/status/2106083041265562075 · Related: phonon-2, whisper-large-v3.