Phonon-2 (open weights, on-device ASR)
English only. Not a new architecture: a compressed NVIDIA Parakeet TDT 0.6B v3 (tokenizer and output conventions unchanged). The launch post's 'more accurate than Whisper large at 1/10 the size' refers to whisper-large-v3-turbo in the vendor's own table; benchmarks not independently reproduced. Fermion Research is a small startup (founder Manan Gupta).
- Input
- audio
- Output
- text
- License
- cc-by-4.0
- Verified
- 2026-09-30
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | FermionResearch/Phonon-2 | huggingface.co/FermionResearch/Phonon-2 | — |
| PyPI (CLI) | fermion-research | pypi.org/project/fermion-research/ | docs |
| Docker | ghcr.io/fermionresearch/phonon-cpu:2.0.2 | — | — |
| Mac app (Detta) | — | www.fermionresearch.com/products/detta | — |
Notable capabilities (2)
- ~2-bit quantised Parakeet for on-device English ASR: A quantisation-aware-trained compression of NVIDIA parakeet-tdt-0.6b-v3: encoder weights at one of five learned levels (~2.1 bits), a 164 MB download vs the 2.5 GB teacher, averaging 5.21% WER on the Open ASR Leaderboard's seven English sets (teacher 4.96%, Whisper large-v3-turbo 6.58%, vendor-run numbers). source
- Fast local transcription: Vendor figures: about 174x realtime on an M5 MacBook Air (an hour of audio in ~20 s), 143x on eight Zen 5 cores, 6,680x on one H100 at batch 128. source
Launch post: https://x.com/yoitsmanan/status/2104990913886031993 (≈37k views). Related: nvidia-parakeet-canary, whisper-large-v3.