Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it holds the accuracy of its 2.5 GB full-precision teacher set for set, beats it on meetings and parliamentary speech, and turns an hour of audio into text in about 20 seconds on a MacBook Air.
We are releasing Phonon-2 today as an open model for English. Its encoder stores every weight as one of five learned levels in about 2.1 bits, and the same file runs on Macs, Linux, Windows and NVIDIA GPUs. The weights are released under CC-BY-4.0, the licence of NVIDIA’s Parakeet TDT 0.6B v3, from which they derive.
Accuracy per byte
Across the Open ASR Leaderboard’s seven English sets Phonon-2 averages 5.21 % word error, and every open model that scores better is at least 5.8 times its size. It reaches 100.8 % of its teacher’s word accuracy on parliamentary speech and beats it on meetings, from a download 15 times smaller.
The lead holds under noise. At every level tested, from a quiet room to background noise as loud as the speaker, Phonon-2 stays ahead of Parakeet Redux, the other low-bit model of its size.
| Model | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Average |
|---|---|---|---|---|---|---|---|---|---|
| Phonon-2 | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 |
| Parakeet TDT 0.6B v3†teacher | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 |
| Parakeet Redux | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 |
| Phonon-1 | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 |
| Canary 180M Flash† | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 |
| Voxtral Mini 4B Realtime† | ≈8,000 MB* | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 |
| Whisper large-v3-turbo† | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 |
| Nemotron 3.5 ASR Streaming 0.6B† | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 |
Phonon-2164 MB
- LS clean
- 1.72
- LS other
- 3.92
- AMI
- 9.37
- Earnings-22
- 6.96
- GigaSpeech
- 8.35
- SPGISpeech
- 3.70
- VoxPopuli
- 2.46
- Average
- 5.21
Parakeet TDT 0.6B v3†teacher2,508 MB
- LS clean
- 1.52
- LS other
- 3.13
- AMI
- 9.42
- Earnings-22
- 5.85
- GigaSpeech
- 7.99
- SPGISpeech
- 3.63
- VoxPopuli
- 3.19
- Average
- 4.96
Parakeet Redux178 MB
- LS clean
- 1.94
- LS other
- 4.35
- AMI
- 9.16
- Earnings-22
- 7.90
- GigaSpeech
- 8.62
- SPGISpeech
- 4.01
- VoxPopuli
- 3.87
- Average
- 5.69
Phonon-1415 MB
- LS clean
- 2.11
- LS other
- 5.03
- AMI
- 10.31
- Earnings-22
- 12.34
- GigaSpeech
- 8.73
- SPGISpeech
- 3.67
- VoxPopuli
- 3.73
- Average
- 6.56
Canary 180M Flash†737 MB
- LS clean
- 1.52
- LS other
- 3.42
- AMI
- 12.09
- Earnings-22
- 8.33
- GigaSpeech
- 8.87
- SPGISpeech
- 2.04
- VoxPopuli
- 3.57
- Average
- 5.69
Voxtral Mini 4B Realtime†≈8,000 MB*
- LS clean
- 1.62
- LS other
- 4.94
- AMI
- 13.34
- Earnings-22
- 9.31
- GigaSpeech
- 8.80
- SPGISpeech
- 2.23
- VoxPopuli
- 2.60
- Average
- 6.12
Whisper large-v3-turbo†1,618 MB
- LS clean
- 2.13
- LS other
- 3.71
- AMI
- 13.88
- Earnings-22
- 8.09
- GigaSpeech
- 8.47
- SPGISpeech
- 2.79
- VoxPopuli
- 7.02
- Average
- 6.58
Nemotron 3.5 ASR Streaming 0.6B†2,368 MB
- LS clean
- 2.83
- LS other
- 6.79
- AMI
- 13.43
- Earnings-22
- 15.30
- GigaSpeech
- 9.86
- SPGISpeech
- 3.27
- VoxPopuli
- 4.24
- Average
- 7.96
Speed on every surface
In the 164 MB file each encoder weight is zero or plus or minus one of two magnitudes stored per output row. The values are packed as base-3 digits five to a byte, with one further bit per non-zero weight selecting the magnitude, for about 2.1 bits per stored weight in all. Each engine unpacks the file once at load, and decoding never reads the packed bytes again.
On Apple silicon the whole model runs on the GPU through MLX. The encoder runs from a 16-bit dense copy built at load, and the greedy transducer decode stays on the GPU, synchronising with the host once every 16 steps rather than once per token. On x86-64 and Arm CPUs the encoder is requantised once to 8-bit integers per row and runs in C on one persistent thread pool. Its dot products use AVX-512 VNNI, AVX2 on older x86 and NEON on Arm, and each layer's normalisation, activation and residual are written straight into the next layer's 8-bit input. The transducer decode loop is also in C, over the 6-bit decoder tables. On NVIDIA GPUs the weights are expanded exactly to a 16-bit encoder, and inputs are padded to one-second buckets whose CUDA graphs are built at load. The label-looping decode advances the whole batch together, 16 steps per graph launch, so throughput on an A100 grows from 267 times realtime for one stream to 3,614 for 128.
Each fast path ships only because its word error stays within noise of the exact path, meaning a paired bootstrap over clips with 4,000 resamples gives a 95 % interval that covers zero. Word error is scored with the leaderboard's normaliser and scorer on the set each row names. The Linux CPU rows use all 2,939 clips of LibriSpeech test-other and the GPU rows use the seven sets. Speed is seconds of audio per second of wall time, one clip at a time, with log-mel included and model load excluded, unless a row says batch.
| Surface | Times realtime | Word error, fast path against exact path |
|---|---|---|
| Apple M5 MacBook Air, GPU (MLX) | 174× | 2.94 % against 2.94 % (400 LibriSpeech utterances) |
| Apple M5 MacBook Air, CPU only | 40× | 2.33 % on a 40-clip check |
| Linux x86-64, eight Zen 5 cores (16 vCPU) | 142.8× | 3.94 % against 3.91 % (LibriSpeech test-other) |
| Linux Arm, eight Google Axion cores | 52.2× | 3.90 % against 3.91 % (LibriSpeech test-other) |
| Windows x64, 8 vCPU | 21.0× | 2.21 % on a 40-clip check |
| NVIDIA A100 80 GB | 267× one stream · 3,614× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |
| NVIDIA H100 80 GB | 465× one stream · 6,680× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |
| Runtime on the same M5 MacBook Air | Times realtime |
|---|---|
| Phonon-2 (MLX) | 174.0× |
| FluidAudio, Parakeet TDT 0.6B v3 (Core ML) | 104.9× |
| FluidAudio, Parakeet Redux (Core ML) | 27.8× |
| Moonshine tiny | 26.7× |
| whisper.cpp, large-v3-turbo (Metal) | 17.0× |
| sherpa-onnx, Parakeet TDT 0.6B v3 (int8) | 16.5× |
On a MacBook Air an hour of audio becomes text in about 20 seconds, 1.7 times faster than any other runtime measured on the same machine. On eight CPU cores the hour takes 25 seconds, and on one H100 a full day of audio takes 13 seconds in batches of 128. A Core ML runtime for Apple devices is coming soon.
Try it
On a Mac with Apple silicon, pip install fermion-research installs the command line and fermion transcribe recording.wav runs Phonon-2, its default speech model; the model page lists the setup and measured speed on every platform. Detta, the dictation app for the Mac, runs Phonon-2 in any text field. The weights are on Hugging Face at FermionResearch/Phonon-2.


