---
title: "Introducing Phonon-2"
description: "The most accurate open speech recognition model under 900 MB. 164 MB, 5.21 % word error on seven public test sets, 174× realtime on a MacBook Air."
canonical: "https://www.fermionresearch.com/research/phonon-2/"
source: "Fermion Research"
---

# Introducing Phonon-2

The most accurate open speech recognition model under 900 MB, in a 164 MB download that turns an hour of audio into text in about 20 seconds on a MacBook Air.

![A field of warm mineral light gathering into a single bright band](https://www.fermionresearch.com/images/launch/phonon-hero.jpg)

![Fermion Research](https://www.fermionresearch.com/images/brand/newlogo.svg)

Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it holds the accuracy of its 2.5 GB full-precision teacher set for set, beats it on meetings and parliamentary speech, and turns an hour of audio into text in about 20 seconds on a MacBook Air.

We are releasing Phonon-2 today as an open model for English. Its encoder stores every weight as one of five learned levels in about 2.1 bits, and the same file runs on Macs, Linux, Windows and NVIDIA GPUs. The weights are released under CC-BY-4.0, the licence of NVIDIA’s Parakeet TDT 0.6B v3, from which they derive.

Figure 1Accuracy against download size. The eight models in Table 1. The dashed line joins the models that no smaller download beats; up and to the left is better.

## Accuracy per byte

Across the Open ASR Leaderboard’s seven English sets Phonon-2 averages 5.21 % word error, and every open model that scores better is at least 5.8 times its size. It reaches 100.8 % of its teacher’s word accuracy on parliamentary speech and beats it on meetings, from a download 15 times smaller.

The lead holds under noise. At every level tested, from a quiet room to background noise as loud as the speaker, Phonon-2 stays ahead of Parakeet Redux, the other low-bit model of its size.

| Model | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Average |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Phonon-2 | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 |
| Parakeet TDT 0.6B v3†teacher | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 |
| Parakeet Redux | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 |
| Phonon-1 | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 |
| Canary 180M Flash† | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 |
| Voxtral Mini 4B Realtime† | ≈8,000 MB* | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 |
| Whisper large-v3-turbo† | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 |
| Nemotron 3.5 ASR Streaming 0.6B† | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 |

- Phonon-2164 MB LS clean1.72 LS other3.92 AMI9.37 Earnings-226.96 GigaSpeech8.35 SPGISpeech3.70 VoxPopuli2.46 Average5.21
- Parakeet TDT 0.6B v3†teacher2,508 MB LS clean1.52 LS other3.13 AMI9.42 Earnings-225.85 GigaSpeech7.99 SPGISpeech3.63 VoxPopuli3.19 Average4.96
- Parakeet Redux178 MB LS clean1.94 LS other4.35 AMI9.16 Earnings-227.90 GigaSpeech8.62 SPGISpeech4.01 VoxPopuli3.87 Average5.69
- Phonon-1415 MB LS clean2.11 LS other5.03 AMI10.31 Earnings-2212.34 GigaSpeech8.73 SPGISpeech3.67 VoxPopuli3.73 Average6.56
- Canary 180M Flash†737 MB LS clean1.52 LS other3.42 AMI12.09 Earnings-228.33 GigaSpeech8.87 SPGISpeech2.04 VoxPopuli3.57 Average5.69
- Voxtral Mini 4B Realtime†≈8,000 MB* LS clean1.62 LS other4.94 AMI13.34 Earnings-229.31 GigaSpeech8.80 SPGISpeech2.23 VoxPopuli2.60 Average6.12
- Whisper large-v3-turbo†1,618 MB LS clean2.13 LS other3.71 AMI13.88 Earnings-228.09 GigaSpeech8.47 SPGISpeech2.79 VoxPopuli7.02 Average6.58
- Nemotron 3.5 ASR Streaming 0.6B†2,368 MB LS clean2.83 LS other6.79 AMI13.43 Earnings-2215.30 GigaSpeech9.86 SPGISpeech3.27 VoxPopuli4.24 Average7.96

Table 1Word error rate, %, lower is better; the best value in each column is underlined. † The Open ASR Leaderboard’s published row; the other rows were scored with its code on the full test sets. * Size from the parameter count at 16 bits. Word accuracy kept is (100 − Phonon-2’s WER) / (100 − the teacher’s WER) per set; the highest is VoxPopuli, 97.54 / 96.81 = 100.75 %.

## Speed on every surface

In the 164 MB file each encoder weight is zero or plus or minus one of two magnitudes stored per output row. The values are packed as base-3 digits five to a byte, with one further bit per non-zero weight selecting the magnitude, for about 2.1 bits per stored weight in all. Each engine unpacks the file once at load, and decoding never reads the packed bytes again.

On Apple silicon the whole model runs on the GPU through MLX. The encoder runs from a 16-bit dense copy built at load, and the greedy transducer decode stays on the GPU, synchronising with the host once every 16 steps rather than once per token. On x86-64 and Arm CPUs the encoder is requantised once to 8-bit integers per row and runs in C on one persistent thread pool. Its dot products use AVX-512 VNNI, AVX2 on older x86 and NEON on Arm, and each layer's normalisation, activation and residual are written straight into the next layer's 8-bit input. The transducer decode loop is also in C, over the 6-bit decoder tables. On NVIDIA GPUs the weights are expanded exactly to a 16-bit encoder, and inputs are padded to one-second buckets whose CUDA graphs are built at load. The label-looping decode advances the whole batch together, 16 steps per graph launch, so throughput on an A100 grows from 267 times realtime for one stream to 3,614 for 128.

Each fast path ships only because its word error stays within noise of the exact path, meaning a paired bootstrap over clips with 4,000 resamples gives a 95 % interval that covers zero. Word error is scored with the leaderboard's normaliser and scorer on the set each row names. The Linux CPU rows use all 2,939 clips of LibriSpeech test-other and the GPU rows use the seven sets. Speed is seconds of audio per second of wall time, one clip at a time, with log-mel included and model load excluded, unless a row says batch.

| Surface | Times realtime | Word error, fast path against exact path |
| --- | --- | --- |
| Apple M5 MacBook Air, GPU (MLX) | 174× | 2.94 % against 2.94 % (400 LibriSpeech utterances) |
| Apple M5 MacBook Air, CPU only | 40× | 2.33 % on a 40-clip check |
| Linux x86-64, eight Zen 5 cores (16 vCPU) | 142.8× | 3.94 % against 3.91 % (LibriSpeech test-other) |
| Linux Arm, eight Google Axion cores | 52.2× | 3.90 % against 3.91 % (LibriSpeech test-other) |
| Windows x64, 8 vCPU | 21.0× | 2.21 % on a 40-clip check |
| NVIDIA A100 80 GB | 267× one stream · 3,614× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |
| NVIDIA H100 80 GB | 465× one stream · 6,680× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |

Table 2One stream at a time unless a row says batch.

| Runtime on the same M5 MacBook Air | Times realtime |
| --- | --- |
| Phonon-2 (MLX) | 174.0× |
| FluidAudio, Parakeet TDT 0.6B v3 (Core ML) | 104.9× |
| FluidAudio, Parakeet Redux (Core ML) | 27.8× |
| Moonshine tiny | 26.7× |
| whisper.cpp, large-v3-turbo (Metal) | 17.0× |
| sherpa-onnx, Parakeet TDT 0.6B v3 (int8) | 16.5× |

Table 3The same 20 dictations, 797 seconds of speech, on the same MacBook Air, one stream at a time with load time excluded, each runtime at its defaults.

On a MacBook Air an hour of audio becomes text in about 20 seconds, 1.7 times faster than any other runtime measured on the same machine. On eight CPU cores the hour takes 25 seconds, and on one H100 a full day of audio takes 13 seconds in batches of 128. A Core ML runtime for Apple devices is coming soon.

## Try it

On a Mac with Apple silicon, `pip install fermion-research` installs the command line and `fermion transcribe recording.wav` runs Phonon-2, its default speech model; the [model page](https://www.fermionresearch.com/models/phonon-2/) lists the setup and measured speed on every platform. [Detta](https://www.fermionresearch.com/products/detta/), the dictation app for the Mac, runs Phonon-2 in any text field. The weights are on Hugging Face at [FermionResearch/Phonon-2](https://huggingface.co/FermionResearch/Phonon-2).

## More from the lab

![An amber cellular orbit gathering around a compact central form](https://www.fermionresearch.com/images/backgrounds/amber-orbit.webp)

AnnouncementJuly 27, 2026

### Introducing the Neutrino-1 models

Three open-weight models trained for a compact ternary format and released with CUDA, Apple-silicon, and x86 runtimes.

![A cobalt cellular membrane becoming progressively finer](https://www.fermionresearch.com/images/backgrounds/cobalt-cell.webp)

ResearchJuly 27, 2026

### Intelligence at one-eighth the bits

One-shot two-bit conversion lands near chance. Training inside the constraint produces 72.1 MMLU from a 3.88 GB artifact. This is what changed inside the weights.

Source: [https://www.fermionresearch.com/research/phonon-2/](https://www.fermionresearch.com/research/phonon-2/)
