Fermion Research

Introducing Phonon-2

The most accurate open speech recognition model under 900 MB, in a 164 MB download that turns an hour of audio into text in about 20 seconds on a MacBook Air.

A field of warm mineral light gathering into a single bright band
Fermion ResearchPhonon-2

Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it holds the accuracy of its 2.5 GB full-precision teacher set for set, beats it on meetings and parliamentary speech, and turns an hour of audio into text in about 20 seconds on a MacBook Air.

We are releasing Phonon-2 today as an open model for English. Its encoder stores every weight as one of five learned levels in about 2.1 bits, and the same file runs on Macs, Linux, Windows and NVIDIA GPUs. The weights are released under CC-BY-4.0, the licence of NVIDIA’s Parakeet TDT 0.6B v3, from which they derive.

45678200 MB500 MB1 GB2 GB5 GB10 GB20 GB50 GB100 GBdownload size, published weight file, log scaleword error rate, seven-set average, %, better is upPareto frontParakeet Redux (Moondream) · 5.69 % · 178 MBnvidia/canary-180m-flash · 5.69 % · 737 MBmistralai/Voxtral-Mini-4B-Realtime-2602 · 6.12 % · 8,000 MBopenai/whisper-large-v3-turbo · 6.58 % · 1,618 MBnvidia/nemotron-3.5-asr-streaming-0.6b · 7.96 % · 2,368 MBPhonon-2 · 5.21 % · 164 MBnvidia/parakeet-tdt-0.6b-v3 · 4.96 % · 2,508 MBPhonon-1 · 6.56 % · 415 MBPhonon-2 · 5.21Parakeet TDT 0.6B v3 · 4.96Parakeet Redux · 5.69Phonon-1 · 6.56Canary 180M Flash · 5.69Voxtral Mini 4B Realtime · 6.12Whisper large-v3-turbo · 6.58Nemotron 3.5 ASR Streaming · 7.96164 MB2.5 GB178 MB415 MB737 MB8.0 GB*1.6 GB2.4 GBPhonon-2 and Phonon-1Parakeet TDT 0.6B v3, the teacherOpen ASR Leaderboard's published rowsreleased since August 2026
Figure 1Accuracy against download size. The eight models in Table 1. The dashed line joins the models that no smaller download beats; up and to the left is better.

Accuracy per byte

Across the Open ASR Leaderboard’s seven English sets Phonon-2 averages 5.21 % word error, and every open model that scores better is at least 5.8 times its size. It reaches 100.8 % of its teacher’s word accuracy on parliamentary speech and beats it on meetings, from a download 15 times smaller.

The lead holds under noise. At every level tested, from a quiet room to background noise as loud as the speaker, Phonon-2 stays ahead of Parakeet Redux, the other low-bit model of its size.

  • Phonon-2164 MB

    LS clean
    1.72
    LS other
    3.92
    AMI
    9.37
    Earnings-22
    6.96
    GigaSpeech
    8.35
    SPGISpeech
    3.70
    VoxPopuli
    2.46
    Average
    5.21
  • Parakeet TDT 0.6B v3†teacher2,508 MB

    LS clean
    1.52
    LS other
    3.13
    AMI
    9.42
    Earnings-22
    5.85
    GigaSpeech
    7.99
    SPGISpeech
    3.63
    VoxPopuli
    3.19
    Average
    4.96
  • Parakeet Redux178 MB

    LS clean
    1.94
    LS other
    4.35
    AMI
    9.16
    Earnings-22
    7.90
    GigaSpeech
    8.62
    SPGISpeech
    4.01
    VoxPopuli
    3.87
    Average
    5.69
  • Phonon-1415 MB

    LS clean
    2.11
    LS other
    5.03
    AMI
    10.31
    Earnings-22
    12.34
    GigaSpeech
    8.73
    SPGISpeech
    3.67
    VoxPopuli
    3.73
    Average
    6.56
  • Canary 180M Flash†737 MB

    LS clean
    1.52
    LS other
    3.42
    AMI
    12.09
    Earnings-22
    8.33
    GigaSpeech
    8.87
    SPGISpeech
    2.04
    VoxPopuli
    3.57
    Average
    5.69
  • Voxtral Mini 4B Realtime†≈8,000 MB*

    LS clean
    1.62
    LS other
    4.94
    AMI
    13.34
    Earnings-22
    9.31
    GigaSpeech
    8.80
    SPGISpeech
    2.23
    VoxPopuli
    2.60
    Average
    6.12
  • Whisper large-v3-turbo†1,618 MB

    LS clean
    2.13
    LS other
    3.71
    AMI
    13.88
    Earnings-22
    8.09
    GigaSpeech
    8.47
    SPGISpeech
    2.79
    VoxPopuli
    7.02
    Average
    6.58
  • Nemotron 3.5 ASR Streaming 0.6B†2,368 MB

    LS clean
    2.83
    LS other
    6.79
    AMI
    13.43
    Earnings-22
    15.30
    GigaSpeech
    9.86
    SPGISpeech
    3.27
    VoxPopuli
    4.24
    Average
    7.96
Table 1Word error rate, %, lower is better; the best value in each column is underlined. † The Open ASR Leaderboard’s published row; the other rows were scored with its code on the full test sets. * Size from the parameter count at 16 bits. Word accuracy kept is (100 − Phonon-2’s WER) / (100 − the teacher’s WER) per set; the highest is VoxPopuli, 97.54 / 96.81 = 100.75 %.

Speed on every surface

In the 164 MB file each encoder weight is zero or plus or minus one of two magnitudes stored per output row. The values are packed as base-3 digits five to a byte, with one further bit per non-zero weight selecting the magnitude, for about 2.1 bits per stored weight in all. Each engine unpacks the file once at load, and decoding never reads the packed bytes again.

On Apple silicon the whole model runs on the GPU through MLX. The encoder runs from a 16-bit dense copy built at load, and the greedy transducer decode stays on the GPU, synchronising with the host once every 16 steps rather than once per token. On x86-64 and Arm CPUs the encoder is requantised once to 8-bit integers per row and runs in C on one persistent thread pool. Its dot products use AVX-512 VNNI, AVX2 on older x86 and NEON on Arm, and each layer's normalisation, activation and residual are written straight into the next layer's 8-bit input. The transducer decode loop is also in C, over the 6-bit decoder tables. On NVIDIA GPUs the weights are expanded exactly to a 16-bit encoder, and inputs are padded to one-second buckets whose CUDA graphs are built at load. The label-looping decode advances the whole batch together, 16 steps per graph launch, so throughput on an A100 grows from 267 times realtime for one stream to 3,614 for 128.

Each fast path ships only because its word error stays within noise of the exact path, meaning a paired bootstrap over clips with 4,000 resamples gives a 95 % interval that covers zero. Word error is scored with the leaderboard's normaliser and scorer on the set each row names. The Linux CPU rows use all 2,939 clips of LibriSpeech test-other and the GPU rows use the seven sets. Speed is seconds of audio per second of wall time, one clip at a time, with log-mel included and model load excluded, unless a row says batch.

SurfaceTimes realtimeWord error, fast path against exact path
Apple M5 MacBook Air, GPU (MLX)174×2.94 % against 2.94 % (400 LibriSpeech utterances)
Apple M5 MacBook Air, CPU only40×2.33 % on a 40-clip check
Linux x86-64, eight Zen 5 cores (16 vCPU)142.8×3.94 % against 3.91 % (LibriSpeech test-other)
Linux Arm, eight Google Axion cores52.2×3.90 % against 3.91 % (LibriSpeech test-other)
Windows x64, 8 vCPU21.0×2.21 % on a 40-clip check
NVIDIA A100 80 GB267× one stream · 3,614× batch 1285.20–5.22 % against 5.20 % (seven sets)
NVIDIA H100 80 GB465× one stream · 6,680× batch 1285.20–5.22 % against 5.20 % (seven sets)
Table 2One stream at a time unless a row says batch.
Runtime on the same M5 MacBook AirTimes realtime
Phonon-2 (MLX)174.0×
FluidAudio, Parakeet TDT 0.6B v3 (Core ML)104.9×
FluidAudio, Parakeet Redux (Core ML)27.8×
Moonshine tiny26.7×
whisper.cpp, large-v3-turbo (Metal)17.0×
sherpa-onnx, Parakeet TDT 0.6B v3 (int8)16.5×
Table 3The same 20 dictations, 797 seconds of speech, on the same MacBook Air, one stream at a time with load time excluded, each runtime at its defaults.

On a MacBook Air an hour of audio becomes text in about 20 seconds, 1.7 times faster than any other runtime measured on the same machine. On eight CPU cores the hour takes 25 seconds, and on one H100 a full day of audio takes 13 seconds in batches of 128. A Core ML runtime for Apple devices is coming soon.

Try it

On a Mac with Apple silicon, pip install fermion-research installs the command line and fermion transcribe recording.wav runs Phonon-2, its default speech model; the model page lists the setup and measured speed on every platform. Detta, the dictation app for the Mac, runs Phonon-2 in any text field. The weights are on Hugging Face at FermionResearch/Phonon-2.

More from the lab

All research