Fermion Research
Speech recognition · Available now

Phonon-2

The most accurate open speech recognition model under 900 MB. 164 MB, 5.21 % word error on seven public test sets, 174× realtime on a MacBook Air.

164 MB
Download, 177 MB on disk
5.21 %
Seven-set word error rate
174×
Realtime, one stream, M5 MacBook Air
6,680×
Realtime, batch 128, H100

Overview

Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it averages 5.21 % word error on the Open ASR Leaderboard’s seven English sets, and every open model that scores better is at least 5.8 times its size. Set for set it holds the accuracy of its 2.5 GB full-precision teacher and beats it on meetings and parliamentary speech. Its encoder stores every weight as one of five learned levels in about 2.1 bits.

It transcribes at 174 times realtime on an M5 MacBook Air, where Parakeet TDT 0.6B v3 in FluidAudio’s Core ML runtime reaches 104.9 on the same audio; at 143 times on eight Zen 5 cores (16 vCPU); and at 6,680 times on one H100 in batches of 128. A Core ML runtime for Apple devices is coming soon. The model writes punctuated, capitalized text, and the weights are released under CC-BY-4.0. Phonon-2 is the model behind Detta, the Fermion Research dictation app for the Mac.

Models

Phonon-2 is a single file. Phonon-1 remains available beside it.

Phonon-2Encoder about 2.1 bits per stored weight, 6-bit tables elsewhere
164 MB download · 177 MB on disk
Phonon-1The earlier model
415 MB download · 455 MB on disk
Language
English
Audio inputmicrophone or file
16 kHz
Output
Punctuated, capitalized text

The Phonon-1 spec

Evaluation

Phonon-2 averages 5.21 % word error on the seven public test sets of the Open ASR Leaderboard, scored with the board’s own code on the full test sets. The table sets it beside its full-precision teacher, Parakeet TDT 0.6B v3, which averages 4.96 in a 2,508 MB download, and six other open models from 178 MB to about 8 GB.

Seven-set comparison

  • Phonon-2Fermion Research · 0.60 B · 164 MB

    LS clean
    1.72
    LS other
    3.92
    AMI
    9.37
    Earnings-22
    6.96
    GigaSpeech
    8.35
    SPGISpeech
    3.70
    VoxPopuli
    2.46
    Average
    5.21
  • parakeet-tdt-0.6b-v3nvidia · 0.60 B · 2,508 MB

    LS clean
    1.52
    LS other
    3.13
    AMI
    9.42
    Earnings-22
    5.85
    GigaSpeech
    7.99
    SPGISpeech
    3.63
    VoxPopuli
    3.19
    Average
    4.96
  • Parakeet ReduxMoondream · 0.60 B · 178 MB

    LS clean
    1.94
    LS other
    4.35
    AMI
    9.16
    Earnings-22
    7.90
    GigaSpeech
    8.62
    SPGISpeech
    4.01
    VoxPopuli
    3.87
    Average
    5.69
  • Phonon-1Fermion Research · 0.78 B · 415 MB

    LS clean
    2.11
    LS other
    5.03
    AMI
    10.31
    Earnings-22
    12.34
    GigaSpeech
    8.73
    SPGISpeech
    3.67
    VoxPopuli
    3.73
    Average
    6.56
  • canary-180m-flashnvidia · 0.18 B · 737 MB

    LS clean
    1.52
    LS other
    3.42
    AMI
    12.09
    Earnings-22
    8.33
    GigaSpeech
    8.87
    SPGISpeech
    2.04
    VoxPopuli
    3.57
    Average
    5.69
  • Voxtral-Mini-4B-Realtime-2602mistralai · 4.00 B · 8,000 MB*

    LS clean
    1.62
    LS other
    4.94
    AMI
    13.34
    Earnings-22
    9.31
    GigaSpeech
    8.80
    SPGISpeech
    2.23
    VoxPopuli
    2.60
    Average
    6.12
  • whisper-large-v3-turboopenai · 0.80 B · 1,618 MB

    LS clean
    2.13
    LS other
    3.71
    AMI
    13.88
    Earnings-22
    8.09
    GigaSpeech
    8.47
    SPGISpeech
    2.79
    VoxPopuli
    7.02
    Average
    6.58
  • nemotron-3.5-asr-streaming-0.6bnvidia · 0.64 B · 2,368 MB

    LS clean
    2.83
    LS other
    6.79
    AMI
    13.43
    Earnings-22
    15.30
    GigaSpeech
    9.86
    SPGISpeech
    3.27
    VoxPopuli
    4.24
    Average
    7.96
Table 1Word error rate, %, on the seven sets, lower is better; bold marks the best value in each column. Leaderboard rows are its published results of 25 September 2026, and the other rows were scored with its code on the same full test sets.

Throughput

One 164 MB file runs on every surface, and the fast path on each keeps the accuracy of the exact one.

SurfaceTimes realtimeWord error, fast path against exact path
Apple M5 MacBook Air, GPU (MLX)174×2.94 % against 2.94 % (400 LibriSpeech utterances)
Apple M5 MacBook Air, CPU only40×2.33 % on a 40-clip check
Linux x86-64, eight Zen 5 cores (16 vCPU)142.8×3.94 % against 3.91 % (LibriSpeech test-other)
Linux Arm, eight Google Axion cores52.2×3.90 % against 3.91 % (LibriSpeech test-other)
Windows x64, 8 vCPU21.0×2.21 % on a 40-clip check
NVIDIA A100 80 GB267× one stream · 3,614× batch 1285.20–5.22 % against 5.20 % (seven sets)
NVIDIA H100 80 GB465× one stream · 6,680× batch 1285.20–5.22 % against 5.20 % (seven sets)
Table 2One stream at a time unless a row says batch.
Runtime on the same M5 MacBook AirTimes realtime
Phonon-2 (MLX)174.0×
FluidAudio, Parakeet TDT 0.6B v3 (Core ML)104.9×
FluidAudio, Parakeet Redux (Core ML)27.8×
Moonshine tiny26.7×
whisper.cpp, large-v3-turbo (Metal)17.0×
sherpa-onnx, Parakeet TDT 0.6B v3 (int8)16.5×
Table 3The same 20 dictations, 797 seconds of speech, on the same MacBook Air, one stream at a time with load time excluded, each runtime at its defaults.
1×3×10×30×100×300×realtime factor, one stream, Apple M5 MacBook Air 16 GB, log scaledownload WERPhonon-2 (MLX, default): 174× realtime, 5.7 ms per second of audio, WER 5.21 (our read, board scorer (accuracy within noise; 10/400 hypotheses differ))Phonon-2 (MLX, default)174× 164 MB 5.21Phonon-2 (MLX, exact decode): 109× realtime, 9.2 ms per second of audio, WER 5.21 (our read, board scorer)Phonon-2 (MLX, exact decode)109× 164 MB 5.21FluidAudio Parakeet TDT 0.6B v3 (Core ML): 104.9× realtime, 9.5 ms per second of audio, WER 4.96 (board's run)FluidAudio Parakeet TDT 0.6B v3 (Core ML)105× 483 MB 4.96Phonon-1 Big (app default): 52× realtime, 19.2 ms per second of audio, WER 6.65 (our read, board scorer)Phonon-1 Big (app default)52× 581 MB 6.65FluidAudio Parakeet Redux (Core ML): 27.8× realtime, 35.9 ms per second of audioFluidAudio Parakeet Redux (Core ML)28× 220 MB —Moonshine tiny (moonshine-voice): 26.7× realtime, 37.4 ms per second of audioMoonshine tiny (moonshine-voice)27× 44 MB —Moonshine base (moonshine-voice): 21.6× realtime, 46.3 ms per second of audioMoonshine base (moonshine-voice)22× 141 MB —whisper.cpp large-v3-turbo (ggml f16): 17× realtime, 58.7 ms per second of audio, WER 6.58 (board's run)whisper.cpp large-v3-turbo (ggml f16)17× 1,625 MB 6.58sherpa-onnx Parakeet TDT 0.6B v3 (int8): 16.5× realtime, 60.5 ms per second of audio, WER 4.96 (board's run)sherpa-onnx Parakeet TDT 0.6B v3 (int8)17× 670 MB 4.96whisper.cpp large-v3-turbo (q5_0): 12.2× realtime, 81.8 ms per second of audio, WER 6.58 (board's run)whisper.cpp large-v3-turbo (q5_0)12× 574 MB 6.58Moonshine medium-streaming (moonshine-voice): 6.1× realtime, 163.5 ms per second of audioMoonshine medium-streaming (moonshine-voice)6× 269 MB —whisper.cpp large-v3 (ggml f16): 3.2× realtime, 309.6 ms per second of audio, WER 6.04 (board's run)whisper.cpp large-v3 (ggml f16)3× 3,095 MB 6.04whisper.cpp large-v3 (q5_0): 2.7× realtime, 363.8 ms per second of audio, WER 6.04 (board's run)whisper.cpp large-v3 (q5_0)3× 1,081 MB 6.04
Figure 1Speed on a MacBook Air. Single-stream realtime factor on an Apple M5 MacBook Air, 16 GB.

Run it

Phonon-2 is the model inside Detta, the dictation app for the Mac.

$ pip install fermion-research
$ pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
$ fermion transcribe recording.wav --model phonon-2

On a Mac with Apple silicon, the second line installs the speech runtime and the model downloads on first use. The name phonon on its own also selects Phonon-2, and --model phonon-1 selects the earlier model.

$ fermion listen --model phonon-2
$ fermion serve --model phonon-2

fermion listen transcribes the microphone live, and fermion serve runs a local OpenAI-compatible transcription endpoint.

$ docker run --rm -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cpu:2.0.2 transcribe /audio/recording.wav --model phonon-2
$ docker run --rm --gpus all -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cuda:1.0.3 transcribe /audio/recording.wav --model phonon-2

The same package runs on Linux (x86-64 and Arm) and Windows CPUs. With Docker, Phonon-2 runs on a CPU or a GPU.

Weights on Hugging Facefermion-research on PyPIEngines on GitHub

Availability

Open weightsCC-BY-4.0Apple siliconNVIDIA GPUsCPUs

License

The weights are released under the Creative Commons Attribution 4.0 licence, which Phonon-2 inherits from NVIDIA’s Parakeet TDT 0.6B v3. The licence permits commercial use, modification and redistribution with attribution. The command line is released under Apache 2.0.

Citation

@misc{fermionresearch2026phonon2,
  title  = {Phonon-2},
  author = {{Fermion Research}},
  year   = {2026},
  url    = {https://fermionresearch.com/models/phonon-2/}
}

Related research

All research