---
title: "Phonon-2"
description: "Phonon-2 downloads, benchmarks and runtimes. A 164 MB open speech recognition model for English at 5.21 % word error, for Mac, Linux, Windows and NVIDIA GPUs."
canonical: "https://www.fermionresearch.com/models/phonon-2/"
source: "Fermion Research"
---

# Phonon-2

The most accurate open speech recognition model under 900 MB. 164 MB, 5.21 % word error on seven public test sets, 174× realtime on a MacBook Air.

## Overview

Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it averages 5.21 % word error on the Open ASR Leaderboard’s seven English sets, and every open model that scores better is at least 5.8 times its size. Set for set it holds the accuracy of its 2.5 GB full-precision teacher and beats it on meetings and parliamentary speech. Its encoder stores every weight as one of five learned levels in about 2.1 bits.

It transcribes at 174 times realtime on an M5 MacBook Air, where Parakeet TDT 0.6B v3 in FluidAudio’s Core ML runtime reaches 104.9 on the same audio; at 143 times on eight Zen 5 cores (16 vCPU); and at 6,680 times on one H100 in batches of 128. A Core ML runtime for Apple devices is coming soon. The model writes punctuated, capitalized text, and the weights are released under CC-BY-4.0. Phonon-2 is the model behind [Detta](https://www.fermionresearch.com/products/detta/), the Fermion Research dictation app for the Mac.

## Models

Phonon-2 is a single file. Phonon-1 remains available beside it.

[The Phonon-1 spec](https://www.fermionresearch.com/models/phonon-1/)

## Evaluation

Phonon-2 averages 5.21 % word error on the seven public test sets of the Open ASR Leaderboard, scored with the board’s own code on the full test sets. The table sets it beside its full-precision teacher, Parakeet TDT 0.6B v3, which averages 4.96 in a 2,508 MB download, and six other open models from 178 MB to about 8 GB.

Seven-set comparison

| Model | Params | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Average |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Phonon-2Fermion Research | 0.60 B | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 |
| parakeet-tdt-0.6b-v3nvidia | 0.60 B | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 |
| Parakeet ReduxMoondream | 0.60 B | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 |
| Phonon-1Fermion Research | 0.78 B | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 |
| canary-180m-flashnvidia | 0.18 B | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 |
| Voxtral-Mini-4B-Realtime-2602mistralai | 4.00 B | 8,000 MB* | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 |
| whisper-large-v3-turboopenai | 0.80 B | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 |
| nemotron-3.5-asr-streaming-0.6bnvidia | 0.64 B | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 |

- Phonon-2Fermion Research · 0.60 B · 164 MB LS clean1.72 LS other3.92 AMI9.37 Earnings-226.96 GigaSpeech8.35 SPGISpeech3.70 VoxPopuli2.46 Average5.21
- parakeet-tdt-0.6b-v3nvidia · 0.60 B · 2,508 MB LS clean1.52 LS other3.13 AMI9.42 Earnings-225.85 GigaSpeech7.99 SPGISpeech3.63 VoxPopuli3.19 Average4.96
- Parakeet ReduxMoondream · 0.60 B · 178 MB LS clean1.94 LS other4.35 AMI9.16 Earnings-227.90 GigaSpeech8.62 SPGISpeech4.01 VoxPopuli3.87 Average5.69
- Phonon-1Fermion Research · 0.78 B · 415 MB LS clean2.11 LS other5.03 AMI10.31 Earnings-2212.34 GigaSpeech8.73 SPGISpeech3.67 VoxPopuli3.73 Average6.56
- canary-180m-flashnvidia · 0.18 B · 737 MB LS clean1.52 LS other3.42 AMI12.09 Earnings-228.33 GigaSpeech8.87 SPGISpeech2.04 VoxPopuli3.57 Average5.69
- Voxtral-Mini-4B-Realtime-2602mistralai · 4.00 B · 8,000 MB* LS clean1.62 LS other4.94 AMI13.34 Earnings-229.31 GigaSpeech8.80 SPGISpeech2.23 VoxPopuli2.60 Average6.12
- whisper-large-v3-turboopenai · 0.80 B · 1,618 MB LS clean2.13 LS other3.71 AMI13.88 Earnings-228.09 GigaSpeech8.47 SPGISpeech2.79 VoxPopuli7.02 Average6.58
- nemotron-3.5-asr-streaming-0.6bnvidia · 0.64 B · 2,368 MB LS clean2.83 LS other6.79 AMI13.43 Earnings-2215.30 GigaSpeech9.86 SPGISpeech3.27 VoxPopuli4.24 Average7.96

Table 1Word error rate, %, on the seven sets, lower is better; bold marks the best value in each column. Leaderboard rows are its published results of 25 September 2026, and the other rows were scored with its code on the same full test sets.

## Throughput

One 164 MB file runs on every surface, and the fast path on each keeps the accuracy of the exact one.

| Surface | Times realtime | Word error, fast path against exact path |
| --- | --- | --- |
| Apple M5 MacBook Air, GPU (MLX) | 174× | 2.94 % against 2.94 % (400 LibriSpeech utterances) |
| Apple M5 MacBook Air, CPU only | 40× | 2.33 % on a 40-clip check |
| Linux x86-64, eight Zen 5 cores (16 vCPU) | 142.8× | 3.94 % against 3.91 % (LibriSpeech test-other) |
| Linux Arm, eight Google Axion cores | 52.2× | 3.90 % against 3.91 % (LibriSpeech test-other) |
| Windows x64, 8 vCPU | 21.0× | 2.21 % on a 40-clip check |
| NVIDIA A100 80 GB | 267× one stream · 3,614× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |
| NVIDIA H100 80 GB | 465× one stream · 6,680× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |

Table 2One stream at a time unless a row says batch.

| Runtime on the same M5 MacBook Air | Times realtime |
| --- | --- |
| Phonon-2 (MLX) | 174.0× |
| FluidAudio, Parakeet TDT 0.6B v3 (Core ML) | 104.9× |
| FluidAudio, Parakeet Redux (Core ML) | 27.8× |
| Moonshine tiny | 26.7× |
| whisper.cpp, large-v3-turbo (Metal) | 17.0× |
| sherpa-onnx, Parakeet TDT 0.6B v3 (int8) | 16.5× |

Table 3The same 20 dictations, 797 seconds of speech, on the same MacBook Air, one stream at a time with load time excluded, each runtime at its defaults.

Figure 1Speed on a MacBook Air. Single-stream realtime factor on an Apple M5 MacBook Air, 16 GB.

## Run it

Phonon-2 is the model inside [Detta](https://www.fermionresearch.com/products/detta/), the dictation app for the Mac.

```text
$ pip install fermion-research
$ pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
$ fermion transcribe recording.wav --model phonon-2
```

On a Mac with Apple silicon, the second line installs the speech runtime and the model downloads on first use. The name phonon on its own also selects Phonon-2, and --model phonon-1 selects the earlier model.

```text
$ fermion listen --model phonon-2
$ fermion serve --model phonon-2
```

fermion listen transcribes the microphone live, and fermion serve runs a local OpenAI-compatible transcription endpoint.

```text
$ docker run --rm -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cpu:2.0.2 transcribe /audio/recording.wav --model phonon-2
$ docker run --rm --gpus all -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cuda:1.0.3 transcribe /audio/recording.wav --model phonon-2
```

The same package runs on Linux (x86-64 and Arm) and Windows CPUs. With Docker, Phonon-2 runs on a CPU or a GPU.

[Weights on Hugging Face](https://huggingface.co/FermionResearch/Phonon-2)[fermion-research on PyPI](https://pypi.org/project/fermion-research/)[Engines on GitHub](https://github.com/fermionresearch/phonon)

## Availability

### License

The weights are released under the Creative Commons Attribution 4.0 licence, which Phonon-2 inherits from NVIDIA’s Parakeet TDT 0.6B v3. The licence permits commercial use, modification and redistribution with attribution. The command line is released under Apache 2.0.

### Citation

```text
@misc{fermionresearch2026phonon2,
  title  = {Phonon-2},
  author = {{Fermion Research}},
  year   = {2026},
  url    = {https://fermionresearch.com/models/phonon-2/}
}
```

## Related research

![A field of warm mineral light gathering into a single bright band](https://www.fermionresearch.com/images/launch/phonon-hero.jpg)

ResearchDate pending

### Introducing Phonon-2

The most accurate open speech recognition model under 900 MB, in a 164 MB download that turns an hour of audio into text in about 20 seconds on a MacBook Air.

![A field of warm mineral light gathering into a single bright band](https://www.fermionresearch.com/images/launch/phonon-hero.jpg)

AnnouncementAugust 28, 2026

### Introducing Phonon-1

An open speech recognition model for English in a 415 MB download. It runs on a laptop or a datacenter GPU and transcribes an hour of audio in about two and a half minutes.

![The Detta window on macOS](https://www.fermionresearch.com/images/launch/detta-window.jpg)

AnnouncementDate pending

### Detta

A dictation app for the Mac that runs Phonon-2 on the device. Hold a key, speak, and the words appear in any text field.

Source: [https://www.fermionresearch.com/models/phonon-2/](https://www.fermionresearch.com/models/phonon-2/)
