Phonon-2 starts much faster and decodes faster on Intel and AMD processors, on Linux and on Intel Macs, and uses less memory while it runs. Transcripts are unchanged, byte for byte.
segments gives the start and end of every sentence or pause-sized stretch of speech, in --json and in the server's verbose_json, in the shape OpenAI clients already read.
pip install fermion-research is now the whole setup on Linux, Windows and Intel Macs. On Apple silicon, pip install "fermion-research[mlx]" adds the MLX runtime.
The CUDA image phonon-cuda:1.0.7 returns segments and word timestamps, and the CPU image phonon-cpu:2.0.8 runs this release on amd64 and arm64.
Speech is recognised on your Mac's Neural Engine. Text appears faster after you speak, with a median of 0.26 seconds from key-up to text, down from 0.33, at the same accuracy.
Detta holds about a sixth of the memory it did while it waits for you to talk.
Your dictionary now steers recognition while it decodes, so names and terms you add are recognised more reliably.
Long recordings are polished from the first word to the last and keep every sentence, and each word keeps its start and end time as the recording plays.
The first time it runs, Detta lets you know while it sets up the Neural Engine for your Mac. Dictation works within about a minute.
Hotwords on every Phonon-2 engine. Give it up to 25 names and terms and it favours them whenever the audio is close. Without hotwords, every transcript is unchanged.
--hotwords on the command line, hotwords= in Python, and a hotwords field on the server, which also reads the OpenAI prompt field as a vocabulary list.
Hotwords work on Apple silicon, on Linux, Windows and macOS CPUs, and in the CUDA image phonon-cuda:1.0.6.
Intel Macs run Phonon-2 on the CPU engine (0.2.8, first published in this release).
Phonon-2 for the Neural Engine, at the accuracy of the reference engine, 5.21 % average word error on the Open ASR Leaderboard's seven English test sets.
On an M5 MacBook Air an hour of speech becomes text in 6 seconds, and the GPU stays free.
Windows from 5 to 35 seconds share one set of weights, and every word carries its start and end time.
A Swift package and command-line tool for macOS 15 and iOS 18 or later, and a Python runner.
The CPU engine reads your processor's features and picks the matching tier, from AVX2, AVX-512 and AMX on x86-64 to dotprod, i8mm and SVE on 64-bit Arm.
Raspberry Pi 3 and 4, other Arm boards without dotprod, and x86-64 processors without AVX2 run Phonon-2 on a baseline tier.
fermion describe shows the features found, the tier chosen and the binaries loaded, with --json for scripts.
An open English speech recognition model in a 164 MB download, averaging 5.21 % word error on the Open ASR Leaderboard's seven test sets.
174 times realtime through MLX on an M5 MacBook Air, and 143 times on eight Zen 5 cores.
fermion-research 0.2.0 makes it the default speech model on Apple silicon and on CPUs across macOS, Linux and Windows, with CPU and CUDA container images.