Fermion Research

Detta now runs on the Neural Engine

Faster dictation, a sixth of the memory at rest, and the same accuracy.

Detta
Fermion ResearchDetta

Today, Detta recognises speech entirely on your Mac’s Neural Engine. Text appears faster after you speak, the app holds a sixth of the memory it did at rest, and accuracy is unchanged. On a MacBook Air, Detta now transcribes and polishes an hour of speech in 23 seconds.

Until this release, Detta recognised speech on the GPU, the same chip that draws your screen, and held 1.8 GB of memory while it waited for you to talk. It now runs on the processor Apple built for neural networks and holds 0.3 GB.

What dictation needs

Dictation has to feel immediate. If text takes more than about a quarter of a second to appear after the key comes up, the delay is noticeable. It also has to stay out of the way on a Mac that is already busy with a browser, a video call and everything else.

Detta meets both on the Mac itself. Speech is recognised on your device, dictation keeps working without a connection, and the app asks for almost nothing while you are not speaking.

Speed

Detta starts listening the moment the key goes down, keeping the very first syllable, so the model is already warm when you finish speaking. Each dictation then costs only as much compute as its length requires, which means a quick "sounds good" returns as fast as the hardware allows rather than paying for a long paragraph. The median time from key-up to text is now 0.26 seconds, down from 0.33.

Long recordings benefit just as much. A 110-minute recording comes back as a finished transcript, polished from the first word to the last, in 49 seconds, and as a raw transcript in under 13. Every word keeps its start and end time, so a transcript in Detta follows along as the recording plays.

Phonon-2 on its own reads an hour of speech in 6.0 seconds, 606 times real time. 24 hours of audio takes under two and a half minutes.

Using the Neural Engine effectively

Every Apple silicon Mac has a Neural Engine. It is fast and efficient, and it is also strict about what it will run, so models are usually approximated to fit it and lose accuracy along the way.

The Neural Engine build of Phonon-2 scores 5.21 % word error on the Open ASR Leaderboard’s seven English test sets, the same as Phonon-2 on every other runtime.

Four sizes of the model, from short phrases to long paragraphs, share a single copy of its weights, so the three extra sizes add less than 2 MB.

On a MacBook Air it worked through 158 hours of benchmark audio in about half an hour. With the model loaded, a few seconds of speech becomes text in about 11 milliseconds. Recognition never touches the GPU, and transcribing an hour of audio costs the CPU about five seconds of work.

The first time Detta runs on a Mac, it sets the Neural Engine up for that chip. Dictation works within about a minute, and on a MacBook Air Detta reaches full speed within about four. From then on, it opens at full speed instantly.

Cleaning up the transcript

Phonon-2 writes down what you said, which is not always what you meant to type. People restart sentences, correct themselves halfway through and repeat words. Say "let’s meet on Tuesday, no, Wednesday at noon", and an accurate transcript keeps every word, although only the correction belongs in the calendar invite.

Gluon handles that step. It is the second model inside Detta, a compact language model that reads Phonon-2’s transcript, removes fillers and false starts, applies corrections, settles punctuation and casing, and writes numbers and dates the way you would type them. The example above reaches your document as "Let’s meet on Wednesday at noon." For a typical sentence the pass adds about 80 milliseconds.

Gluon tidies your words without dropping them. If an edit would remove a whole run of speech rather than a filler or a restart, Detta keeps that stretch exactly as you said it. Gluon also leaves memory completely ten minutes after your last sentence and returns when you speak again.

Phonon-2 runs on the Neural Engine and Gluon on the GPU, so the two models work side by side and never queue behind each other.

Your dictionary now steers Phonon-2 while it decodes, so a word the model has never seen, a product name, a colleague’s surname, a function in your codebase, is recognised rather than repaired afterwards. Recall of dictionary terms rose from 0.54 to 0.79 with zero false substitutions, and with an empty dictionary the transcript is exactly what it was before.

The same hotword biasing ships today in the open Phonon-2 runtime on every decoder, as --hotwords on the command line, hotwords= in Python and a hotwords field on the server, which also reads the OpenAI-compatible prompt field as a vocabulary list.

Memory

For an app that is always running, the number that matters most is the memory Detta holds while you are not dictating.

Detta previously held 1.8 GB at rest. It now holds 0.6 GB with Gluon loaded, and with Gluon off it rests at under 0.3 GB. A typical dictation peaks at 0.81 GB, and even the heaviest, ten-minute dictations peak at 1.1 GB, down from 3.1 GB.

At rest
Detta (previous)
1.8 GB
Detta
0.3 GB
At rest, Gluon loaded
Detta (previous)
1.8 GB
Detta
0.6 GB
Peak of a dictation
Detta (previous)
3.1 GB
Detta
1.1 GB
Figure 1Memory on an M5 MacBook Air with 16 GB.
Detta (previous)Detta
Text ready after the key comes up, median0.33 s0.26 s
One hour of audio, transcribed and polished31 s23 s
One hour of audio, Phonon-2 alone20 s6.0 s
One hour of audio, word error2.6 %2.2 %
Table 1One M5 MacBook Air with 16 GB. The third row times Phonon-2 alone with Gluon off, on the GPU in the previous Detta and on the Neural Engine now.

Availability

Detta is free for Macs with Apple silicon, and existing copies update from the menu bar. New users can download it here. For developers, the Core ML package of Phonon-2 is open on Hugging Face today, so any app on the Mac can put the Neural Engine to work on speech.

Detta
Phonon-2
Phonon-2 Core ML package on Hugging Face
Phonon-2 Core ML runtime on GitHub
Phonon runtime on GitHub

More from the lab

All research