Fermion Research

Changelog

Every release across Detta, the Phonon runtime, Core ML and our models, newest first.

October 2026

· Phonon runtime

fermion-research 0.2.10

Faster on Intel and AMD, with segment timestamps

  • Phonon-2 starts much faster and decodes faster on Intel and AMD processors, on Linux and on Intel Macs, and uses less memory while it runs. Transcripts are unchanged, byte for byte.
  • segments gives the start and end of every sentence or pause-sized stretch of speech, in --json and in the server's verbose_json, in the shape OpenAI clients already read.
  • pip install fermion-research is now the whole setup on Linux, Windows and Intel Macs. On Apple silicon, pip install "fermion-research[mlx]" adds the MLX runtime.
  • The CUDA image phonon-cuda:1.0.7 returns segments and word timestamps, and the CPU image phonon-cpu:2.0.8 runs this release on amd64 and arm64.

Release notes · PyPI · Docs

· Detta

Detta 1.0.26

Detta now runs on the Neural Engine

  • Speech is recognised on your Mac's Neural Engine. Text appears faster after you speak, with a median of 0.26 seconds from key-up to text, down from 0.33, at the same accuracy.
  • Detta holds about a sixth of the memory it did while it waits for you to talk.
  • Your dictionary now steers recognition while it decodes, so names and terms you add are recognised more reliably.
  • Long recordings are polished from the first word to the last and keep every sentence, and each word keeps its start and end time as the recording plays.
  • The first time it runs, Detta lets you know while it sets up the Neural Engine for your Mac. Dictation works within about a minute.

Read the post · Download · Release notes

· Core ML

phonon-coreml 1.1.1

Any audio file, with word timings

  • The command-line tool reads any sample rate and the common audio formats, m4a included.
  • --words prints every word on its own line with its start and end time.
  • A clearer notice on the first run, and clearer error messages.

Release · Run it

· Phonon runtime

fermion-research 0.2.9

Hotwords, and Phonon-2 on Intel Macs

  • Hotwords on every Phonon-2 engine. Give it up to 25 names and terms and it favours them whenever the audio is close. Without hotwords, every transcript is unchanged.
  • --hotwords on the command line, hotwords= in Python, and a hotwords field on the server, which also reads the OpenAI prompt field as a vocabulary list.
  • Hotwords work on Apple silicon, on Linux, Windows and macOS CPUs, and in the CUDA image phonon-cuda:1.0.6.
  • Intel Macs run Phonon-2 on the CPU engine (0.2.8, first published in this release).

Release notes · Hotwords guide · Docs

· Core ML

Phonon-2 Core ML · phonon-coreml 1.1.0

Phonon-2 on the Apple Neural Engine

  • Phonon-2 for the Neural Engine, at the accuracy of the reference engine, 5.21 % average word error on the Open ASR Leaderboard's seven English test sets.
  • On an M5 MacBook Air an hour of speech becomes text in 6 seconds, and the GPU stays free.
  • Windows from 5 to 35 seconds share one set of weights, and every word carries its start and end time.
  • A Swift package and command-line tool for macOS 15 and iOS 18 or later, and a Python runner.

Hugging Face · GitHub · Run it

· Detta

Detta 1.0.24

Stays out of your way

  • The app you are dictating into keeps focus while Detta works from the menu bar.
  • Check for Updates in the menu, and a notice you will see when a new version is ready.
  • Report a problem from inside Detta. What you wrote is kept if it cannot be sent.
  • Performance improvements and bug fixes.

Release notes · Download

· Phonon runtime

fermion-research 0.2.7

Word timestamps

  • fermion transcribe phonon-2 clip.wav --json returns every word with its start and end in seconds, placed correctly on long audio.
  • fermion serve honours timestamp_granularities, so verbose_json carries words and segments.
  • Both Phonon-2 engines return timings, on Apple silicon and on CPUs. Transcripts are unchanged.

PyPI · Docs

· Phonon runtime

fermion-research 0.2.6

Clearer messages

  • fermion serve with a name that is not a published model stops in one line, before any network request.
  • fermion transcribe checks the audio file exists before it loads a model.
  • fermion models and the help texts show each command with its model.

PyPI · Docs

· Phonon runtime

fermion-research 0.2.5

The right engine for every CPU

  • The CPU engine reads your processor's features and picks the matching tier, from AVX2, AVX-512 and AMX on x86-64 to dotprod, i8mm and SVE on 64-bit Arm.
  • Raspberry Pi 3 and 4, other Arm boards without dotprod, and x86-64 processors without AVX2 run Phonon-2 on a baseline tier.
  • fermion describe shows the features found, the tier chosen and the binaries loaded, with --json for scripts.

PyPI · Docs

September 2026

· Phonon runtime

fermion-research 0.2.4

Every command names its model

  • Commands name their model, as in fermion transcribe phonon-2 meeting.wav, and phonon, phonon-2 and phonon-1 each run a fixed model.
  • fermion serve takes --unix-socket for an owner-only local socket and --threads for the CPU engine.
  • The CPU and CUDA containers take the model first and keep it in a cache volume.

PyPI · Docs

· Detta

Detta 1.0.21

More complete long recordings

  • Transcribe a File cuts long recordings at natural pauses, so more of every recording comes back as text.
  • Long dictations finish faster.
  • Detta opens at login on new installs.
  • If the speech engine cannot start, Detta says why and lets you copy the details for support.

Release notes · Download

· Detta

Detta 1.0.14 to 1.0.20

Your keys, your way

  • Choose any key or key combination as your hold key (1.0.14).
  • Double-tap to lock can use its own key or key combination (1.0.16).
  • Set more than one hold key (1.0.17).

Release notes · Download

· Detta

Detta 1.0

Detta, dictation for the Mac

  • Hold a key and talk. Detta types what you say where your cursor is, in any app, punctuated and capitalised.
  • Phonon-2 runs on the Mac itself, and Detta removes ums, repeated words and false starts.
  • Quick notes keep their recordings, and Transcribe a File turns wav, mp3, m4a, flac, mp4, mov and more into text.
  • Free for Apple silicon Macs, and it keeps itself up to date.

Read the post · Download

· Models

Phonon-2 · fermion-research 0.2.0

Introducing Phonon-2

  • An open English speech recognition model in a 164 MB download, averaging 5.21 % word error on the Open ASR Leaderboard's seven test sets.
  • 174 times realtime through MLX on an M5 MacBook Air, and 143 times on eight Zen 5 cores.
  • fermion-research 0.2.0 makes it the default speech model on Apple silicon and on CPUs across macOS, Linux and Windows, with CPU and CUDA container images.
  • Weights released under CC-BY-4.0.

Read the post · Model · Hugging Face

· Phonon runtime

fermion-research 0.1.23

Recordings of any length

  • Long recordings are cut into windows at pauses and transcribed in full, from the command line and from the server.
  • --json adds the start and end of each segment, and verbose_json on the server fills segments.

PyPI · Docs

August 2026

· Models

Phonon-1 · Phonon-1 Micro

Introducing Phonon-1

  • An open English speech recognition model in a 415 MB download that runs on a laptop or a datacenter GPU.
  • It transcribes an hour of audio in about two and a half minutes.
  • Phonon-1 Micro, the smallest model of the family, joins it.

Read the post · Phonon-1 · Phonon-1 Micro

July 2026

· Models

Neutrino-1 8B · 0.6B · 0.6B-Chat

Introducing the Neutrino-1 models

  • Three open-weight language models trained for a compact ternary format.
  • Neutrino-1 8B scores 72.1 on MMLU from a 2.56 GB download.
  • pip install fermion-research runs them on CUDA, Apple silicon and x86, with a local OpenAI-compatible server and a llama.cpp fork.
  • The 0.6B model also serves as the draft for speculative decoding.

Read the post · Models