---
title: "Changelog"
description: "Recent changes to the public fermion package, verified against the shipped package source."
canonical: "https://www.fermionresearch.com/docs/changelog/"
source: "Fermion Research"
---

Docs / Reference

# Changelog

Recent changes to the public fermion package, verified against the shipped package source.

## 0.1.24

- Fixes a startup failure on systems with a broken torchaudio install; the runtime now starts and tells you how to remove the broken package.
- Requirements: on Linux x86-64 the CPU speech engine needs glibc 2.34 or newer (Ubuntu 22.04, Debian 12, RHEL 9 and later) and about 7 GB of free RAM. macOS and Windows requirements are unchanged.

## 0.1.23

- Long recordings are transcribed in full. `fermion transcribe` and `POST /v1/audio/transcriptions` decoded a whole file in one pass with a token budget sized for a single utterance, so anything over about two minutes came back truncated and anything over about three degenerated into a repeated letter. Audio longer than 35 s is now cut into 25 to 35 s windows at pauses and each window is decoded on its own, on both the MLX and the CPU engine. Audio up to 35 s takes the unchanged single-pass path, so short-utterance transcripts are byte-identical to 0.1.22.
- `--json` adds `wall_seconds` next to the decode-only `decode_seconds`, per-segment start and end times, `segment_count`, and `truncated`. `--verbose` prints the same, labelled. `verbose_json` on the server fills `segments` and reports `truncated`.
- One always-on line at engine start names the decode backend (mlx or cpu) and the model, and says on the first run in a new environment that the run compiles once and later runs are fast.
- A Python running under Rosetta on an Apple-silicon Mac is told so, with the fix (an arm64 Python), instead of being refused as an Intel Mac.

## 0.1.22

- The CPU speech engine now also runs on ARM Linux (aarch64): `fermion transcribe` and a speech `fermion serve` work on 64-bit ARM Linux machines with the same install line as x86-64 Linux. The models are the same verified downloads, and transcription stays deterministic.
- The CPU engine's first run no longer prints a spurious advisory line from the model backend's import.
- `fermion models`: one transcribe example in the usage footer.

## 0.1.21

- Speech on plain CPU: `fermion transcribe` and a speech `fermion serve` now run on x86-64 Linux and Windows, and on a Mac without the MLX runtime. `pip install fermion-research torch safetensors soundfile scipy zstandard` is the whole setup; the models are the same verified downloads. Apple silicon keeps the MLX engine as its default, and `FERMION_DEVICE=cpu` selects the CPU engine explicitly.
- Transcription decodes deterministically on every engine; transcripts across engines are same-accuracy, not word-identical.
- `fermion models` reports which speech engine this machine runs.

## 0.1.20

- Package documentation copy edits; no behavior change.
- `fermion models`: shorter speech model summaries.

## 0.1.19

- Package documentation refreshed for the Phonon-1 release.

## 0.1.18

- Streaming speech over WebSocket: a speech `fermion serve` also mounts `GET /v1/audio/stream`, taking one JSON config frame and then raw mono audio, and answering with partial hypotheses, final segments, and the full transcript on close. `--api-key` applies, as a Bearer header or `?api_key=` for browser clients.
- `fermion serve --model` pointed at a local speech model directory now serves speech; it previously failed down the chat path.

## 0.1.17

- `fermion listen`: live dictation from the microphone. Partial hypotheses update one terminal line, a pause finalizes the segment, and Ctrl-C prints the full transcript to standard output.
- `fermion listen --wav FILE` streams a file through the same path paced to real time.
- `fermion --version` reads the installed package metadata instead of a maintained literal.

## 0.1.16

- Fresh-install fix: `mlx-lm` added to the checked speech runtime and the documented install line. Without it, a first run crashed after the model download instead of printing the install hint.

## 0.1.15

- Speech support: `fermion transcribe` and speech serving with `/v1/audio/transcriptions`.
- Three speech models, downloaded on first use and selected with `--model`: `FermionResearch/Phonon-1` (the default), `Phonon-1-Big`, and `Phonon-1-Micro`.

## 0.1.14

- Native `--draft` routing with no PyTorch model loaded on the optimized path.
- `--draft auto` resolves to Neutrino-0.6B.
- `fermion verify` defaults to float32 so PASS matches the identity guarantee.
- Clearer reduced-precision near-tie messages.

## 0.1.13

- More tolerant multi-call envelope parser and private think-block guard.
- Serve-side stuck detector and tool-request sampling profiles.
- Warnings for unsupported OpenAI sampling fields.
- TF32 disabled during CUDA float32 verification.

## 0.1.12

- Int8 KV column in `inspect`.
- Clearer YaRN scope and per-model tool-calling guidance.

## 0.1.11 and earlier

- 0.1.11 added free-space preflight, named disk errors, `FERMION_CACHE_DIR`, and post-expansion transport cleanup.
- 0.1.10 added KV dtypes, experimental YaRN controls, and session memory guards.
- 0.1.9 added coded Apple Silicon downloads and live prose streaming for tool requests.
- 0.1.8 added unique tool-name aliasing, multiple KV conversations, and session statistics.
- 0.1.7 added OpenAI tools, persistent KV sessions, context clamps, and longer chat output.

Source: [https://www.fermionresearch.com/docs/changelog/](https://www.fermionresearch.com/docs/changelog/)
