Fermion Research
Available now

Neutrino-1 0.6B

A 596M-parameter model in a 238 MB download, measured at 225 tok/s on the native CPU path and certified to draft for Neutrino-1 8B without changing its greedy output.

Overview

Draft and compact generation model

A 596M-parameter model in a 238 MB download, measured at 225 tok/s on the native CPU path and certified to draft for Neutrino-1 8B without changing its greedy output.

Parameters
596M
Download
238 MB
Layers
28
Context
40,960

Performance

Fast alone. Exact when drafting.

It runs independently or proposes six-token drafts for Neutrino-1 8B. The verifier emits only the prefix that matches ordinary greedy decoding.

1,177 tok/s
H100 80 GB
225 tok/s
Native CPU path
0 divergences
Across 27,648 drafted tokens
shared process and runtime binariesNeutrino-1 0.6B, 328 MBNeutrino-1 8B, 3.88 GBsix proposed tokensaccepted prefix
The drafted pair drawn to byte scale, with the residency each side costs.

Install

Run locally.

First run downloads and verifies the model. Later runs load from the local cache.

$ pip install fermion-research
$ fermion chat --model fermionresearch/Neutrino-0.6B

Model download: 238 MB.

Compare all three Neutrino-1 models.

Model family