
Neutrino-1 0.6B
A 596M-parameter model in a 238 MB download, measured at 225 tok/s on the native CPU path and certified to draft for Neutrino-1 8B without changing its greedy output.
Overview
Draft and compact generation model
A 596M-parameter model in a 238 MB download, measured at 225 tok/s on the native CPU path and certified to draft for Neutrino-1 8B without changing its greedy output.
- Parameters
- 596M
- Download
- 238 MB
- Layers
- 28
- Context
- 40,960
Performance
Fast alone. Exact when drafting.
It runs independently or proposes six-token drafts for Neutrino-1 8B. The verifier emits only the prefix that matches ordinary greedy decoding.
- 1,177 tok/s
- H100 80 GB
- 225 tok/s
- Native CPU path
- 0 divergences
- Across 27,648 drafted tokens
Install
Run locally.
First run downloads and verifies the model. Later runs load from the local cache.
$ pip install fermion-research$ fermion chat --model fermionresearch/Neutrino-0.6B
Model download: 238 MB.
Compare all three Neutrino-1 models.
Model family