---
title: "Neutrino-1 8B"
description: "An 8.19B-parameter open-weight model in a 2.56 GB download, scoring 72.1 on five-shot MMLU and 68.9 on BFCL v3, with runtimes for CUDA, Apple silicon, and x86."
canonical: "https://www.fermionresearch.com/models/neutrino-8b/"
source: "Fermion Research"
---

# Neutrino-1 8B

An 8.19B-parameter open-weight model in a 2.56 GB download, scoring 72.1 on five-shot MMLU and 68.9 on BFCL v3, with runtimes for CUDA, Apple silicon, and x86.

## Overview

Neutrino-1 8B is a local 8.19B-parameter model for chat, tool calling, and OpenAI-compatible applications. The fermion runtime serves it on CUDA, Apple silicon, and x86 from the same 3.88 GB artifact, with streaming and a persistent KV session enabled by default.

Its measured behavior is not limited to multiple-choice recall. On BFCL it composes simple, multiple, parallel, and parallel-multiple function calls, and declines an irrelevant function nearly as reliably as it selects the right one. On GSM8K, flexible and strict stated-format grading differ by 1.67 points.

Its 252 transformer projection matrices are stored in a packed ternary-family format. The weights stay bit-packed at rest and are decoded inside the matrix kernels. The same container serves each supported platform without a conversion pass.

The base is Qwen3-8B. Fermion Research continued it with ternary quantization-aware training, so the projection weights adapted to the representation used by the shipping runtime instead of being rounded once after training. The base is credited; the training method is ours.

## Architecture

A dense decoder-only transformer. Grouped-query attention holds the KV cache at a quarter of the query width. fermion 0.1.10 stores it in fp16 by default at 144 KiB per token, so an 8K-token session occupies 1.13 GiB beside the 3.88 GB of weights.

Geometry as read from the shipped container's header.

### Artifact byte composition

Only the transformer linears carry the coded format. The two embedding tensors stay int8 because their rows are read one token at a time, not multiplied against the full activation stream, and the normalization weights are too small to be worth coding. A third of the file is vocabulary.

Byte budget of the 3,875,404,812-byte container, by tensor class.

### Zero-state occupancy

Across the 6.95B coded weights, 62.63% are zero. Occupancy varies with depth: the gate and down feed-forward projections reach 70 to 72% zeros in layers 1 through 3, while all four attention projections stay within about one point of 62% from layer 0 to layer 35.

Share of coded weights at zero, per projection, across all 36 layers.

## Format

The 2.56 GB download is a losslessly coded transport of the 3.88 GB serving artifact. Expansion is bit-exact, and each supported runtime loads the same stored model weights.

The weights, manifest, GGUF pack, MLX implementation, and llama.cpp fork are public under their stated open-source licenses. The optimized fermion kernels are closed. The manifest binds tensor layout, sizes, and hashes so the loader can verify the artifact before execution.

### Distribution surfaces

#### pip engine

24.9 tok/s on the native CPU path, 9 threads

Installing fermion-research provides the loader and the platform-matching native binary. CPU runtimes are available for macOS arm64 and Linux x86-64, with a bit-exact torch reference implementation.

#### GGUF pack + llama.cpp fork

30.7 tok/s on an NVIDIA L4, 4.68 GiB at 4k context

The container converted to GGUF with the FV5 tensor type. It requires the public Fermion Research llama.cpp fork: stock llama.cpp, Ollama, and LM Studio do not load this tensor type. The fermion-fv5 branch runs on CPU/CUDA; fermion-fv5-metal carries the Apple-GPU kernels.

#### MLX pack

33.7 tok/s on the optimized MLX path

A Python-native Apple-silicon implementation with custom Metal kernels. The artifact is memory-mapped and the packed planes are decoded inside the GEMV kernels.

## Evaluation

The release battery runs on the published artifact with thinking disabled. Each row states the shot count, grading mode, and item count.

Measured on standard public harnesses, July 2026. Methodology on the model card.

The BFCL result covers thirteen suites rather than a single function-call template. The model handles simple, multiple, parallel, and parallel-multiple calls against public API schemas with varied argument names and nested fields. Its relevance suites also test the opposite behavior: declining an offered function when none applies.

GSM8K is reported twice to expose output discipline. Flexible extraction accepts the last number and scores 53.4; stated-format extraction accepts only the requested answer position and scores 51.73. The 1.67-point gap separates formatting failures from arithmetic failures instead of folding both into one number.

## Throughput

These measurements use single-stream decode with one active generation. Every row uses identical model weights through a different runtime backend.

Single-stream decode rates by platform and surface, July 2026.

The optimized MLX path reaches 33.7 tokens per second and the native CPU path reaches 24.9. The same weights reach 396 tokens per second on one H100 before drafting. These are different kernels over the same 3.88 GB model, not separately converted checkpoints.

## Speculative decoding

Neutrino-1 0.6B drafts a run of tokens, the 8B scores the whole run in one forward pass, and the agreeing prefix is kept. A draft token is accepted only when it equals the 8B's own argmax, so the output stream is the plain greedy stream: on the shipping configuration, 27,648 consecutive tokens matched with zero divergences.

Speculative-decoding speed depends on how many proposed tokens the 8B accepts, so results are reported per prompt class against the 396 tok/s plain rate. On counting prompts the 8B accepts the full six-token draft on every pass, about seven tokens emitted per 8B forward; on factual prompts acceptance holds at 96.5%. A dynamic controller sizes each draft from recent acceptance. Every measured class exceeds the plain decode rate.

H100, three-round median per prompt class, over 396 tok/s plain decode.

### Memory requirements

Both models use the same format and runtime binaries. The draft loads in the verifier process without a conversion step. The draft artifact uses 328 MB beside the 8B's 3.88 GB, and 4k tokens of shared context allocate exactly one gibibyte of cache across the pair.

The drafted pair drawn to byte scale, with the residency each side costs.

### Drafting through MLX

Both models load into one MLX process and peak at 4.3 GiB together, with the draft accounting for 0.53 GiB. The exactness gate returns 6 of 6 prompts token-identical with drafting on and off, and on factual prompts the drafted rate is 25.71 tok/s against 22.00 plain at an acceptance of 0.744.

## Run it

### Install and run Neutrino-1 8B.

```text
$ pip install fermion-research
$ fermion chat --model fermionresearch/Neutrino-8B
```

On the first 8B run, fermion downloads a 2.56 GB coded transport with a progress bar, verifies its SHA-256, and unpacks it locally. Later runs load from cache.

- Long context with bounded KV memory<div><span class="select-none text-paper/30">$ </span>fermion chat --model fermionresearch/Neutrino-8B --kv-dtype int8 --yarn-factor 4</div> The native window is 40,960 tokens. --yarn-factor 4 extends the loaded window toward 160K without changing the artifact; --kv-dtype int8 holds 131,072 tokens in about 9.3 GiB. fp16 is the default, fp32 preserves byte-identical 0.1.9 output, and --yarn-factor 1.0 is exact identity.
- A local OpenAI-compatible server<div><span class="select-none text-paper/30">$ </span>fermion serve --model fermionresearch/Neutrino-8B</div> Serves at http://127.0.0.1:8000/v1. Point any OpenAI client at that base URL; the key can be any string. Tool calling, streaming, and a persistent KV session are on by default.
- Plain Transformers<div><span class="select-none text-paper/30">$ </span>hf download fermionresearch/Neutrino-8B --local-dir Neutrino-8B \</div><div> --exclude "gguf/*" --exclude "*.tv4z"</div> <div>import fermion</div><div>from transformers import AutoModelForCausalLM, AutoTokenizer</div><div>model = AutoModelForCausalLM.from_pretrained("Neutrino-8B")</div><div>tokenizer = AutoTokenizer.from_pretrained("Neutrino-8B")</div> Import fermion first to register the trtc_v4 model type. Load from the downloaded local directory; passing the Hub repository id directly to from_pretrained does not work.
- GGUF, through our llama.cpp fork only<div><span class="select-none text-paper/30">$ </span>git clone https://github.com/fermionresearch/llama.cpp && cd llama.cpp</div><div><span class="select-none text-paper/30">$ </span>git checkout fermion-fv5</div><div><span class="select-none text-paper/30">$ </span>cmake -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DGGML_METAL=OFF</div><div><span class="select-none text-paper/30">$ </span>cmake --build build -j --target llama-completion</div> The pack uses the FV5 tensor type: stock llama.cpp, Ollama, and LM Studio cannot load it. The fermion-fv5 branch is the CPU/CUDA build and requires -DGGML_METAL=OFF. For Apple-GPU kernels, use the live fermion-fv5-metal branch and omit that flag.
- MLX, on Apple silicon<div><span class="select-none text-paper/30">$ </span>cd Neutrino-8B/mlx</div><div><span class="select-none text-paper/30">$ </span>pip install -r requirements.txt</div><div><span class="select-none text-paper/30">$ </span>python -m fermion_mlx --model ../neutrino-8b_v4.bin --mode chat --tokenizer .</div> Run this exact sequence after downloading the repository. The working directory and relative model path are required.

#### The full CLI

Every command takes --model with a Hub id or local path.

[Fermion documentation](https://www.fermionresearch.com/docs/)

Get the models

Weights, native binaries, GGUF pack, MLX pack.

```text
$ fermion chat --model fermionresearch/Neutrino-8B
```

Run Neutrino-1 8B

```text
$ pip install fermion-research
$ fermion chat --model fermionresearch/Neutrino-8B
```

First run downloads a 2.56 GB coded transport, shows progress, verifies its SHA-256, and unpacks locally. Later runs load from cache.

The draft model, and a fast model on its own.

```text
$ fermion chat --model fermionresearch/Neutrino-0.6B
```

Run Neutrino-1 0.6B

```text
$ pip install fermion-research
$ fermion chat --model fermionresearch/Neutrino-0.6B
```

First run downloads the 0.24 GB coded transport, shows progress, verifies its SHA-256, and unpacks locally. Later runs load from cache.

The conversational small model.

```text
$ fermion chat --model fermionresearch/Neutrino-0.6B-Chat
```

Run Neutrino-1 0.6B-Chat

```text
$ pip install fermion-research
$ fermion chat --model fermionresearch/Neutrino-0.6B-Chat
```

First run downloads about 0.33 GB of raw model files with progress and verifies them. This model does not use the coded transport. Later runs load from cache.

The engine, the loader, and the chat front end.

```text
$ pip install fermion-research
```

Adds the format to the llama.cpp toolchain.

```text
$ git clone https://github.com/fermionresearch/llama.cpp && git checkout fermion-fv5
```

## Availability

Neutrino-1 8B ships in a public repository with the model artifact, manifest, GGUF pack, MLX implementation, and installation instructions. The optimized fermion kernels distributed by the pip package are closed.

### License

Open weights under the Apache License 2.0. Commercial use, modification, fine-tuning, and redistribution are permitted, with no access request and no acceptance form. The model is a derivative of Qwen3-8B, itself Apache-2.0. The pip package is Apache-2.0 too; the llama.cpp fork is MIT, following upstream llama.cpp.

### Citation

```text
@misc{fermionresearch2026neutrino,
  title  = {Neutrino-1 8B},
  author = {{Fermion Research}},
  year   = {2026},
  url    = {https://www.fermionresearch.com/models/neutrino-8b/}
}
```

## Related research

![An amber cellular orbit gathering around a compact central form](https://www.fermionresearch.com/images/backgrounds/amber-orbit.webp)

AnnouncementJuly 27, 2026

### Introducing the Neutrino-1 models

Three open-weight models trained for a compact ternary format and released with CUDA, Apple-silicon, and x86 runtimes.

![A honey-toned filament carrying energy across a tactile surface](https://www.fermionresearch.com/images/backgrounds/honey-filament.webp)

ResearchJuly 27, 2026

### The Neutrino Engine

An A100 can move the weights for one measured decoder step in 104 microseconds. The step takes 694. The Neutrino Engine is built around the gap.

Source: [https://www.fermionresearch.com/models/neutrino-8b/](https://www.fermionresearch.com/models/neutrino-8b/)
