---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- mixture-of-experts
- moe
- small-language-model
- edge-inference
- int4
- from-scratch
- distributed-training
base_model: vovaRL/NanoColibri-Instruct
model-index:
- name: Colibri-Nano
results:
- task:
type: text-generation
dataset:
name: LAMBADA
type: lambada
metrics:
- type: accuracy
value: 26.3
name: zero-shot accuracy
- task:
type: text-generation
dataset:
name: PIQA
type: piqa
metrics:
- type: accuracy
value: 62.7
name: zero-shot accuracy
- task:
type: text-generation
dataset:
name: WinoGrande
type: winogrande
metrics:
- type: accuracy
value: 49.2
name: zero-shot accuracy
- task:
type: text-generation
dataset:
name: ARC-easy
type: arc_easy
metrics:
- type: accuracy
value: 43.4
name: zero-shot accuracy
- task:
type: text-generation
dataset:
name: ARC-challenge
type: arc_challenge
metrics:
- type: accuracy
value: 22.8
name: zero-shot accuracy
- task:
type: text-generation
dataset:
name: OpenBookQA
type: openbookqa
metrics:
- type: accuracy
value: 22.0
name: zero-shot accuracy
- task:
type: text-generation
dataset:
name: HellaSwag
type: hellaswag
metrics:
- type: accuracy
value: 31.1
name: zero-shot accuracy
---
# Colibri-Nano
*A 2.7B-parameter Mixture-of-Experts that runs on a 4 GB mini-PC — trained from
scratch by volunteers passing a GPU baton, for about the price of a game console.*
| | |
|---|---|
| **Parameters** | **2.7B** total · 0.34B active per token (~12%) |
| **Training tokens** | **~5.4B** (pretrain + chat SFT) |
| **Container** | **~1.2 GB** int4, streams experts from disk |
| **Compute** | **1 GPU** at a time · ~50–100 H100-hours for the whole run |
| **Cost** | **~$180–260** for the whole run |
| **License** | **Apache-2.0** — code, checkpoints, and a ledger of who trained what |
**Architecture.** 24 layers (1 dense + 23 MoE) · 64 experts, top-2 + a 2048-wide
always-resident shared expert · GQA 4× · per-head QK-norm · sigmoid+bias router
(DeepSeek-V3-style aux-loss-free balancing) · SmolLM2 tokenizer, 49 152 vocab ·
4 096 context.
---
## How it was built
**1 — Relay pretraining.** Pretraining is sequential, so contributors take turns:
claim a lease on a HuggingFace repo (a real compare-and-swap, not a convention),
pull the checkpoint, train a leg, push, hand the baton on. 20 000 updates ≈ 5.2B
tokens of FineWeb-Edu, then a 1 500-update chat SFT.
**2 — Engine-native.** The architecture was chosen *for the serving engine*
(hy3.c): int4 experts stream from disk under a fixed RAM budget, while the dense
backbone and the shared expert stay resident. No conversion shims.
**3 — The bet paid off.** At a matched token budget, the MoE beats dense
references that spend *more* compute per token.
---
## Evaluation 1 — beats dense at the same token budget
Zero-shot accuracy (%). Colibri-Nano (2.7B total / 0.34B active) against two
token-matched dense models.
| Task | Colibri-Nano
0.34B act / 5.4B tok | Pythia-410M
@ 6.3B tok (dense 405M) | Cerebras-GPT-256M
@ 5.1B tok (dense) | Random |
|---|---|---|---|---|
| LAMBADA | 26.3 | 26.3 | **29.3** | 0 |
| PIQA | **62.7** | 59.8 | 61.3 | 50 |
| WinoGrande † | 49.2 | 50.9 | 51.1 | 50 |
| ARC-easy | **43.4** | 41.3 | 41.0 | 25 |
| ARC-challenge | **22.8** | 18.8 | 17.0 | 25 |
| OpenBookQA | **22.0** | 15.6 | 15.8 | 25 |
| HellaSwag | **31.1** | 27.0 | 27.4 | 25 |
Identical prompts and scoring for all three columns (in-repo lm-eval-compatible
harness; Cerebras row from [arXiv:2304.03208](https://arxiv.org/abs/2304.03208)
Table 8, same protocol). **Colibri-Nano wins 5 of 7 against both token-matched
dense models — ties LAMBADA, WinoGrande within noise — with fewer active
parameters per token.**
> † WinoGrande differences are within noise (±1.4 pts).
>
> ‡ The Pythia checkpoint (`step3000`) is early in its LR schedule, which
> handicaps it somewhat — the fully-annealed Cerebras-GPT-256M column is the
> conservative comparison, and Nano wins that one too.
## Evaluation 2 — quality vs training tokens
Average zero-shot accuracy over the 7 tasks above. Random-guess average on this
task set: 28.6.
| Model | Tokens (B) | Avg accuracy | Kind |
|---|---|---|---|
| Cerebras-GPT-111M | 2.2 | 31.5 | dense |
| Cerebras-GPT-256M | 5.1 | 34.7 | dense |
| **Colibri-Nano** | **5.4** | **36.8** | **MoE (ours)** |
| Pythia-410M @ step3000 ‡ | 6.3 | 34.2 | dense |
| Cerebras-GPT-590M | 11.8 | 37.1 | dense |
| OPT-125M | 180 | 37.1 | dense |
| OPT-350M | 180 | 39.5 | dense |
| Pythia-160M | 300 | 37.4 | dense |
| Pythia-410M | 300 | 41.9 | dense |
At 5.4B tokens, Colibri-Nano matches OPT-125M (33× the tokens) and
Cerebras-GPT-590M (~4× the training compute), and clearly beats everything at its
own budget. Models further down the table just ate more data.
## Evaluation 3 — the efficiency frontier (context, not a target)
> **Different metric family.** These are `acc_norm` numbers under the SmolLM2-card
> protocol and are **not comparable** with the accuracies in Evaluation 1–2.
| Model | Active (M) | Total (M) | Tokens (B) | HellaSwag | ARC avg | PIQA | WinoGrande | OpenBookQA |
|---|---|---|---|---|---|---|---|---|
| **Colibri-Nano** | **341** | **2700** | **5.4** | 35.6 | 32.8 | 62.7 | 49.2 | 29.6 |
| SmolLM2-135M | 135 | 135 | 2000 | 42.1 | 43.9 | 68.4 | 51.3 | 34.6 |
| SmolLM2-360M | 360 | 360 | 4000 | 54.5 | 53.0 | 71.7 | 52.5 | 37.4 |
| Qwen2.5-0.5B | 494 | 494 | 18000 | 51.2 | 45.4 | 69.9 | 54.1 | 37.4 |
The trillions-of-tokens ceiling for this size class — context, not a target. The
gap is data volume (370–3300× more tokens), not architecture.
---
## Reproducing every number on this card
```bash
python -m scripts.bench_nano --indir checkpoints/nano_chat
```
Every number here regenerates from the repo (`docs/BENCH.md`, `docs/RESULTS.md`).
**Train it yourself:** the one-line `setup.sh` + `./train.sh` relay (see README in
the training repo).
## Limitations
- Undertrained by design of budget, not ambition: at ~5.4B tokens the model sits
370–3300× below the token counts of the frontier models in its size class, and
Evaluation 3 shows the resulting gap.
- Speed benchmarks across target devices (tokens/s vs engine memory budget) are
not yet published here; disk-pressure benchmarks (memory-capped cgroups, larger
containers) are the successor project's first milestone.
- Small-model failure modes apply: factual errors, repetition, and weak
multi-step reasoning.
## What's next
A 7B chat+code pilot with RL post-training, then a 24–28B flagship whose
container is deliberately bigger than a 16 GB laptop's memory — streaming experts
from NVMe at 20+ tok/s (plan: `docs/NEXT_MODEL.md`).
---
License Apache-2.0 · HF repo [`vovaRL/NanoColibri-Instruct`](https://huggingface.co/vovaRL/NanoColibri-Instruct) · engine: hy3.c