BarunLM-35M / README.md
harrrshall's picture
BarunLM-35M: 35M-parameter base language model trained on 5.7B tokens
66ccdde verified
|
Raw
History Blame Contribute Delete
8.46 kB
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: pytorch
tags:
- base-model
- causal-lm
- small-language-model
- text-generation
datasets:
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/finemath
- mlfoundations/dclm-baseline-1.0-parquet
- code-search-net/code_search_net
- codeparrot/codeparrot-clean
---
# BarunLM-35M
BarunLM-35M is a compact, decoder-only base language model built for strong
capability per parameter. With exactly **35,072,768 parameters**, it reaches
**41.01%** on a fixed, decontaminated nine-task zero-shot suite and exceeds
[LFM2.5-230M-Base](https://huggingface.co/LiquidAI/LFM2.5-230M-Base) while
using **6.55× fewer parameters**.
[Download the model](https://huggingface.co/harrrshall/BarunLM-35M) ·
[Inspect benchmark data](benchmark_results.json) · [License](LICENSE)
## Highlights
- **Parameter-efficient hybrid attention.** Three local-attention layers feed
each full-attention layer, combining a 256-token working window with periodic
global information exchange.
- **Selective residual routing.** A learned convex selector every four layers
chooses between the block-group input and its transformed state, providing a
lightweight path for preserving useful representations.
- **Stable small-model optimization.** Grouped-query attention, partial RoPE,
QK normalization, gated attention outputs, and bounded SwiGLU are integrated
into a single 35M-parameter design.
- **Capacity-aligned pretraining.** The model was trained on 5.70B tokens—about
162.5 tokens per parameter—from a curated blend of educational web text,
synthetic exposition, mathematics, general web text, and code.
- **Contamination-aware evaluation.** Reported scores exclude 1,854 samples
identified by a correctness-blind exact 13-token scan over the complete
training history.
## Results
All rows below were evaluated zero-shot with LM Evaluation Harness 0.4.12 on
ARC-Challenge, ARC-Easy, BoolQ, HellaSwag, LAMBADA OpenAI, OpenBookQA, PIQA,
SciQ, and WinoGrande. The reported value is the unweighted macro average after
applying the same frozen decontamination decisions to every model.
| Model | Loaded parameters | Macro accuracy | BarunLM lead |
|---|---:|---:|---:|
| **BarunLM-35M** | **35.1M** | **41.01%** | — |
| [LFM2.5-230M-Base](https://huggingface.co/LiquidAI/LFM2.5-230M-Base) | 229.7M | 39.20% | **+1.81 pp** |
| [Pythia-160M](https://huggingface.co/EleutherAI/pythia-160m-deduped) | 162.3M | 37.35% | **+3.66 pp** |
| [Stentor-30M](https://huggingface.co/StentorLabs/Stentor-30M) | 30.4M | 36.46% | **+4.55 pp** |
| [TinyStories-33M](https://huggingface.co/roneneldan/TinyStories-33M)¹ | 68.5M | 33.16% | **+7.85 pp** |
| [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m-deduped) | 70.4M | 31.71% | **+9.30 pp** |
The paired 10,000-resample bootstrap interval for the BarunLM minus LFM2.5
macro difference is **[+0.92, +2.71] percentage points**. Exact revisions,
task scores, confidence intervals, sample counts, and evidence hashes are in
[`benchmark_results.json`](benchmark_results.json).
¹ TinyStories is included as a narrow-domain diagnostic rather than a
general-purpose peer. Parameter counts are computed from the loaded models,
not inferred from repository names.
These results establish parameter efficiency on this evaluation suite; they
do not imply universal superiority across tasks or deployment settings.
## Architecture
| Component | Configuration |
|---|---|
| Parameters | 35,072,768 |
| Layers / width | 12 / 448 |
| Attention | 7 query heads, 1 key/value head |
| Attention rhythm | 3 local layers, then 1 full layer |
| Local window | 256 tokens |
| Position encoding | 50% partial RoPE |
| Feed-forward width | 1,228 |
| Residual selection | Every 4 layers |
| Vocabulary | 16,384 byte-level BPE tokens |
| Context length | 2,048 tokens |
| Embeddings | Input/output weights tied |
The design treats global attention as a periodic communication layer rather
than a cost paid at every depth. Local layers concentrate computation on nearby
structure, while the global layer propagates information across the sequence.
The residual selector then gives each four-layer group a learned preservation
path with negligible parameter overhead.
## Quick start
BarunLM uses a small native PyTorch implementation. It is a completion model,
not a chat model, and does not require `trust_remote_code`.
For the shortest path from this model repository:
```bash
hf download harrrshall/BarunLM-35M --local-dir BarunLM-35M
cd BarunLM-35M
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python generate.py \
--prompt "The future of efficient language models is" \
--max-new-tokens 48 \
--temperature 0.8
```
The maintained source package and tests are available in the
[GitHub repository](https://github.com/harrrshall/barunlm-35m):
```bash
git clone https://github.com/harrrshall/barunlm-35m.git
cd barunlm-35m
python -m venv .venv
source .venv/bin/activate
pip install -e .
python examples/generate.py \
--prompt "The future of efficient language models is" \
--max-new-tokens 48 \
--temperature 0.8
```
Both examples verify the three required model artifacts and select CUDA
automatically when available. For deterministic greedy decoding, pass
`--temperature 0`.
To download the release without running it:
```bash
hf download harrrshall/BarunLM-35M --local-dir BarunLM-35M
sha256sum -c BarunLM-35M/SHA256SUMS
```
## Training
BarunLM-35M was pretrained on **5,699,985,408 realized tokens** at sequence
length 2,048. The final 4B-token continuation used Muon with a peak learning
rate of `1e-4`, weight decay `0.1`, batch size 48, and 40,690 optimizer steps
on one NVIDIA H200.
The complete corpus combines the following pinned public sources:
- [FineWeb-Edu and Cosmopedia v2](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus)
- [FineMath-4+](https://huggingface.co/datasets/HuggingFaceTB/finemath)
- [DCLM-Baseline](https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet)
- [CodeSearchNet Python](https://huggingface.co/datasets/code-search-net/code_search_net)
- [CodeParrot Clean](https://huggingface.co/datasets/codeparrot/codeparrot-clean)
Documents were deduplicated across stages before admission. Source revisions
and upstream licensing notices are recorded in [`NOTICE`](NOTICE); no training
data is redistributed in this repository or the model release.
## Release integrity
| Artifact | SHA-256 |
|---|---|
| `model.safetensors` | `f2a7c88b9f2c2e3584809081407ab136795d82e30e89b730e007781c45d01447` |
| `barun_config.json` | `9b3a1d71baa95a198744d250f9629231738d942570b8685c44307fd83dd33565` |
| `tokenizer.json` | `70ded9605fccd09c2340ca7e225361eab0ae8b4dbbb0d6e26343ab5183979db6` |
The checkpoint contains 35,072,768 unique trainable parameters. All stored
floating-point tensors passed an independent finiteness and shape audit.
## Intended use and limitations
BarunLM-35M is intended for research on compact language models, controlled
text-generation experiments, education, and local prototyping.
- It is a base model, not instruction-tuned, and should be prompted as a text
continuation model.
- Its 2,048-token context and 35M-parameter capacity limit factual recall,
multi-step reasoning, instruction following, and long-context synthesis.
- Training and evaluation are English-centric.
- Outputs may be inaccurate, biased, unsafe, repetitive, or fabricated.
- The model has not undergone a comprehensive safety evaluation and should not
be used for medical, legal, financial, or other high-stakes decisions.
Users should evaluate the model in their own domain and apply appropriate
guardrails before deployment.
## Acknowledgements
BarunLM builds on ideas developed across the open language-model ecosystem,
including [grouped-query attention](https://arxiv.org/abs/2305.13245),
[rotary position embeddings](https://arxiv.org/abs/2104.09864), and
[SwiGLU](https://arxiv.org/abs/2002.05202). We thank the teams behind PyTorch,
Hugging Face, LM Evaluation Harness, FineWeb-Edu, Cosmopedia, FineMath, DCLM,
CodeSearchNet, and CodeParrot for the open infrastructure and data that made
this release possible.
## License
The source code and released weights are available under the
[Apache License 2.0](LICENSE). Upstream datasets retain their own licenses and
terms as described in [`NOTICE`](NOTICE).