File size: 8,455 Bytes
66ccdde
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
---
license: apache-2.0
language:
  - en
pipeline_tag: text-generation
library_name: pytorch
tags:
  - base-model
  - causal-lm
  - small-language-model
  - text-generation
datasets:
  - HuggingFaceTB/smollm-corpus
  - HuggingFaceTB/finemath
  - mlfoundations/dclm-baseline-1.0-parquet
  - code-search-net/code_search_net
  - codeparrot/codeparrot-clean
---

# BarunLM-35M

BarunLM-35M is a compact, decoder-only base language model built for strong
capability per parameter. With exactly **35,072,768 parameters**, it reaches
**41.01%** on a fixed, decontaminated nine-task zero-shot suite and exceeds
[LFM2.5-230M-Base](https://huggingface.co/LiquidAI/LFM2.5-230M-Base) while
using **6.55× fewer parameters**.

[Download the model](https://huggingface.co/harrrshall/BarunLM-35M) ·
[Inspect benchmark data](benchmark_results.json) · [License](LICENSE)

## Highlights

- **Parameter-efficient hybrid attention.** Three local-attention layers feed
  each full-attention layer, combining a 256-token working window with periodic
  global information exchange.
- **Selective residual routing.** A learned convex selector every four layers
  chooses between the block-group input and its transformed state, providing a
  lightweight path for preserving useful representations.
- **Stable small-model optimization.** Grouped-query attention, partial RoPE,
  QK normalization, gated attention outputs, and bounded SwiGLU are integrated
  into a single 35M-parameter design.
- **Capacity-aligned pretraining.** The model was trained on 5.70B tokens—about
  162.5 tokens per parameter—from a curated blend of educational web text,
  synthetic exposition, mathematics, general web text, and code.
- **Contamination-aware evaluation.** Reported scores exclude 1,854 samples
  identified by a correctness-blind exact 13-token scan over the complete
  training history.

## Results

All rows below were evaluated zero-shot with LM Evaluation Harness 0.4.12 on
ARC-Challenge, ARC-Easy, BoolQ, HellaSwag, LAMBADA OpenAI, OpenBookQA, PIQA,
SciQ, and WinoGrande. The reported value is the unweighted macro average after
applying the same frozen decontamination decisions to every model.

| Model | Loaded parameters | Macro accuracy | BarunLM lead |
|---|---:|---:|---:|
| **BarunLM-35M** | **35.1M** | **41.01%** | — |
| [LFM2.5-230M-Base](https://huggingface.co/LiquidAI/LFM2.5-230M-Base) | 229.7M | 39.20% | **+1.81 pp** |
| [Pythia-160M](https://huggingface.co/EleutherAI/pythia-160m-deduped) | 162.3M | 37.35% | **+3.66 pp** |
| [Stentor-30M](https://huggingface.co/StentorLabs/Stentor-30M) | 30.4M | 36.46% | **+4.55 pp** |
| [TinyStories-33M](https://huggingface.co/roneneldan/TinyStories-33M)¹ | 68.5M | 33.16% | **+7.85 pp** |
| [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m-deduped) | 70.4M | 31.71% | **+9.30 pp** |

The paired 10,000-resample bootstrap interval for the BarunLM minus LFM2.5
macro difference is **[+0.92, +2.71] percentage points**. Exact revisions,
task scores, confidence intervals, sample counts, and evidence hashes are in
[`benchmark_results.json`](benchmark_results.json).

¹ TinyStories is included as a narrow-domain diagnostic rather than a
general-purpose peer. Parameter counts are computed from the loaded models,
not inferred from repository names.

These results establish parameter efficiency on this evaluation suite; they
do not imply universal superiority across tasks or deployment settings.

## Architecture

| Component | Configuration |
|---|---|
| Parameters | 35,072,768 |
| Layers / width | 12 / 448 |
| Attention | 7 query heads, 1 key/value head |
| Attention rhythm | 3 local layers, then 1 full layer |
| Local window | 256 tokens |
| Position encoding | 50% partial RoPE |
| Feed-forward width | 1,228 |
| Residual selection | Every 4 layers |
| Vocabulary | 16,384 byte-level BPE tokens |
| Context length | 2,048 tokens |
| Embeddings | Input/output weights tied |

The design treats global attention as a periodic communication layer rather
than a cost paid at every depth. Local layers concentrate computation on nearby
structure, while the global layer propagates information across the sequence.
The residual selector then gives each four-layer group a learned preservation
path with negligible parameter overhead.

## Quick start

BarunLM uses a small native PyTorch implementation. It is a completion model,
not a chat model, and does not require `trust_remote_code`.

For the shortest path from this model repository:

```bash
hf download harrrshall/BarunLM-35M --local-dir BarunLM-35M
cd BarunLM-35M
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python generate.py \
  --prompt "The future of efficient language models is" \
  --max-new-tokens 48 \
  --temperature 0.8
```

The maintained source package and tests are available in the
[GitHub repository](https://github.com/harrrshall/barunlm-35m):

```bash
git clone https://github.com/harrrshall/barunlm-35m.git
cd barunlm-35m
python -m venv .venv
source .venv/bin/activate
pip install -e .
python examples/generate.py \
  --prompt "The future of efficient language models is" \
  --max-new-tokens 48 \
  --temperature 0.8
```

Both examples verify the three required model artifacts and select CUDA
automatically when available. For deterministic greedy decoding, pass
`--temperature 0`.

To download the release without running it:

```bash
hf download harrrshall/BarunLM-35M --local-dir BarunLM-35M
sha256sum -c BarunLM-35M/SHA256SUMS
```

## Training

BarunLM-35M was pretrained on **5,699,985,408 realized tokens** at sequence
length 2,048. The final 4B-token continuation used Muon with a peak learning
rate of `1e-4`, weight decay `0.1`, batch size 48, and 40,690 optimizer steps
on one NVIDIA H200.

The complete corpus combines the following pinned public sources:

- [FineWeb-Edu and Cosmopedia v2](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus)
- [FineMath-4+](https://huggingface.co/datasets/HuggingFaceTB/finemath)
- [DCLM-Baseline](https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet)
- [CodeSearchNet Python](https://huggingface.co/datasets/code-search-net/code_search_net)
- [CodeParrot Clean](https://huggingface.co/datasets/codeparrot/codeparrot-clean)

Documents were deduplicated across stages before admission. Source revisions
and upstream licensing notices are recorded in [`NOTICE`](NOTICE); no training
data is redistributed in this repository or the model release.

## Release integrity

| Artifact | SHA-256 |
|---|---|
| `model.safetensors` | `f2a7c88b9f2c2e3584809081407ab136795d82e30e89b730e007781c45d01447` |
| `barun_config.json` | `9b3a1d71baa95a198744d250f9629231738d942570b8685c44307fd83dd33565` |
| `tokenizer.json` | `70ded9605fccd09c2340ca7e225361eab0ae8b4dbbb0d6e26343ab5183979db6` |

The checkpoint contains 35,072,768 unique trainable parameters. All stored
floating-point tensors passed an independent finiteness and shape audit.

## Intended use and limitations

BarunLM-35M is intended for research on compact language models, controlled
text-generation experiments, education, and local prototyping.

- It is a base model, not instruction-tuned, and should be prompted as a text
  continuation model.
- Its 2,048-token context and 35M-parameter capacity limit factual recall,
  multi-step reasoning, instruction following, and long-context synthesis.
- Training and evaluation are English-centric.
- Outputs may be inaccurate, biased, unsafe, repetitive, or fabricated.
- The model has not undergone a comprehensive safety evaluation and should not
  be used for medical, legal, financial, or other high-stakes decisions.

Users should evaluate the model in their own domain and apply appropriate
guardrails before deployment.

## Acknowledgements

BarunLM builds on ideas developed across the open language-model ecosystem,
including [grouped-query attention](https://arxiv.org/abs/2305.13245),
[rotary position embeddings](https://arxiv.org/abs/2104.09864), and
[SwiGLU](https://arxiv.org/abs/2002.05202). We thank the teams behind PyTorch,
Hugging Face, LM Evaluation Harness, FineWeb-Edu, Cosmopedia, FineMath, DCLM,
CodeSearchNet, and CodeParrot for the open infrastructure and data that made
this release possible.

## License

The source code and released weights are available under the
[Apache License 2.0](LICENSE). Upstream datasets retain their own licenses and
terms as described in [`NOTICE`](NOTICE).