--- license: apache-2.0 datasets: - HuggingFaceFW/finephrase - HuggingFaceFW/fineweb-edu - mlfoundations/dclm-baseline-1.0 - nampdn-ai/tiny-strange-textbooks - ByteDance-Seed/mga-fineweb-edu - nvidia/OpenMathInstruct-2 - AxiomicLabs/NPset-2-Python-Edu language: - en tags: - slm - custom-architecture - custom-tokenizer - byte-level - small - fromzero - negative-v1.0 --- # Negative-v1.0 Negative-v1.0 is a 67K-parameter small language model (SLM) featuring a custom architecture inspired by Needle2. Trained entirely on CPU over 600M tokens, Negative-v1.0 utilizes a byte-level tokenizer with 4 special tokens (``, ``, ``, ``), resulting in a compact vocabulary size of 260. ## Architecture Negative-v1.0 employs a compact, parameter-efficient architecture incorporating Engram memory, Hadamard FFNs (parameter-free) with SwiGLU intervals, and an 8-stream topology powered by mHC. - Vocab Size: `260` - Max Position Embeddings: `96` - Hidden Size: `32` - Intermediate Size (for SwiGLU): `64` - Total Number of Layers: `9` - Hadamard Layers: `7` - SwiGLU Layers: `2` - Number of Heads: `4` - Number of KV Heads: `2` - Dimensions Per Head: `8` - Use Per-Head Gating: `false` - Use XSA: `false` - Number of mHC Streams: `8` - Use Engram: `true` - Number of Engram Entries: `196` - Engram Orders: `(4, 8)` ## Training Dataset Negative was trained on 600 million tokens of a diverse dataset mixture comprising general web text, educational content, synthetic data, normalized code, and mathematics. | Dataset | Share | | :--- | :---: | | **FineWeb-Edu** | 36.0% | | **DCLM Baseline 1.0** | 22.9% | | **FinePhrase** | 13.4% | | **MGA FineWeb-Edu** | 10.3% | | **Tiny Strange Textbooks** | 8.2% | | **OpenMathInstruct-2** | 7.6% | | **NPset-2 Python-Edu** | 1.6% | ## Benchmark Results We benchmaked Negative-v1.0 on five tasks: Arc_Easy, Arc_Challenge, HellaSwag, PiQA, and ArithMark-3.0. | Task | Metric | Score | | :--- | :--- | :---: | | **ARC Challenge** | `acc_norm` | 22.95% | | **ARC Easy** | `acc_norm` | 27.65% | | **HellaSwag** | `acc_norm` | 25.94% | | **PIQA** | `acc_norm` | 49.62% | | **ArithMark-3.0** | `acc_norm` | 31.50% | Despite its compact size, Negative exhibits surprisingly competitive performance on knowledge-intensive and mathematical benchmarks within its parameter class. ## License Apache 2.0. ## Citation ``` @misc{negative-v1.0, title = {Negative-v1.0}, organization = {FromZero}, authors = {Paul Courneya}, year = {2026}, url = {https://huggingface.co/fromziro/Negative-v1.0] } ```