Text Generation
MLX
GGUF
1b
1b-active
5b
7b
allenai
android
apple-silicon
attested
calibration-aware-pruning
chain-of-custody
chinese
consumer-gpu
cryptographically-verified
edge-inference
embedded
english
expert-pruning
forge-alloy
fully-open
general
general-purpose
ggml
iphone
llama-cpp
lm-studio
local-inference
macbook
mixture-of-experts
mobile
Mixture of Experts
multilingual
ollama
olmoe
on-device
q5-k-m
q5_k_m
quantized
raspberry-pi
reproducible
sparse-moe
versatile
conversational
Instructions to use continuum-ai/olmoe-1b-7b-compacted-5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use continuum-ai/olmoe-1b-7b-compacted-5b with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("continuum-ai/olmoe-1b-7b-compacted-5b") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use continuum-ai/olmoe-1b-7b-compacted-5b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M # Run inference directly in the terminal: llama cli -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M # Run inference directly in the terminal: llama cli -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M # Run inference directly in the terminal: ./llama-cli -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
Use Docker
docker model run hf.co/continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
- LM Studio
- Jan
- vLLM
How to use continuum-ai/olmoe-1b-7b-compacted-5b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "continuum-ai/olmoe-1b-7b-compacted-5b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "continuum-ai/olmoe-1b-7b-compacted-5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
- Ollama
How to use continuum-ai/olmoe-1b-7b-compacted-5b with Ollama:
ollama run hf.co/continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
- Unsloth Studio
How to use continuum-ai/olmoe-1b-7b-compacted-5b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for continuum-ai/olmoe-1b-7b-compacted-5b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for continuum-ai/olmoe-1b-7b-compacted-5b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for continuum-ai/olmoe-1b-7b-compacted-5b to start chatting
- Atomic Chat new
- MLX LM
How to use continuum-ai/olmoe-1b-7b-compacted-5b with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "continuum-ai/olmoe-1b-7b-compacted-5b"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "continuum-ai/olmoe-1b-7b-compacted-5b" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "continuum-ai/olmoe-1b-7b-compacted-5b", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use continuum-ai/olmoe-1b-7b-compacted-5b with Docker Model Runner:
docker model run hf.co/continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
- Lemonade
How to use continuum-ai/olmoe-1b-7b-compacted-5b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull continuum-ai/olmoe-1b-7b-compacted-5b:Q5_K_M
Run and chat with the model
lemonade run user.olmoe-1b-7b-compacted-5b-Q5_K_M
List all available models
lemonade list
File size: 7,595 Bytes
2c9f32a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 | ---
tags:
- 1b
- 1b-active
- 5b
- 7b
- allenai
- android
- apple-silicon
- attested
- calibration-aware-pruning
- chain-of-custody
- chinese
- consumer-gpu
- cryptographically-verified
- edge-inference
- embedded
- english
- expert-pruning
- forge-alloy
- fully-open
- general
- general-purpose
- ggml
- gguf
- iphone
- llama-cpp
- lm-studio
- local-inference
- macbook
- mixture-of-experts
- mlx
- mobile
- moe
- multilingual
- ollama
- olmoe
- on-device
- q5-k-m
- q5_k_m
- quantized
- raspberry-pi
- reproducible
- sparse-moe
- text-generation
- versatile
base_model: allenai/OLMoE-1B-7B-0924-Instruct
pipeline_tag: text-generation
license: apache-2.0
---
# 25% Experts Pruned, 36.0 HUMANEVAL (base 40.9)
**OLMoE-1B-7B-0924-Instruct** compacted via per-layer-normalized MoE expert pruning against the unmodified teacher.
- **HUMANEVAL**: 36.0 (base 40.9, Δ -4.9)
- **HUMANEVAL+PLUS**: 31.7 (base 36.6, Δ -4.9)
<p align="center">
<a href="https://cambriantech.github.io/forge-alloy/verify/#bba0a92ff0c8bebb">
<img src="alloy-qr.png" alt="Verify Chain of Custody" width="160"/>
</a>
</p>
<p align="center">
<a href="https://cambriantech.github.io/forge-alloy/verify/#bba0a92ff0c8bebb"><b>Every claim on this card is verified</b></a><br>
<b>Trust: self-attested</b> · 2 benchmarks · 1 device tested<br>
<a href="https://github.com/CambrianTech/forge-alloy">ForgeAlloy</a> chain of custody · <a href="olmoe-1b-7b-compacted-5b.alloy.json">Download alloy</a> · Merkle-chained
</p>
---
## About this model
Cross-architecture validation artifact for the §4.1.3.4 calibration-aware expert importance methodology. OLMoE-1B-7B-0924-Instruct (the smallest serious MoE on HuggingFace, fully-open Allen AI release) compacted from 64 experts per layer to 48 via per-layer-normalized activation-count importance ranking on a held-out Python code calibration corpus. Hardware-measured 36.0 HumanEval / 31.7 HumanEval+ vs the unmodified base's 40.9 / 36.6 — within −4.9 / −4.9 of the base anchor. The negative-baseline broad-corpus variant scored 28.0 / 26.2 (Δ −12.9 / −10.4); the +8.0 / +5.5 swing from changing only the calibration corpus is the second empirical anchor for §4.1.3.4 (the first was Qwen3-Coder-30B-A3B with a +9.7 swing). Two architectures (`Qwen3MoeForCausalLM` and `OlmoeForCausalLM`) now empirically validate the cross-architecture invariance claim: the metric is architecture-invariant, the calibration-corpus alignment is the lever.
## Benchmarks
| Benchmark | Score | Base | Δ | Verified |
|---|---|---|---|---|
| **humaneval** | **36.0** | 40.9 | -4.9 | ✅ Result hash |
| **humaneval_plus** | **31.7** | 36.6 | -4.9 | ✅ Result hash |
## What Changed (Base → Forged)
| | Base | Forged | Delta |
|---|---|---|---|
| **Pipeline** | | expert-activation-profile → expert-prune → quant → eval | 1 cycles |
## Runs On
| Device | Format | Size | Speed |
|--------|--------|------|-------|
| **NVIDIA GeForce RTX 5090** | Q5_K_M | 3.6GB | Verified |
| MacBook Pro 32GB | fp16 | 3.6GB | Expected |
| MacBook Air 16GB | Q8_0 | ~1.8GB | Expected |
| MacBook Air 8GB | Q4_K_M | ~1.1GB | Expected |
| iPhone / Android | Q4_K_M | ~1.1GB | Expected |
## Quick Start
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("continuum-ai/olmoe-1b-7b-compacted-5b",
torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("continuum-ai/olmoe-1b-7b-compacted-5b")
inputs = tokenizer("def merge_sort(arr):", return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```
## How It Was Made
```
expert-activation-profile → expert-prune → quant → eval (1 cycles)
```
- **expert-activation-profile**
> Same script unchanged from the Qwen3-Coder-30B-A3B forge — first cross-architecture validation that the activation-count importance metric ports across MoE families. The hooks register on `model.layers.{L}.mlp.gate` for both Qwen3MoE and OlmoeForCausalLM (same module path).
- **Expert pruning**: 0% of MoE experts removed pre-load
> Same script unchanged. Identical regex layout (unfused per-expert tensors at `model.layers.{L}.mlp.experts.{K}.{gate,up,down}_proj.weight`). Cross-arch portability confirmed: OlmoeForCausalLM and Qwen3MoeForCausalLM share the same prunable-unit module structure, so the script works without modification.
- **quant**
- **Calibrated evaluation**: anchored against `OLMoE-1B-7B-0924-Instruct` (published None, measured 40.9, ±3.0pt tolerance)
> Self-anchor calibration. HumanEval is not OLMoE's natural benchmark — OLMoE is general-purpose, not coder-specific. The 40.9 base / 36.0 student numbers are methodology validation, not tier-leading absolute quality. The artifact's value is the structural finding (cross-architecture portability + +8.0 swing from calibration alignment), not the absolute number.
- **Hardware**: NVIDIA GeForce RTX 5090
- **Forge tool**: [Continuum](https://github.com/CambrianTech/continuum) Factory + [sentinel-ai](https://github.com/CambrianTech/sentinel-ai)
## Limitations
- **HumanEval is not OLMoE's natural benchmark.** OLMoE is general-purpose (Allen AI), not coder-specific. The 40.9 base / 36.0 student numbers are methodology validation, not tier-leading absolute quality. For a tier-leading code model, see [`qwen3-coder-30b-a3b-compacted-19b-256k`](https://huggingface.co/continuum-ai/qwen3-coder-30b-a3b-compacted-19b-256k).
- **Validates §4.1.3.4 cross-architecture; does NOT compete on absolute numbers.** This is the second empirical anchor for the methodology paper, alongside the Qwen3-Coder-30B-A3B v1. Together they demonstrate that the activation-count importance metric is architecture-invariant across two structurally distinct MoE families.
- Calibration corpus was 300 Python code examples. For non-code workloads (math/reasoning/general), the methodology will preserve OLMoE's general capability if profiled on a matching corpus — but that's a separate forge run.
- Single GGUF tier shipped (Q5_K_M, 3.6 GB). Q4_K_M and Q8_0 will be added in v1.1 if there's demand.
## Chain of Custody
Scan the QR or [verify online](https://cambriantech.github.io/forge-alloy/verify/#bba0a92ff0c8bebb). Download the [alloy file](olmoe-1b-7b-compacted-5b.alloy.json) to verify independently.
| What | Proof |
|------|-------|
| Forged on | NVIDIA GeForce RTX 5090, ? |
| Published | [huggingface](https://huggingface.co/continuum-ai/olmoe-1b-7b-compacted-5b) — 2026-04-08T16:36:55.037319+00:00 |
| Trust level | [`self-attested`](https://github.com/CambrianTech/forge-alloy/blob/main/docs/ATTESTATION.md) |
| Spec | [ForgeAlloy](https://github.com/CambrianTech/forge-alloy) — Rust/Python/TypeScript |
## Make Your Own
Forged with [Continuum](https://github.com/CambrianTech/continuum) — a distributed AI world that runs on your hardware.
<p align="center">
<a href="https://github.com/CambrianTech/continuum"><img src="https://raw.githubusercontent.com/CambrianTech/continuum/main/docs/images/factory.png" alt="Continuum Model Factory" width="400"/></a>
</p>
The Factory configurator lets you design and forge custom models visually — context extension, pruning, LoRA, quantization, vision/audio modalities. Pick your target devices, the system figures out what fits.
[GitHub](https://github.com/CambrianTech/continuum) · [All Models](https://huggingface.co/continuum-ai) · [Forge-Alloy](https://github.com/CambrianTech/forge-alloy)
## License
apache-2.0
|