File size: 6,530 Bytes
2f389f3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bda7c36
 
 
 
 
 
2f389f3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1b8c040
 
 
 
70eb70f
3eab76b
1b8c040
 
 
 
e5924e6
1b8c040
e5924e6
1b8c040
 
 
 
2f389f3
78cb7f8
 
 
 
 
 
2f389f3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
---
license: openmdw-1.1
license_link: https://huggingface.co/poolside/Laguna-S-2.1/blob/main/LICENSE.md
base_model: poolside/Laguna-S-2.1
base_model_relation: quantized
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- oq
- quantized
- moe
- laguna
---
# Laguna-S-2.1-oQ5e

Calibrated 5-bit MLX quantization of [poolside/Laguna-S-2.1](https://huggingface.co/poolside/Laguna-S-2.1)
(118B total, 8B activated per token), produced with [oMLX](https://github.com/jundot/omlx) oQ at
level 5 enhanced — **5.30 bits/weight effective**, 78 GB on disk. Data-driven mixed precision:
bits are allocated per tensor from an imatrix-calibrated sensitivity map, not a fixed rule.
For Apple Silicon.

- **78 GB** on disk, down from 235 GB BF16
- 48 layers, 47 of them MoE with 256 routed experts + 1 shared, top-10 (L0 is a dense MLP);
  interleaved attention (12 global with YaRN to 1M context, 36 sliding-window 512)
- Peak memory in my tests: **73.5 GB** at 1k context, 76.6 GB at 64k — fits a 96 GB Mac
- Converted and tested on a **Macbook Pro M5 Max 128GB 40 GPU**

## Requirements

mlx-lm doesn't support the `laguna` architecture yet — there's an open PR:
[mlx-lm#1223](https://github.com/ml-explore/mlx-lm/pull/1223). Until it lands, use **mlx-vlm**
(0.6.3+), which implements laguna as a text-only model:

```bash
uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-S-2.1-oQ5e --prompt "..."
```

oMLX serves it directly from **0.5.3** on — it vendors that PR and patches it into mlx-lm at
import, so no model setting is needed. On earlier builds, discovery decides between the mlx-lm and
mlx-vlm loaders by looking for a vision sub-config, and laguna has none — so it lands on mlx-lm and
fails with `Model type laguna not supported`. Set **`model_type_override: "vlm"`** in the model's
settings, then refresh discovery (`omlx restart`): the load failure is cached per entry until the
next discovery pass, so setting the override alone won't clear it.

## Quantization

oQ5e allocates bits per tensor from an importance-matrix calibration pass over calibration data.
The 5-bit base lands on the experts; the dense spine — attention, embeddings, `lm_head`, routers,
386 tensors in total — came out mixed, 175 at 8 bits and 211 at 6. Output is standard MLX affine
quantization — no custom kernels or runtime required.

Unlike the smaller variants, this build was calibrated with omlx's newer adaptive imatrix
collection: up to 1024 samples, extended until every routed expert is covered, instead of a
fixed 128 samples.

## How it was quantized

oQ at level 5 enhanced — imatrix-calibrated, group size 128. I had to patch omlx to route laguna
through the mlx-vlm loader. At 235 GB the model doesn't fit in 128 GB of RAM, so calibration ran
against a uniform 4-bit proxy on disk rather than the FP weights, which shifts the bit allocation
slightly.

## Conversion check

Smoke-tested after conversion with `mlx_vlm.generate`: coherent — solved `17 * 24 = 408` and
verified it with the standard algorithm, no repetition loop. 55.5 tok/s generation on a short
prompt, peak 78.2 GB.

## Performance

Measured with oMLX's benchmark harness on a **Macbook Pro M5 Max 128GB 40 GPU**, single request,
128 generated tokens:

| prompt | gen tok/s | prefill tok/s | TTFT ms | peak GB |
|---|---|---|---|---|
| 1k | 57.5 | 960.4 | 1067 | 73.46 |
| 4k | 55.9 | 1020.9 | 4013 | 73.60 |
| 8k | 53.4 | 900.6 | 9098 | 73.80 |
| 16k | 50.0 | 813.0 | 20153 | 74.16 |
| 32k | 45.5 | 775.6 | 42250 | 74.95 |
| 64k | 38.1 | 696.3 | 94120 | 76.58 |

Continuous batching at 1k prompt / 128 generated:

| batch | tg tok/s | speedup | TTFT ms | E2E s |
|---|---|---|---|---|
| 1 | 57.5 | 1.00x | 1067 | 3.30 |
| 2 | 78.8 | 1.37x | 2079 | 5.33 |
| 4 | 102.7 | 1.79x | 3630 | 8.70 |
| 8 | 125.1 | 2.18x | 5374 | 15.06 |

## Benchmarks & Variants

mmlu_pro, mathqa and winogrande, n=300 seeded samples each, thinking off, identical questions across
every variant. The bf16 row is the hosted API, measured the same way. Standard error at this n is
around 2.5 points, so oQ4e through oQ6e aren't separated by this run.

![Accuracy vs bits per weight, three benchmarks, n=300](ladder.png)

| Variant | Size | bpw | gen tok/s (1k → 64k) | mmlu_pro | mathqa | winogrande |
|---|---|---|---|---|---|---|
| [Laguna-S-2.1-oQ2e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ2e-fast) | 35 GB | 2.60 | 78.8 → 48.6 | 0.700 | 0.850 | 0.713 |
| [Laguna-S-2.1-oQ2e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ2e) | 36 GB | 2.70 | 61.5 → 38.8 | 0.703 | 0.840 | 0.707 |
| [Laguna-S-2.1-oQ3e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ3e-fast) | 49 GB | 3.56 | 77.2 → 48.4 | 0.750 | 0.887 | 0.760 |
| [Laguna-S-2.1-oQ3e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ3e) | 49 GB | 3.59 | 67.5 → 40.1 | 0.750 | 0.880 | 0.760 |
| [Laguna-S-2.1-oQ4e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ4e-fast) | 63 GB | 4.54 | 69.3 → 45.5 | 0.787 | 0.873 | 0.777 |
| [Laguna-S-2.1-oQ4e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ4e) | 64 GB | 4.60 | 55.9 → 39.8 | 0.757 | 0.887 | 0.777 |
| [**Laguna-S-2.1-oQ5e**](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ5e) (this repo) | 78 GB | 5.30 | 57.5 → 38.1 | 0.773 | 0.883 | 0.797 |
| [Laguna-S-2.1-oQ6e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ6e) | 92 GB | 6.27 | 53.0 → 32.9 | 0.763 | 0.873 | 0.777 |
| [Laguna S 2.1 (API, bf16)](https://openrouter.ai/poolside/laguna-s-2.1) | — | 16 | — | 0.773 | 0.880 | 0.810 |

Treat this as a rough sighting, not a verdict. Three benchmarks at n=300 cover a narrow slice of what
the model does — no long-context work, no agentic loops, no real code — and at this sample size most
of the ladder above 3.6 bpw sits inside the error bars. I ran them to size the drop between levels,
not to rank the variants against each other. Test the one you're considering on your own workload
before trusting any of it.

## Usage

```bash
# mlx-vlm — plain mlx-lm doesn't support the laguna architecture
uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-S-2.1-oQ5e \
  --prompt "Explain Bayes' theorem in two sentences." --max-tokens 300

# oMLX — discovers the model from the HF cache; set model_type_override: "vlm" first
omlx serve
```

## License

[OpenMDW-1.1](https://huggingface.co/poolside/Laguna-S-2.1/blob/main/LICENSE.md), inherited from
the base model. Refer to the original model card for architecture, benchmarks, and intended use.