--- license: openmdw-1.1 license_link: https://huggingface.co/poolside/Laguna-S-2.1/blob/main/LICENSE.md base_model: poolside/Laguna-S-2.1 base_model_relation: quantized library_name: mlx pipeline_tag: text-generation tags: - mlx - oq - quantized - moe - laguna --- # Laguna-S-2.1-oQ5e Calibrated 5-bit MLX quantization of [poolside/Laguna-S-2.1](https://huggingface.co/poolside/Laguna-S-2.1) (118B total, 8B activated per token), produced with [oMLX](https://github.com/jundot/omlx) oQ at level 5 enhanced — **5.30 bits/weight effective**, 78 GB on disk. Data-driven mixed precision: bits are allocated per tensor from an imatrix-calibrated sensitivity map, not a fixed rule. For Apple Silicon. - **78 GB** on disk, down from 235 GB BF16 - 48 layers, 47 of them MoE with 256 routed experts + 1 shared, top-10 (L0 is a dense MLP); interleaved attention (12 global with YaRN to 1M context, 36 sliding-window 512) - Peak memory in my tests: **73.5 GB** at 1k context, 76.6 GB at 64k — fits a 96 GB Mac - Converted and tested on a **Macbook Pro M5 Max 128GB 40 GPU** ## Requirements mlx-lm doesn't support the `laguna` architecture yet — there's an open PR: [mlx-lm#1223](https://github.com/ml-explore/mlx-lm/pull/1223). Until it lands, use **mlx-vlm** (0.6.3+), which implements laguna as a text-only model: ```bash uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-S-2.1-oQ5e --prompt "..." ``` oMLX serves it directly from **0.5.3** on — it vendors that PR and patches it into mlx-lm at import, so no model setting is needed. On earlier builds, discovery decides between the mlx-lm and mlx-vlm loaders by looking for a vision sub-config, and laguna has none — so it lands on mlx-lm and fails with `Model type laguna not supported`. Set **`model_type_override: "vlm"`** in the model's settings, then refresh discovery (`omlx restart`): the load failure is cached per entry until the next discovery pass, so setting the override alone won't clear it. ## Quantization oQ5e allocates bits per tensor from an importance-matrix calibration pass over calibration data. The 5-bit base lands on the experts; the dense spine — attention, embeddings, `lm_head`, routers, 386 tensors in total — came out mixed, 175 at 8 bits and 211 at 6. Output is standard MLX affine quantization — no custom kernels or runtime required. Unlike the smaller variants, this build was calibrated with omlx's newer adaptive imatrix collection: up to 1024 samples, extended until every routed expert is covered, instead of a fixed 128 samples. ## How it was quantized oQ at level 5 enhanced — imatrix-calibrated, group size 128. I had to patch omlx to route laguna through the mlx-vlm loader. At 235 GB the model doesn't fit in 128 GB of RAM, so calibration ran against a uniform 4-bit proxy on disk rather than the FP weights, which shifts the bit allocation slightly. ## Conversion check Smoke-tested after conversion with `mlx_vlm.generate`: coherent — solved `17 * 24 = 408` and verified it with the standard algorithm, no repetition loop. 55.5 tok/s generation on a short prompt, peak 78.2 GB. ## Performance Measured with oMLX's benchmark harness on a **Macbook Pro M5 Max 128GB 40 GPU**, single request, 128 generated tokens: | prompt | gen tok/s | prefill tok/s | TTFT ms | peak GB | |---|---|---|---|---| | 1k | 57.5 | 960.4 | 1067 | 73.46 | | 4k | 55.9 | 1020.9 | 4013 | 73.60 | | 8k | 53.4 | 900.6 | 9098 | 73.80 | | 16k | 50.0 | 813.0 | 20153 | 74.16 | | 32k | 45.5 | 775.6 | 42250 | 74.95 | | 64k | 38.1 | 696.3 | 94120 | 76.58 | Continuous batching at 1k prompt / 128 generated: | batch | tg tok/s | speedup | TTFT ms | E2E s | |---|---|---|---|---| | 1 | 57.5 | 1.00x | 1067 | 3.30 | | 2 | 78.8 | 1.37x | 2079 | 5.33 | | 4 | 102.7 | 1.79x | 3630 | 8.70 | | 8 | 125.1 | 2.18x | 5374 | 15.06 | ## Benchmarks & Variants mmlu_pro, mathqa and winogrande, n=300 seeded samples each, thinking off, identical questions across every variant. The bf16 row is the hosted API, measured the same way. Standard error at this n is around 2.5 points, so oQ4e through oQ6e aren't separated by this run. ![Accuracy vs bits per weight, three benchmarks, n=300](ladder.png) | Variant | Size | bpw | gen tok/s (1k → 64k) | mmlu_pro | mathqa | winogrande | |---|---|---|---|---|---|---| | [Laguna-S-2.1-oQ2e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ2e-fast) | 35 GB | 2.60 | 78.8 → 48.6 | 0.700 | 0.850 | 0.713 | | [Laguna-S-2.1-oQ2e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ2e) | 36 GB | 2.70 | 61.5 → 38.8 | 0.703 | 0.840 | 0.707 | | [Laguna-S-2.1-oQ3e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ3e) | 49 GB | 3.59 | 67.5 → 40.1 | 0.750 | 0.880 | 0.760 | | [Laguna-S-2.1-oQ4e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ4e) | 64 GB | 4.60 | 55.9 → 39.8 | 0.757 | 0.887 | 0.777 | | [**Laguna-S-2.1-oQ5e**](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ5e) (this repo) | 78 GB | 5.30 | 57.5 → 38.1 | 0.773 | 0.883 | 0.797 | | [Laguna-S-2.1-oQ6e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ6e) | 92 GB | 6.27 | 53.0 → 32.9 | 0.763 | 0.873 | 0.777 | | [Laguna S 2.1 (API, bf16)](https://openrouter.ai/poolside/laguna-s-2.1) | — | 16 | — | 0.773 | 0.880 | 0.810 | Treat this as a rough sighting, not a verdict. Three benchmarks at n=300 cover a narrow slice of what the model does — no long-context work, no agentic loops, no real code — and at this sample size most of the ladder above 3.6 bpw sits inside the error bars. I ran them to size the drop between levels, not to rank the variants against each other. Test the one you're considering on your own workload before trusting any of it. ## Usage ```bash # mlx-vlm — plain mlx-lm doesn't support the laguna architecture uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-S-2.1-oQ5e \ --prompt "Explain Bayes' theorem in two sentences." --max-tokens 300 # oMLX — discovers the model from the HF cache; set model_type_override: "vlm" first omlx serve ``` ## License [OpenMDW-1.1](https://huggingface.co/poolside/Laguna-S-2.1/blob/main/LICENSE.md), inherited from the base model. Refer to the original model card for architecture, benchmarks, and intended use.