Text Generation
MLX
Safetensors
laguna
oq
quantized
Mixture of Experts
conversational
custom_code
5-bit
Instructions to use mlx-community/Laguna-S-2.1-oQ5e with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Laguna-S-2.1-oQ5e with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/Laguna-S-2.1-oQ5e") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/Laguna-S-2.1-oQ5e with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/Laguna-S-2.1-oQ5e"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/Laguna-S-2.1-oQ5e" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use mlx-community/Laguna-S-2.1-oQ5e with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/Laguna-S-2.1-oQ5e"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/Laguna-S-2.1-oQ5e
Run Hermes
hermes
- OpenClaw new
How to use mlx-community/Laguna-S-2.1-oQ5e with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/Laguna-S-2.1-oQ5e"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/Laguna-S-2.1-oQ5e" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use mlx-community/Laguna-S-2.1-oQ5e with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "mlx-community/Laguna-S-2.1-oQ5e"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "mlx-community/Laguna-S-2.1-oQ5e" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/Laguna-S-2.1-oQ5e", "messages": [ {"role": "user", "content": "Hello"} ] }'
| license: openmdw-1.1 | |
| license_link: https://huggingface.co/poolside/Laguna-S-2.1/blob/main/LICENSE.md | |
| base_model: poolside/Laguna-S-2.1 | |
| base_model_relation: quantized | |
| library_name: mlx | |
| pipeline_tag: text-generation | |
| tags: | |
| - mlx | |
| - oq | |
| - quantized | |
| - moe | |
| - laguna | |
| # Laguna-S-2.1-oQ5e | |
| Calibrated 5-bit MLX quantization of [poolside/Laguna-S-2.1](https://huggingface.co/poolside/Laguna-S-2.1) | |
| (118B total, 8B activated per token), produced with [oMLX](https://github.com/jundot/omlx) oQ at | |
| level 5 enhanced — **5.30 bits/weight effective**, 78 GB on disk. Data-driven mixed precision: | |
| bits are allocated per tensor from an imatrix-calibrated sensitivity map, not a fixed rule. | |
| For Apple Silicon. | |
| - **78 GB** on disk, down from 235 GB BF16 | |
| - 48 layers, 47 of them MoE with 256 routed experts + 1 shared, top-10 (L0 is a dense MLP); | |
| interleaved attention (12 global with YaRN to 1M context, 36 sliding-window 512) | |
| - Peak memory in my tests: **73.5 GB** at 1k context, 76.6 GB at 64k — fits a 96 GB Mac | |
| - Converted and tested on a **Macbook Pro M5 Max 128GB 40 GPU** | |
| ## Requirements | |
| mlx-lm doesn't support the `laguna` architecture yet — there's an open PR: | |
| [mlx-lm#1223](https://github.com/ml-explore/mlx-lm/pull/1223). Until it lands, use **mlx-vlm** | |
| (0.6.3+), which implements laguna as a text-only model: | |
| ```bash | |
| uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-S-2.1-oQ5e --prompt "..." | |
| ``` | |
| oMLX serves it directly from **0.5.3** on — it vendors that PR and patches it into mlx-lm at | |
| import, so no model setting is needed. On earlier builds, discovery decides between the mlx-lm and | |
| mlx-vlm loaders by looking for a vision sub-config, and laguna has none — so it lands on mlx-lm and | |
| fails with `Model type laguna not supported`. Set **`model_type_override: "vlm"`** in the model's | |
| settings, then refresh discovery (`omlx restart`): the load failure is cached per entry until the | |
| next discovery pass, so setting the override alone won't clear it. | |
| ## Quantization | |
| oQ5e allocates bits per tensor from an importance-matrix calibration pass over calibration data. | |
| The 5-bit base lands on the experts; the dense spine — attention, embeddings, `lm_head`, routers, | |
| 386 tensors in total — came out mixed, 175 at 8 bits and 211 at 6. Output is standard MLX affine | |
| quantization — no custom kernels or runtime required. | |
| Unlike the smaller variants, this build was calibrated with omlx's newer adaptive imatrix | |
| collection: up to 1024 samples, extended until every routed expert is covered, instead of a | |
| fixed 128 samples. | |
| ## How it was quantized | |
| oQ at level 5 enhanced — imatrix-calibrated, group size 128. I had to patch omlx to route laguna | |
| through the mlx-vlm loader. At 235 GB the model doesn't fit in 128 GB of RAM, so calibration ran | |
| against a uniform 4-bit proxy on disk rather than the FP weights, which shifts the bit allocation | |
| slightly. | |
| ## Conversion check | |
| Smoke-tested after conversion with `mlx_vlm.generate`: coherent — solved `17 * 24 = 408` and | |
| verified it with the standard algorithm, no repetition loop. 55.5 tok/s generation on a short | |
| prompt, peak 78.2 GB. | |
| ## Performance | |
| Measured with oMLX's benchmark harness on a **Macbook Pro M5 Max 128GB 40 GPU**, single request, | |
| 128 generated tokens: | |
| | prompt | gen tok/s | prefill tok/s | TTFT ms | peak GB | | |
| |---|---|---|---|---| | |
| | 1k | 57.5 | 960.4 | 1067 | 73.46 | | |
| | 4k | 55.9 | 1020.9 | 4013 | 73.60 | | |
| | 8k | 53.4 | 900.6 | 9098 | 73.80 | | |
| | 16k | 50.0 | 813.0 | 20153 | 74.16 | | |
| | 32k | 45.5 | 775.6 | 42250 | 74.95 | | |
| | 64k | 38.1 | 696.3 | 94120 | 76.58 | | |
| Continuous batching at 1k prompt / 128 generated: | |
| | batch | tg tok/s | speedup | TTFT ms | E2E s | | |
| |---|---|---|---|---| | |
| | 1 | 57.5 | 1.00x | 1067 | 3.30 | | |
| | 2 | 78.8 | 1.37x | 2079 | 5.33 | | |
| | 4 | 102.7 | 1.79x | 3630 | 8.70 | | |
| | 8 | 125.1 | 2.18x | 5374 | 15.06 | | |
| ## Benchmarks & Variants | |
| mmlu_pro, mathqa and winogrande, n=300 seeded samples each, thinking off, identical questions across | |
| every variant. The bf16 row is the hosted API, measured the same way. Standard error at this n is | |
| around 2.5 points, so oQ4e through oQ6e aren't separated by this run. | |
|  | |
| | Variant | Size | bpw | gen tok/s (1k → 64k) | mmlu_pro | mathqa | winogrande | | |
| |---|---|---|---|---|---|---| | |
| | [Laguna-S-2.1-oQ2e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ2e-fast) | 35 GB | 2.60 | 78.8 → 48.6 | 0.700 | 0.850 | 0.713 | | |
| | [Laguna-S-2.1-oQ2e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ2e) | 36 GB | 2.70 | 61.5 → 38.8 | 0.703 | 0.840 | 0.707 | | |
| | [Laguna-S-2.1-oQ3e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ3e-fast) | 49 GB | 3.56 | 77.2 → 48.4 | 0.750 | 0.887 | 0.760 | | |
| | [Laguna-S-2.1-oQ3e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ3e) | 49 GB | 3.59 | 67.5 → 40.1 | 0.750 | 0.880 | 0.760 | | |
| | [Laguna-S-2.1-oQ4e-fast](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ4e-fast) | 63 GB | 4.54 | 69.3 → 45.5 | 0.787 | 0.873 | 0.777 | | |
| | [Laguna-S-2.1-oQ4e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ4e) | 64 GB | 4.60 | 55.9 → 39.8 | 0.757 | 0.887 | 0.777 | | |
| | [**Laguna-S-2.1-oQ5e**](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ5e) (this repo) | 78 GB | 5.30 | 57.5 → 38.1 | 0.773 | 0.883 | 0.797 | | |
| | [Laguna-S-2.1-oQ6e](https://huggingface.co/mlx-community/Laguna-S-2.1-oQ6e) | 92 GB | 6.27 | 53.0 → 32.9 | 0.763 | 0.873 | 0.777 | | |
| | [Laguna S 2.1 (API, bf16)](https://openrouter.ai/poolside/laguna-s-2.1) | — | 16 | — | 0.773 | 0.880 | 0.810 | | |
| Treat this as a rough sighting, not a verdict. Three benchmarks at n=300 cover a narrow slice of what | |
| the model does — no long-context work, no agentic loops, no real code — and at this sample size most | |
| of the ladder above 3.6 bpw sits inside the error bars. I ran them to size the drop between levels, | |
| not to rank the variants against each other. Test the one you're considering on your own workload | |
| before trusting any of it. | |
| ## Usage | |
| ```bash | |
| # mlx-vlm — plain mlx-lm doesn't support the laguna architecture | |
| uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-S-2.1-oQ5e \ | |
| --prompt "Explain Bayes' theorem in two sentences." --max-tokens 300 | |
| # oMLX — discovers the model from the HF cache; set model_type_override: "vlm" first | |
| omlx serve | |
| ``` | |
| ## License | |
| [OpenMDW-1.1](https://huggingface.co/poolside/Laguna-S-2.1/blob/main/LICENSE.md), inherited from | |
| the base model. Refer to the original model card for architecture, benchmarks, and intended use. | |