--- base_model: poolside/Laguna-S-2.1 library_name: gguf pipeline_tag: text-generation tags: - gguf - quantized - laguna - mixture-of-experts - q2_k - q3_k - ds4 --- # Laguna S 2.1 GGUF This repository contains a reduced-memory Laguna S 2.1 quantization for DwarfStar. ## Mixed Q2_K/Q3_K variant `laguna-s-2.1-RoutedQ2_K-Last27Q3_K.gguf` keeps every non-routed tensor byte-identical to Poolside's `laguna-s-2.1-Q4_K_M.gguf` at revision `706fa69799926b6afde1af9e24ca2a4923f110a1`. Only routed expert tensors were requantized using the source importance matrix: - routed layers 1 through 20: Q2_K gate, up, and down - routed layers 21 through 47: Q3_K gate, up, and down - all other tensors: unchanged from the official Q4_K_M GGUF The file is 48,260,803,968 bytes (44.946 GiB), intended for full-residency inference on 64 GiB systems. Runtime memory also depends on context size and KV-cache allocation. SHA-256: ```text 61fc66596597985cb9408a8530de6322d9e0d5b1d2ad4ed6503938018e0ce903 ``` ## DwarfStar Q3_K routed Laguna inference is supported starting with DwarfStar commit `938227a2`. ```sh ./download_model.sh laguna-q2-q3 ./ds4 -m gguf/laguna-s-2.1-RoutedQ2_K-Last27Q3_K.gguf -p "Hello" ``` On an Apple M5 Max, the tested model reached approximately 514 tokens/second for a 4096-token prefill and 63 tokens/second steady-state generation. Against 100 official continuation vectors, the mixed model obtained average NLL 0.2583, 87/100 first-token matches, and average matching-prefix length 9.50 tokens. The corresponding full Q4_K_M measurements were 0.2352, 92/100, and 10.86.