This is a GGUF quantization of Poolside's Laguna S 2.1 using an imatrix based on Bartowski's dataset.

It was generated with llama.cpp 846e991e plus the changes from PR 25165, which at the time of building included commits up to 54f214a09b8

DFlash draft model is available at wimmmm/poolside-Laguna-S-2.1-DFlash-GGUF

To run with DFlash enabled, you can use the following command line (change quantizations as needed):

llama serve -hf wimmmm/poolside-Laguna-S-2.1-GGUF:Q4_K_M -hfd wimmmm/poolside-Laguna-S-2.1-DFlash-GGUF:Q4_0 --spec-type draft-dflash --spec-draft-n-max 15 -fa on --jinja --port 8000

Have fun!

Downloads last month
7,013
GGUF
Model size
118B params
Architecture
laguna
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for wimmmm/poolside-Laguna-S-2.1-GGUF

Quantized
(78)
this model