# Quality gate — Moonlight-16B-A3B-Instruct (retrofit) Applied after the fact to quants built before this repo had a quality-gate process. **Read-only**: the existing GGUFs below were not re-quantized, only tested — real generation on a fixed prompt suite, and (where a working chat template exists) tool-calling reliability. This does not change what's published, only documents it honestly. **No tool-calling test for this model: this model's chat template has no tool-call support at all (verified: no tools/tool_calls handling in the embedded template, no tool-related special tokens in the tokenizer, and no mention of function/tool calling anywhere in the model card — Moonlight is a Muon-optimizer research release, not trained for tool use).** Not a quant defect — a limitation of the base model's chat template. | size | wikitext PPL | vs fp | coherent | |---|---|---|---| | fp | 8.71 | 1.00x | — | | handroll | 8.82 | 1.01x | yes | | i-quality | 8.77 | 1.01x | yes |