Myric's picture
Upload GATE.md with huggingface_hub
d76fdfe verified
|
Raw
History Blame Contribute Delete
990 Bytes

Quality gate — Moonlight-16B-A3B-Instruct (retrofit)

Applied after the fact to quants built before this repo had a quality-gate process. Read-only: the existing GGUFs below were not re-quantized, only tested — real generation on a fixed prompt suite, and (where a working chat template exists) tool-calling reliability. This does not change what's published, only documents it honestly.

No tool-calling test for this model: this model's chat template has no tool-call support at all (verified: no tools/tool_calls handling in the embedded template, no tool-related special tokens in the tokenizer, and no mention of function/tool calling anywhere in the model card — Moonlight is a Muon-optimizer research release, not trained for tool use). Not a quant defect — a limitation of the base model's chat template.

size wikitext PPL vs fp coherent
fp 8.71 1.00x
handroll 8.82 1.01x yes
i-quality 8.77 1.01x yes