malaiwah's picture
Add files using upload-large-folder tool
4727bb7 verified
|
Raw
History Blame Contribute Delete
2.62 kB
metadata
license: mit
base_model:
  - mindlab-research/Macaron-V1-Venti
  - zai-org/GLM-5.2
base_model_relation: quantized
library_name: safetensors
tags:
  - lora
  - fp8
  - nvfp4
  - glm
  - reference-artifact

Macaron-V1-Venti FP8 and NVFP4 LoRA factors

This repository contains unofficial FP8 E4M3FN and NVIDIA NVFP4 W4A16 quantizations of the four BF16 rank-16 LoRA adapters (L0L3) from mindlab-research/Macaron-V1-Venti, pinned at revision 3d6f30eea38663a7b9320f3a6b28822ed4aa7ac4. The upstream base model is zai-org/GLM-5.2.

These are reference-format artifacts, not drop-in PEFT or stock-vLLM adapters. A compatible runtime must understand weight_packed, weight_scale, and (for NVFP4) weight_scale_2. The artifacts do not modify or contain the EXL3 base checkpoint.

Layout

fp8/{L0,L1,L2,L3}/
nvfp4-w4a16/{L0,L1,L2,L3}/
validation/
conversion/

FP8 uses one symmetric FP32 scale per factor and E4M3FN stored factors. NVFP4 uses packed E2M1 values, an E4M3FN scale for every 16 values, and one FP32 double scale per factor. NVFP4 conversion uses NVIDIA Model Optimizer revision 87c9f8cf83021957d1a1a575c90c9a4eaaf7ef0c.

The validation/ directory reports deterministic sampled comparisons of BF16 and dequantized B(Ax) outputs. Checksums and complete source provenance are included in the repository.

Offline validation

Each row uses 256 deterministic LoRA A/B pairs selected by tensor-name hash. Relative RMSE is measured on B(Ax) for four seeded synthetic input columns.

Adapter FP8 mean relative RMSE FP8 mean cosine NVFP4 mean relative RMSE NVFP4 mean cosine
L0 3.77% 0.999290 13.73% 0.990591
L1 3.81% 0.999270 13.83% 0.990437
L2 3.76% 0.999294 13.69% 0.990665
L3 3.78% 0.999288 13.80% 0.990509

These are adapter-only numerical checks, not end-to-end benchmark scores. NVFP4 should be treated as experimental until base-model evaluation confirms that task quality is retained.

Limitations

  • No claim of stock vLLM, PEFT, or EXL3 runtime compatibility is made.
  • Quantization is lossy; consult the validation reports before integration.
  • End-to-end model quality must still be measured after runtime support lands.
  • This project is not endorsed by MindLab Research or Zhipu AI.

License and attribution

The upstream repositories publish these weights under MIT. See THIRD_PARTY_NOTICES.md for the full notice and both upstream copyright statements.