gte-Qwen2-1.5B-instruct — MLX Q6

This checkpoint passed the bounded local representation-fidelity gate described below.

This is the Q uniform affine quantization checkpoint from a matched local embedding-quantization experiment. It is published with explicit lineage, calibration evidence where applicable, and the bounded evaluation result that accompanied the conversion.

Provenance and lineage

  • Upstream model: Alibaba-NLP/gte-Qwen2-1.5B-instruct
  • Upstream revision recorded for publication: a9af15a6372d7d6b25e9fb07c2ccb9e1fe645644
  • Revision evidence: upstream revision verified at publication time; historical local snapshot metadata was not retained
  • Direct parent: TiGa-RCE/gte-Qwen2-1.5B-instruct-MLX-BF16
  • Conversion rule: every quantized checkpoint branches directly from the family MLX BF16 checkpoint; no lossy checkpoint was used to create another.
  • Quantization: Q uniform affine quantization, nominal 6-bit, group size 64
  • Local conversion stack: oMLX 0.5.3, mlx-lm 0.31.3, MLX 0.32.0
  • Full collection: MLX Embedding Quantization Matrix

PROVENANCE.json contains machine-readable lineage and SHA-256 hashes for the published weight files. No importance matrix was used for this checkpoint.

Bounded local evaluation

Metric Result
Top-1 retrieval 1.000
MRR 1.000
Mean aligned cosine vs BF16 0.997786
Minimum aligned cosine vs BF16 0.996132
Score RMSE vs BF16 0.004667
Queries with rank change 0
Predeclared gate PASS

The evaluation used 24 frozen query/document pairs, the upstream query instruction recipe, last-token pooling, L2 normalization, and direct comparison with vectors from the family BF16 checkpoint. This is an engineering smoke test, not MTEB and not a claim of universal quality. Retrieval success and representation fidelity are reported separately.

Runtime scope

This checkpoint targets Apple Silicon through MLX/oMLX. CUDA and PyTorch results are a separate control lane and must not be interpreted as measurements of MLX/Metal kernel performance.

License and attribution

Apache-2.0, following the upstream model card. The original model authors retain attribution for the upstream model; this repository contains a local MLX conversion or quantized derivative prepared by TiGa-RCE for reproducibility research.

Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including TiGa-RCE/gte-Qwen2-1.5B-instruct-MLX-Q6