Feature Extraction
MLX
Safetensors
qwen3
apple-silicon
quantized
mixed-precision
axquant
axq
development
8bit
8-bit precision
v2
embedding
sentence-similarity
6-bit
Instructions to use AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
| license: apache-2.0 | |
| library_name: mlx | |
| base_model: Qwen/Qwen3-Embedding-0.6B | |
| base_model_relation: quantized | |
| pipeline_tag: feature-extraction | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - quantized | |
| - mixed-precision | |
| - axquant | |
| - axq | |
| - development | |
| - qwen3 | |
| - 8bit | |
| - 8-bit | |
| - v2 | |
| - embedding | |
| - sentence-similarity | |
| # AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit | |
| An **AXQuant (AXQ)** mixed-precision MLX checkpoint for Apple Silicon, converted directly from | |
| the BF16 source model. The language path is quantized under AXQuant protection floors (embeddings, norms, and other protected tensors remain higher precision). | |
| > **Development evidence — not a certified AXQuant release.** This package has conversion and | |
| > artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed, | |
| > or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim. | |
| > **Stable-name v2.** `main` serves the audited v2 artifact for backward compatibility. The same revision is tagged `v2`; the replaced artifact remains recoverable at `legacy-pre-v2`. | |
| ## Model details | |
| | Property | Value | | |
| | --- | --- | | |
| | Base model | [Qwen/Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B/tree/97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3) | | |
| | Source revision | `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3` | | |
| | Product family | `qwen3` | | |
| | Source architecture | `Qwen3ForCausalLM` (dense); text path optimized | | |
| | Main-model parameters | 595.78M logical parameters | | |
| | Quantizer | AXQuant `1.2.0` | | |
| | Hub budget class | `8bit` | | |
| | Artifact edition | `v2` | | |
| | AXQuant base precision class | `8bit` | | |
| | Planned storage-adjusted BPW | 7.9992 | | |
| | Measured main-model BPW | 8.0003 | | |
| | Measured total BPW | **8.0003** | | |
| | Safetensors weight size | 0.60 GB | | |
| | Approximate complete download | 0.61 GB | | |
| | Configured maximum context | 32,768 tokens; practical limits depend on unified memory | | |
| | MLX-LM compatibility | Standard text inference, compatibility level B | | |
| | AX Engine native execution | Not established; no validated native manifest is included | | |
| | MTP present | `False` | | |
| | Vision sidecar present | `False` | | |
| This repository contains MLX Safetensors. It does **not** contain PyTorch or GGUF weights. | |
| ## Choosing an AXQ pack | |
| AXQ names describe a **storage-budget product class**, not one uniform precision applied to every | |
| tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. | |
| In particular, a `6bit`-named mixed plan may retain `4bit` as its base precision while selecting | |
| 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection | |
| floors can also raise a `4bit`-named pack close to (or above) a `6bit` budget on small or heavily | |
| protected models. | |
| | Sibling | Intended trade-off | | |
| | --- | --- | | |
| | [4bit sibling](https://huggingface.co/AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-4bit) | Lower-storage AXQ budget; check its exact BPW | | |
| | [8bit sibling](https://huggingface.co/AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit) | Higher average precision near the 8-BPW budget | | |
| See the [AutomatosX MLX model catalog](https://huggingface.co/collections/AutomatosX/automatosx-mlx-model-catalog) | |
| for related MLX and OptiQ alternatives. | |
| ## Download | |
| ```bash | |
| python -m pip install -U huggingface_hub | |
| hf download AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit --local-dir ./AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit | |
| ``` | |
| Allow at least 0.61 GB of free disk space. Pin the resulting Hub commit in reproducible | |
| deployments rather than relying indefinitely on `main`. | |
| ## Run with MLX-LM | |
| ```bash | |
| python -m pip install -U mlx-lm | |
| mlx_lm.generate \ | |
| --model AutomatosX/AX-Qwen3-Embedding-0.6B-MLX-AXQ-8bit \ | |
| --prompt "Explain mixed-precision quantization in three sentences." \ | |
| --max-tokens 128 \ | |
| --temp 0.0 | |
| ``` | |
| MLX-LM compatibility covers standard **text/backbone inference**. It may ignore AXQuant runtime | |
| metadata and optional sidecars (`vision.safetensors`, `mtp.safetensors`); this command therefore | |
| does not establish MTP acceleration or vision-language quality. The artifact records MLX | |
| `0.32.0` and MLX-LM `0.31.3` from conversion. | |
| ## AX Engine status | |
| This package does **not** include a validated native `model-manifest.json`, so AX Engine execution | |
| is not established by this release. The AX Engine fields in `axquant_runtime.json` describe the | |
| intended compatibility contract, not observed runtime evidence. Use the MLX-LM path above for | |
| standard text/backbone inference. The artifact records AX Engine version | |
| `not recorded`, but version discovery alone is not a runtime check. | |
| ## Quantization layout | |
| | Main-weight precision | Parameters | Share | | |
| | --- | ---: | ---: | | |
| | `6bit` | 201.33M | 33.79% | | |
| | `8bit` | 394.38M | 66.20% | | |
| | `bf16` | 65,536 | 0.01% | | |
| - Quantization methods: `affine, bf16`. | |
| - Group sizes used by quantized assignments: `32, 64`. | |
| - MTP sidecar: not included. | |
| - Vision sidecar: not included. | |
| - Optimization scope: `text-path`. | |
| - Support tier: `convertible`. | |
| BF16 sidecars, when present, are included in total download size. Their presence does not by itself | |
| establish MTP acceleration or vision-language quality. | |
| ## Evidence and validation status | |
| | Check | Status | | |
| | --- | --- | | |
| | Planning evidence | `architecture_prior` | | |
| | Calibration | none; the allocation is based on architecture priors | | |
| | Quantizer execution | 197/197 recorded module conversions succeeded; 0 fallbacks | | |
| | AX Engine native manifest | not included | | |
| | Quality versus BF16 or uniform baselines | Not published; no quality-retention claim | | |
| | MTP acceptance and speed | not measured; no MTP speedup claim | | |
| | AX Engine kernel evidence | `unmeasured` | | |
| | Vision-language quality | Not applicable (no vision sidecar in this package) | | |
| | Long-context quality | 32,768-token capacity is config metadata, not a validated claim | | |
| | Release certification | **Not certified**; formal AXQuant M0-M8 gates are not closed | | |
| ## Intended use and limitations | |
| - Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes. | |
| - No minimum unified-memory figure is claimed; loadability depends on model size, context length, | |
| KV-cache policy, runtime buffers, and other processes using unified memory. | |
| - Architecture-prior allocation is not measured sensitivity. It must not be presented as measured | |
| model quality. | |
| - The configured context window can require substantially more memory as the KV cache grows. | |
| - AX Engine execution is not established because this package has no validated native manifest. | |
| - Upstream capabilities, limitations, biases, and responsible-use guidance still apply. | |
| ## Provenance and audit files | |
| - [`axquant_manifest.json`](axquant_manifest.json): package identity, byte accounting, runtime | |
| contract, software versions, and file checksums. | |
| - [`axquant_plan.json`](axquant_plan.json): per-tensor precision decisions and planning evidence. | |
| - [`axquant_quantizer_execution.json`](axquant_quantizer_execution.json): conversion coverage and | |
| fallback records. | |
| - [`axquant_runtime.json`](axquant_runtime.json): declared AX Engine and MLX-LM compatibility metadata; runtime checks remain separate evidence. | |
| All published provenance uses repository-relative paths. Local source paths are stripped before | |
| publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ | |
| artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have | |
| identical BPW or quality. | |
| ## License | |
| The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See | |
| the [Qwen/Qwen3-Embedding-0.6B model card](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B/tree/97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3) for license terms, model | |
| limitations, and responsible-use guidance. | |