Any-to-Any
Transformers
Safetensors
gemma4
image-text-to-text
abliterated
uncensored
4-bit precision
Instructions to use mlx-community/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated-4bit-msq with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mlx-community/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated-4bit-msq with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("mlx-community/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated-4bit-msq") model = AutoModelForMultimodalLM.from_pretrained("mlx-community/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated-4bit-msq", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated — MLX 4.3 BPW
Mixed-precision MLX quantization of huihui-ai/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated, quantized with MLX Smart Quantize (MSQ) — my own sensitivity-based mixed-precision quantization method for Apple Silicon. It measures per-layer NMSE and assigns optimal bit widths automatically, combining architecture knowledge with measured data.
Details
- Type: Vision (VLM)
- Average: 4.30 bits per weight
- Method: MLX Smart Quantize (MSQ)
- AWQ scaling: applied to 30 groups
- Downloads last month
- 69
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for mlx-community/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated-4bit-msq
Base model
google/gemma-4-26B-A4B Finetuned
google/gemma-4-26B-A4B-it