--- language: - en tags: - gptq - auto-gptq - autogptq - pytorch - causal-lm - autoround - auto-round - intel-autoround - intel - woq - weights-only-quantization license: apache-2.0 library_name: transformers model_name: Minerva 1B base v1.0 base_model: sapienzanlp/Minerva-1B-base-v1.0 inference: false model_creator: sapienzanlp pipeline_tag: text-generation prompt_template: '{prompt} ' quantized_by: fbaldassarri --- ## Model Information Quantized version of [sapienzanlp/Minerva-1B-base-v1.0](https://huggingface.co/sapienzanlp/Minerva-1B-base-v1.0) using `torch.bfloat16` for quantization tuning. - 4 bits (INT4) - group size = 64 - Symmetrical Quantization - Method: WoQ — GPTQ (AutoGPTQ algorithm) Fast and low memory, 2-3X speedup (slight accuracy drop at W4G64) Quantization framework: [Intel AutoRound](https://github.com/intel/auto-round) v0.13.0 Note: this INT4 version of `Minerva 1B base v1.0` has been quantized for inference on Intel CPU, Intel iGPU (Arc) via intel-extension-for-pytorch, Intel NPU (AI Boost on Core Ultra series) via OpenVINO. ## Replication Recipe The recommended way to reproduce this exact quantization is via the [auto-round-pipeline](https://git.epicdynamic.com/auto-round-pipeline) — the same orchestration tool that produced this artifact. ### Step 1 — Bootstrap the auto-round-pipeline Set up a dedicated conda environment using the pipeline's `setup.sh`. Any of the install modes below produces an environment that can reproduce this quantization; pick the one that matches your goals: ``` git clone https://git.epicdynamic.com/auto-round-pipeline cd auto-round-pipeline # Pinned PyPI wheel (fastest; matches what this pipeline used by default): bash setup.sh --pip-version 0.13.0 # Or build from intel/auto-round at the same tag (byte-identical reproducibility): bash setup.sh --source-tag v0.13.0 # Intel Arc iGPU acceleration (e.g. Core Ultra 185H) — append to either of the above: # ... --intel-xpu # NVIDIA / AMD opt-in: --cuda / --rocm ``` The script prints the resulting conda env name (something like `auto-round-pipeline-v0.13.0[-src][-xpu|-cuda|-rocm]`) at the end. ### Step 2 — Quantize just this model Activate the env that `setup.sh` created, then invoke the runner with the same job filters that produced this artifact: ``` conda activate python runner.py \ --model 'sapienzanlp/Minerva-1B-base-v1.0' \ --quant 'INT4-gs64' \ --format auto_gptq \ --no-upload # drop this to also push to HuggingFace Hub ``` ### Step 3 — (Optional) standalone Python recipe If you'd rather call auto-round directly without the orchestration wrapper, this is the exact call the pipeline made: ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer from auto_round import AutoRound model_name = "sapienzanlp/Minerva-1B-base-v1.0" model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16) tokenizer = AutoTokenizer.from_pretrained(model_name) bits, group_size, sym = 4, 64, True autoround = AutoRound( model, tokenizer, bits=bits, group_size=group_size, sym=sym, device_map="cpu", nsamples=128, iters=200, seqlen=512, batch_size=4, ) autoround.quantize_and_save("./AutoRound/sapienzanlp_Minerva-1B-base-v1.0-auto_gptq-int4-gs64-sym", format="auto_gptq") ``` ## Actual Run Conditions Recorded by the auto-round-pipeline at quantization time: | Field | Value | |---|---| | Intel auto-round version | 0.13.0 | | transformers version | 4.55.3 | | torch version | 2.12.0+cpu | | torch_dtype (load) | torch.bfloat16 | | calibration device | `cpu` | | calibration samples | 128 | | tuning iterations | 200 | | calibration seq len | 512 | | calibration batch size | 4 | | quantization duration | 6770.9s (112.8 min) | | completed at (UTC) | 2026-06-14T20:27:10.055854+00:00 | ## License [Apache 2.0 License](https://choosealicense.com/licenses/apache-2.0/) ## Disclaimer This quantized model comes with no warranty. It has been developed only for research purposes.