Qwen3.5-4B-bnb-4bit / README.md
techwithsergiu's picture
upd visual tower info in readme
633d7a7 verified
|
Raw
History Blame
2.7 kB
---
tags:
- techwithsergiu
library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-4B/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:
- Qwen/Qwen3.5-4B
---
# Qwen3.5-4B-bnb-4bit
<img width="400px" src="https://qianwen-res.oss-accelerate.aliyuncs.com/logo_qwen3.5.png">
BNB NF4 4-bit quantization of [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B).
Retains the full visual tower — this is a **VLM-capable** model (image + text input).
Primary use-case: Unsloth LoRA fine-tuning when you need image understanding in the
fine-tuned result.
> If you only need text fine-tuning, use
> [techwithsergiu/Qwen3.5-text-4B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-text-4B-bnb-4bit)
> instead — same backbone, visual tower removed, lighter VRAM footprint.
## What was changed
- Quantized with `bitsandbytes` NF4 double-quant (`bnb_4bit_quant_type=nf4`, `bnb_4bit_compute_dtype=bfloat16`)
- Visual tower layers kept at **bf16** (`llm_int8_skip_modules`) — required for correct image inference
- `lm_head.weight` kept at **bf16** for output quality
## Model family
![](diagrams/diagram_01.png)
| Model | Type | Base model |
|---|---|---|
| [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) | f16 · VLM · source | — |
| **[techwithsergiu/Qwen3.5-4B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-4B-bnb-4bit)** | BNB NF4 · VLM | Qwen/Qwen3.5-4B |
| [techwithsergiu/Qwen3.5-text-4B](https://huggingface.co/techwithsergiu/Qwen3.5-text-4B) | bf16 · text-only | Qwen/Qwen3.5-4B |
| [techwithsergiu/Qwen3.5-text-4B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-text-4B-bnb-4bit) | BNB NF4 · text-only | Qwen3.5-text-4B |
| [techwithsergiu/Qwen3.5-text-4B-GGUF](https://huggingface.co/techwithsergiu/Qwen3.5-text-4B-GGUF) | GGUF quants | Qwen3.5-text-4B |
The visual tower is a bf16 overhead that scales with model size (~0.19 GB for 0.8B, ~0.62 GB for 2B/4B, ~0.85 GB for 9B).
BNB-quantized models are roughly 40% of the original f16 size (exact ratio varies by size).
## Fine-tuning
For VLM (image + text) fine-tuning with Unsloth, refer to the official guide:
[unsloth.ai/docs/models/qwen3.5/fine-tune](https://unsloth.ai/docs/models/qwen3.5/fine-tune)
## Pipeline diagram
![](diagrams/diagram_02.png)
## Acknowledgements
Based on [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)
by the Qwen Team. If you use this model in research, please cite the original:
```bibtex
@misc{qwen3.5,
title = {{Qwen3.5}: Towards Native Multimodal Agents},
author = {{Qwen Team}},
month = {February},
year = {2026},
url = {https://qwen.ai/blog?id=qwen3.5}
}
```