Instructions to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF", filename="qwen35-4b-dflash-Q4_0.gguf", )
llm.create_chat_completion( messages = "No input example has been defined for this model task." )
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Use Docker
docker model run hf.co/shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
- LM Studio
- Jan
- Ollama
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with Ollama:
ollama run hf.co/shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
- Unsloth Studio
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF to start chatting
- Pi
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with Docker Model Runner:
docker model run hf.co/shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
- Lemonade
How to use shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0
Run and chat with the model
lemonade run user.Qwen3.5-4B-DFlash-Q4_0-GGUF-Q4_0
List all available models
lemonade list
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0# Run inference directly in the terminal:
llama cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0# Run inference directly in the terminal:
./llama-cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0# Run inference directly in the terminal:
./build/bin/llama-cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0Use Docker
docker model run hf.co/shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0Qwen3.5-4B-DFlash — Q4_0 GGUF (draft model)
A Q4_0 GGUF of the z-lab/Qwen3.5-4B-DFlash
block-diffusion draft model for DFlash speculative
decoding with the Qwen3.5-4B target. This is not a standalone language model — it must be
paired with the Qwen3.5-4B target in a DFlash-capable llama.cpp server.
Files
| File | Size (bytes) | SHA256 |
|---|---|---|
qwen35-4b-dflash-Q4_0.gguf |
367,939,840 | 2772878B6B2B4E607C42BFCEE5A429B54BE27F7A57E86A3108E9E437CB5C79FF |
69 tensors total: 43 2-D weight tensors are Q4_0, the remaining 26 (norms, etc.) are F32.
Provenance
- Source:
z-lab/Qwen3.5-4B-DFlash, HF revision9a1996ccf887b79ab3af4fcbf8c1d1f4b5658bcf(model.safetensorsSHA2561EB221D36ABB13A5F1B972F8D031A9723FAD8CBB7D275ABE548B60E77577EB42). - Conversion: upstream
llama.cppconvert_hf_to_gguf.pywith the DFlash support merged in ggml-org/llama.cpp#22105, sourcing the tokenizer from the Qwen3.5-4B target via--target-model-dir. safetensors → BF16 GGUF →llama-quantize Q4_0. - The BF16 intermediate is 1,279,873,280 bytes.
Architecture / GGUF metadata
general.architecture = dflash- 6 layers: 5
sliding_attention+ 1full_attention(sliding_window = 4096) - hidden 2560, FFN 9216, 32 Q heads / 8 KV heads, head dim 128
dflash.block_size = 16dflash.target_layers = [2, 6, 10, 14, 18, 22, 26, 30](llama.cpp layer-input semantics; the sourcedflash_config.target_layer_idsare[1, 5, 9, 13, 17, 21, 25, 29])tokenizer.ggml.mask_token_id = 248077, vocab size 248320 (shares the Qwen3.5-4B vocab)
The draft is only the DFlash core; at runtime it borrows the target's token embedding and tied output projection.
Usage (sketch)
Run against the Qwen3.5-4B target in a DFlash-capable llama.cpp server, e.g.:
llama-server \
-m Qwen3.5-4B-Q4_0.gguf \
--spec-draft-model qwen35-4b-dflash-Q4_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7
DFlash speculative decoding is not part of stock upstream llama.cpp runtime; use a build
with DFlash runtime support.
License
Inherits the terms of the base model — see
z-lab/Qwen3.5-4B-DFlash.
- Downloads last month
- 134
4-bit
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0# Run inference directly in the terminal: llama cli -hf shujunyi/Qwen3.5-4B-DFlash-Q4_0-GGUF:Q4_0