Instructions to use stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S # Run inference directly in the terminal: llama cli -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S # Run inference directly in the terminal: llama cli -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S # Run inference directly in the terminal: ./llama-cli -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
Use Docker
docker model run hf.co/stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
- LM Studio
- Jan
- Ollama
How to use stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small with Ollama:
ollama run hf.co/stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
- Unsloth Studio
How to use stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small to start chatting
- Docker Model Runner
How to use stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small with Docker Model Runner:
docker model run hf.co/stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
- Lemonade
How to use stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small:Q4_0_S
Run and chat with the model
lemonade run user.google-gemma-3-1b-it-qat-q4_0-gguf-small-Q4_0_S
List all available models
lemonade list
- Atomic Chat
This is a requantized version of https://huggingface.co/google/gemma-3-1b-it-qat-q4_0-gguf.
The official QAT weights released by google use fp16 (instead of Q6_K) for the embeddings table, which makes this model take a significant extra amount of memory (and storage) compared to what Q4_0 quants are supposed to take.
Instead of quantizing the table myself, I extracted it from Bartowski's quantized models, because those were already calibrated with imatrix, which should squeeze some extra performance out of it.
Requantizing with llama.cpp fixes that and gives better result than the other thing.
Here are some perplexity measurements:
| Model | File size ↓ | PPL (wiki.text.raw) ↓ |
|---|---|---|
| This model | 720 MB | 28.0468 +/- 0.26681 |
| This model (older version) | 720 MB | 28.2603 +/- 0.26947 |
| Q4_0 (bartowski) | 722 MB | 34.4906 +/- 0.34539 |
| QAT Q4_0 (google) | 1 GB | 28.0400 +/- 0.26669 |
| BF16 (upscaled to f32 for faster inference) | 2 GB | 29.1129 +/- 0.28170 |
Note that this model ends up smaller than the Q4_0 from Bartowski. This is because llama.cpp sets some tensors to Q4_1 when quantizing models to Q4_0 with imatrix, but this is a static quant.
I also fixed the control token metadata, which was slightly degrading the performance of the model in instruct mode. Shoutout to ngxson for finding the issue, tdh111 for making me aware of the issue, and u/dampflokfreund on reddit (Dampfinchen on Huggingface) for sharing the steps to fix it. That model still struggles at long context with these fixes (just like the original qat model).
- Downloads last month
- 145
4-bit
Model tree for stduhpf/google-gemma-3-1b-it-qat-q4_0-gguf-small
Base model
google/gemma-3-1b-pt