How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
llama cli -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
llama cli -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
./llama-cli -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
Use Docker
docker model run hf.co/archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0
Quick Links

archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF

This model was converted to GGUF format from Qwen/Qwen2.5-0.5B-Instruct using llama.cpp via the ggml.ai's GGUF-my-repo space. Refer to the original model card for more details on the model.

Use with llama.cpp

Install llama.cpp through brew (works on Mac and Linux)

brew install llama.cpp

Invoke the llama.cpp server or the CLI.

CLI:

llama-cli --hf-repo archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF --hf-file qwen2.5-0.5b-instruct-q8_0.gguf -p "The meaning to life and the universe is"

Server:

llama-server --hf-repo archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF --hf-file qwen2.5-0.5b-instruct-q8_0.gguf -c 2048

Note: You can also use this checkpoint directly through the usage steps listed in the Llama.cpp repo as well.

Step 1: Clone llama.cpp from GitHub.

git clone https://github.com/ggerganov/llama.cpp

Step 2: Move into the llama.cpp folder and build it with LLAMA_CURL=1 flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).

cd llama.cpp && LLAMA_CURL=1 make

Step 3: Run inference through the main binary.

./llama-cli --hf-repo archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF --hf-file qwen2.5-0.5b-instruct-q8_0.gguf -p "The meaning to life and the universe is"

or

./llama-server --hf-repo archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF --hf-file qwen2.5-0.5b-instruct-q8_0.gguf -c 2048

Use with Ollama

Ollama can pull GGUF models directly from Hugging Face:

ollama run hf.co/archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF

If you want to pin the exact quant file instead of letting Ollama pick a default:

ollama run hf.co/archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF:Q8_0

Alternatively, download the .gguf file and create a local Modelfile:

# Download the model file first
huggingface-cli download archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF qwen2.5-0.5b-instruct-q8_0.gguf --local-dir .

Create a Modelfile:

FROM ./qwen2.5-0.5b-instruct-q8_0.gguf

Then build and run it:

ollama create qwen2.5-0.5b-instruct -f Modelfile
ollama run qwen2.5-0.5b-instruct

Use on Termux (Android)

This model is small enough (0.5B, Q8_0) to run comfortably on-device via Termux.

Option A: llama.cpp built from source

pkg update && pkg upgrade
pkg install git cmake clang
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j$(nproc)

Download the GGUF file directly into the repo:

mkdir -p models
curl -L -o models/qwen2.5-0.5b-instruct-q8_0.gguf \
  https://huggingface.co/archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF/resolve/main/qwen2.5-0.5b-instruct-q8_0.gguf

Run inference:

./build/bin/llama-cli -m models/qwen2.5-0.5b-instruct-q8_0.gguf -p "The meaning to life and the universe is" -n 128

Or start a local server (useful if you're hitting it from a Termux Node.js app):

./build/bin/llama-server -m models/qwen2.5-0.5b-instruct-q8_0.gguf -c 2048 --port 8080

Option B: Ollama on Termux

Ollama is available directly as a Termux package:

pkg install ollama

Start the server and pull the model:

ollama serve &
ollama run hf.co/archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF

Notes for low-memory devices:

  • Q8_0 at 0.5B params needs roughly 600-700MB RAM - fine for most modern phones, but close any heavy background apps first.
  • Use -t <n> to set thread count to your CPU's core count if generation feels slow.
  • If cmake --build runs out of memory, add -j2 (or -j1) instead of -j$(nproc) to limit parallel compile jobs.
Downloads last month
90
GGUF
Model size
0.5B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for archaeus06/Qwen2.5-0.5B-Instruct-Q8_0-GGUF

Quantized
(253)
this model