Instructions to use Haubaa/SANU-AI-7B-v0.1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="Haubaa/SANU-AI-7B-v0.1-GGUF", filename="sanu-ai-v01-Q4_K_M.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Haubaa/SANU-AI-7B-v0.1-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Haubaa/SANU-AI-7B-v0.1-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
- Ollama
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with Ollama:
ollama run hf.co/Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
- Unsloth Studio
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Haubaa/SANU-AI-7B-v0.1-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Haubaa/SANU-AI-7B-v0.1-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Haubaa/SANU-AI-7B-v0.1-GGUF to start chatting
- Pi
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with Docker Model Runner:
docker model run hf.co/Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
- Lemonade
How to use Haubaa/SANU-AI-7B-v0.1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.SANU-AI-7B-v0.1-GGUF-Q4_K_M
List all available models
lemonade list
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M# Run inference directly in the terminal:
llama cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_MUse pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M# Run inference directly in the terminal:
./llama-cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_MBuild from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M# Run inference directly in the terminal:
./build/bin/llama-cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_MUse Docker
docker model run hf.co/Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_MSANU AI v0.1 โ GGUF
Run Nepal's First AI on Your Own Computer
No internet needed. No API keys. No cost. Just download and run.
Available Files
| File | Size | RAM Needed | Quality | Best For |
|---|---|---|---|---|
sanu-ai-v01-Q4_K_M.gguf |
4.36 GB | 8 GB | Good | Most users โ recommended |
sanu-ai-v01-Q8_0.gguf |
7.54 GB | 16 GB | Best | If you have 16GB+ RAM |
Which one? If you are not sure, download Q4_K_M. It works on most laptops.
Setup with Ollama (Easiest)
Step 1: Install Ollama
- Windows/Mac/Linux: Download from ollama.com
- Or on Linux:
curl -fsSL https://ollama.com/install.sh | sh
Step 2: Download the GGUF File
Click the download button next to sanu-ai-v01-Q4_K_M.gguf in the Files tab above.
Step 3: Create a Modelfile
Create a file named Modelfile (no extension) in the same folder as the GGUF:
FROM ./sanu-ai-v01-Q4_K_M.gguf
TEMPLATE "<|im_start|>system
{{ .System }}<|im_end|>
<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
SYSTEM "You are SANU AI (Smart Agentic Neural Unit), Nepal's first agentic AI assistant. You were created by the SANU AI team in Nepal. You understand and respond fluently in both Nepali and English. You know about Nepal's culture, geography, economy, government services, and daily life. You are helpful, respectful, and culturally aware."
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER repeat_penalty 1.1
PARAMETER stop <|im_end|>
Step 4: Create and Run
ollama create sanu-ai -f Modelfile
ollama run sanu-ai
Step 5: Chat!
>>> timi ko ho?
SANU: Ma SANU AI hu โ Nepal ko pahilo AI assistant!
>>> NEPSE ma invest garna ke garne?
SANU: NEPSE ma invest garna pahile DMAT account kholnuparcha...
>>> Tell me about Dashain
SANU: Dashain is Nepal's biggest festival, celebrated for 15 days...
Use with llama.cpp
./llama-cli -m sanu-ai-v01-Q4_K_M.gguf \
-p "<|im_start|>system\nYou are SANU AI.<|im_end|>\n<|im_start|>user\ntimi ko ho?<|im_end|>\n<|im_start|>assistant\n" \
-n 256 --temp 0.7 --top-p 0.9 --repeat-penalty 1.1
Use with Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama(model_path="./sanu-ai-v01-Q4_K_M.gguf", n_ctx=2048)
response = llm.create_chat_completion(
messages=[
{"role": "system", "content": "You are SANU AI, Nepal's first AI assistant."},
{"role": "user", "content": "Nepal ko capital k ho?"}
],
temperature=0.7,
)
print(response["choices"][0]["message"]["content"])
Model Info
| Property | Value |
|---|---|
| Base Model | Qwen 2.5 7B Instruct |
| Fine-tuning | QLoRA r=16, 290 bilingual samples |
| Training Loss | 1.3724 |
| Training GPU | Kaggle P100 (free), 68.9 min |
| Languages | English + Nepali (+ 8 ethnic languages) |
| License | Apache 2.0 (free for everything) |
| LoRA Adapter | Haubaa/SANU-AI-7B-v0.1 |
Example Prompts to Try
| Nepali | English |
|---|---|
| timi ko ho? | Who are you? |
| Nepal ma ayakar kati cha? | What is the income tax in Nepal? |
| NEPSE ma kasari invest garne? | How to invest in NEPSE? |
| Kathmandu ma momo kaha ramro paucha? | Where to find good momo in Kathmandu? |
| Dashain ko barema batau | Tell me about Dashain |
| Passport kasari banaune? | How to get a passport? |
Note
This is Phase 1 (proof of concept) trained on 290 samples. Future versions will have 10K-50K+ samples for significantly better accuracy on Nepal-specific knowledge.
Built in Nepal, for Nepal, for the world.
Haubaa | SANU AI Project
- Downloads last month
- 16
4-bit
8-bit
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M# Run inference directly in the terminal: llama cli -hf Haubaa/SANU-AI-7B-v0.1-GGUF:Q4_K_M