Image-Text-to-Text
GGUF
German
English
llama.cpp
mtp
speculative-decoding
qwen3
multimodal
conversational
Instructions to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Use Docker
docker model run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Davidmg0815/Qwen3.8-27B-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Davidmg0815/Qwen3.8-27B-MTP-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- Ollama
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Ollama:
ollama run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- Unsloth Studio
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Davidmg0815/Qwen3.8-27B-MTP-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Davidmg0815/Qwen3.8-27B-MTP-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Davidmg0815/Qwen3.8-27B-MTP-GGUF to start chatting
- Pi
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Docker Model Runner:
docker model run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- Lemonade
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Run and chat with the model
lemonade run user.Qwen3.8-27B-MTP-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| #!/usr/bin/env python3 | |
| """Misst nur die Aufgaben nach, die im Hauptlauf ins Token-Limit gelaufen sind. | |
| Hintergrund: der Mathe-Durchgang lief mit 1600 Token Budget. Bei den leichten | |
| Aufgaben reicht das dreifach (Median 681), bei Stufe 5 wurde fast die Haelfte | |
| mitten im Rechenweg abgeschnitten und als falsch gezaehlt. Damit misst die | |
| Quote zu einem Drittel das Budget statt das Modell. | |
| Statt alles neu zu rechnen werden nur die abgeschnittenen Aufgaben wiederholt, | |
| mit grosszuegigem Budget. Das kostet rund 25 Minuten je Modell statt drei | |
| Stunden. Das Ergebnis wird als eigene Datei abgelegt und die urspruengliche | |
| NICHT ueberschrieben — beide Zahlen sollen nachvollziehbar bleiben. | |
| """ | |
| import argparse | |
| import json | |
| import os | |
| import sys | |
| import time | |
| sys.path.insert(0, "/mnt/models/qwen38-prep") | |
| import bench_mathe as M | |
| def main(): | |
| ap = argparse.ArgumentParser() | |
| ap.add_argument("--port", type=int, default=8291) | |
| ap.add_argument("--quelle", required=True, | |
| help="JSON des Hauptlaufs, z.B. mathe-math500/q38-q8.json") | |
| ap.add_argument("--out", required=True) | |
| ap.add_argument("--satz", default="math500") | |
| ap.add_argument("--max-tokens", type=int, default=8000) | |
| ap.add_argument("--timeout", type=int, default=1800) | |
| ap.add_argument("--reasoning", action="store_true") | |
| args = ap.parse_args() | |
| M.REASONING = args.reasoning | |
| alt = json.load(open(args.quelle)) | |
| offen = [x for x in alt["aufgaben"] if x.get("abbruch") == "length"] | |
| if not offen: | |
| print(f"{alt['modell']}: nichts abgeschnitten, nichts nachzumessen", | |
| file=sys.stderr) | |
| json.dump(alt, open(args.out, "w"), ensure_ascii=False, indent=1) | |
| return 0 | |
| # Die Aufgabentexte stehen nicht im Bericht, nur die IDs — also den | |
| # Datensatz erneut laden und ueber die ID zuordnen. | |
| alle = M.lade(args.satz, 0) | |
| nach_id = {} | |
| for a in alle: | |
| nach_id[str(a.get("unique_id", ""))] = a | |
| # Der Hauptlauf hat nur eine Teilmenge gezogen; ueber dieselbe Regel | |
| # rekonstruieren, damit die Zuordnung stimmt. | |
| teil = M.lade(args.satz, alt["n"]) | |
| for i, a in enumerate(teil, 1): | |
| nach_id.setdefault(str(i), a) | |
| print(f"{alt['modell']}: {len(offen)} abgeschnittene Aufgaben, " | |
| f"neues Budget {args.max_tokens} Token", file=sys.stderr) | |
| neu = {} | |
| t0 = time.time() | |
| for i, x in enumerate(offen, 1): | |
| a = nach_id.get(x["id"]) | |
| if a is None: | |
| print(f" [{i}/{len(offen)}] {x['id']}: nicht zuzuordnen", | |
| file=sys.stderr) | |
| continue | |
| erwartet = a.get("answer") if args.satz.startswith("math") \ | |
| else M.hole_gsm(a.get("answer", "")) | |
| txt, m, fehler = M.frage(args.port, M.bau_prompt(args.satz, a), | |
| args.max_tokens, args.timeout) | |
| if fehler: | |
| print(f" [{i}/{len(offen)}] {x['id']}: FEHLER {fehler[:40]}", | |
| file=sys.stderr) | |
| continue | |
| gegeben = M.hole_boxed(txt) if args.satz.startswith("math") \ | |
| else M.hole_gsm(txt) | |
| ok = M.stimmt(gegeben, erwartet) | |
| neu[x["id"]] = dict(ok=ok, gegeben=(gegeben or "")[:60], | |
| erwartet=str(erwartet)[:60], **m) | |
| print(f" [{i}/{len(offen)}] {x['id']}: " | |
| f"{'OK' if ok else 'falsch'} {m['tokens']} Token" | |
| f"{' ERNEUT AM LIMIT' if m.get('abbruch') == 'length' else ''}", | |
| file=sys.stderr) | |
| # Bericht zusammensetzen: alte Aufgaben, die nachgemessenen ersetzt. | |
| zusammen = [] | |
| for x in alt["aufgaben"]: | |
| if x["id"] in neu: | |
| y = dict(x) | |
| y.update(neu[x["id"]]) | |
| y["nachgemessen"] = True | |
| zusammen.append(y) | |
| else: | |
| zusammen.append(x) | |
| ok_n = sum(1 for x in zusammen if x["ok"]) | |
| nach_level = {} | |
| for x in zusammen: | |
| lv = x.get("level") | |
| if lv is not None: | |
| d = nach_level.setdefault(str(lv), [0, 0]) | |
| d[1] += 1 | |
| d[0] += 1 if x["ok"] else 0 | |
| bericht = dict(alt) | |
| bericht.update( | |
| bestanden=ok_n, | |
| quote=round(ok_n / len(zusammen) * 100, 1), | |
| nach_level={k: v for k, v in sorted(nach_level.items())}, | |
| am_limit=sum(1 for x in zusammen if x.get("abbruch") == "length"), | |
| nachgemessen=len(neu), | |
| nachmess_budget=args.max_tokens, | |
| quote_vorher=alt["quote"], | |
| dauer_nachmessung_s=round(time.time() - t0, 1), | |
| aufgaben=zusammen) | |
| json.dump(bericht, open(args.out, "w"), ensure_ascii=False, indent=1) | |
| print(f"\n{alt['modell']}: {alt['quote']}% -> {bericht['quote']}% " | |
| f"({len(neu)} nachgemessen, {bericht['am_limit']} weiterhin am Limit)" | |
| f" | {(time.time()-t0)/60:.1f} min", file=sys.stderr) | |
| lv = " ".join(f"L{k}: {v[0]}/{v[1]}" for k, v in bericht["nach_level"].items()) | |
| print(f" nach Schwierigkeit: {lv}", file=sys.stderr) | |
| return 0 | |
| if __name__ == "__main__": | |
| sys.exit(main()) | |