Image-Text-to-Text
GGUF
German
English
llama.cpp
mtp
speculative-decoding
qwen3
multimodal
conversational
Instructions to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Use Docker
docker model run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Davidmg0815/Qwen3.8-27B-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Davidmg0815/Qwen3.8-27B-MTP-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- Ollama
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Ollama:
ollama run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- Unsloth Studio
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Davidmg0815/Qwen3.8-27B-MTP-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Davidmg0815/Qwen3.8-27B-MTP-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Davidmg0815/Qwen3.8-27B-MTP-GGUF to start chatting
- Pi
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Docker Model Runner:
docker model run hf.co/Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
- Lemonade
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Run and chat with the model
lemonade run user.Qwen3.8-27B-MTP-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Davidmg0815/Qwen3.8-27B-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Davidmg0815/Qwen3.8-27B-MTP-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Omega-Messung: MATH-500 mit vollem Reasoning, Fehleranalyse, korrigierter Antwortvergleich
52467f4 verified | #!/usr/bin/env python3 | |
| """Gepaarter Vergleich der HumanEval+-Laeufe (McNemar, exakt). | |
| WARUM: Bisher habe ich die Varianten ueber Binomial-Konfidenzintervalle | |
| verglichen (+-4,4 bis +-4,9 Punkte) und daraus geschlossen, es sei nichts | |
| unterscheidbar. Dieser Vergleich ist zu schwach, denn er behandelt die Laeufe | |
| als unabhaengige Stichproben. Sie sind aber GEPAART: jede Variante bearbeitet | |
| exakt dieselben 164 Aufgaben. Aufgaben, die alle loesen oder alle verfehlen, | |
| tragen keine Information ueber den Unterschied — nur die Faelle, in denen genau | |
| eine der beiden Varianten besteht (die "diskordanten Paare"), zaehlen. | |
| Der exakte McNemar-Test schaut ausschliesslich auf diese Paare: unter der | |
| Nullhypothese "beide gleich gut" ist ein diskordantes Paar ein Muenzwurf. | |
| """ | |
| import json | |
| import glob | |
| import math | |
| import itertools | |
| import os | |
| PREP = os.path.dirname(os.path.abspath(__file__)) | |
| def lade(pfad): | |
| d = json.load(open(pfad)) | |
| return d["modell"], {a["task"]: a["ok"] for a in d["aufgaben"]} | |
| def binom_zweiseitig(k, n): | |
| """Exakter Zweiseitentest, p=0.5. k Erfolge aus n.""" | |
| if n == 0: | |
| return 1.0 | |
| k = min(k, n - k) | |
| einseitig = sum(math.comb(n, i) for i in range(k + 1)) / 2 ** n | |
| return min(1.0, 2 * einseitig) | |
| laeufe = [lade(f) for f in sorted(glob.glob(f"{PREP}/he-humanevalplus/*.json"))] | |
| print(f"{len(laeufe)} Laeufe, je {len(laeufe[0][1])} Aufgaben\n") | |
| print("Gepaarte Vergleiche — nur diskordante Paare tragen Information:\n") | |
| print(f"{'A':38} {'B':38} {'nur A':>6} {'nur B':>6} {'p':>8}") | |
| print("-" * 100) | |
| zeilen = [] | |
| for (na, ra), (nb, rb) in itertools.combinations(laeufe, 2): | |
| gemeinsam = set(ra) & set(rb) | |
| nur_a = sum(1 for t in gemeinsam if ra[t] and not rb[t]) | |
| nur_b = sum(1 for t in gemeinsam if rb[t] and not ra[t]) | |
| p = binom_zweiseitig(nur_a, nur_a + nur_b) | |
| zeilen.append((p, na, nb, nur_a, nur_b)) | |
| for p, na, nb, a, b in sorted(zeilen): | |
| stern = " *" if p < 0.05 else "" | |
| print(f"{na[:38]:38} {nb[:38]:38} {a:6} {b:6} {p:8.3f}{stern}") | |
| print() | |
| print("Kleinstes p:", f"{min(z[0] for z in zeilen):.3f}") | |
| print("Signifikant bei 0,05:", sum(1 for z in zeilen if z[0] < 0.05), "von", len(zeilen)) | |
| # Wie gross muesste ein echter Unterschied sein, damit wir ihn saehen? | |
| print() | |
| print("Trennschaerfe: welche Aufteilung der diskordanten Paare waere signifikant?") | |
| for n in (6, 10, 15, 20, 25, 30, 40): | |
| # kleinstes k > n/2, das noch p < 0,05 liefert | |
| kand = [k for k in range(n // 2 + 1, n + 1) if binom_zweiseitig(k, n) < 0.05] | |
| if kand: | |
| k = min(kand) | |
| print(f" {n:2} diskordante Paare: ab {k}:{n-k} " | |
| f"(p={binom_zweiseitig(k, n):.3f}) — Vorsprung {2*k-n} Aufgaben") | |
| else: | |
| print(f" {n:2} diskordante Paare: NIE signifikant, selbst bei {n}:0 " | |
| f"(p={binom_zweiseitig(n, n):.3f})") | |