How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for josephmayo/Qwen2.5-agentic-7B-SLM-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for josephmayo/Qwen2.5-agentic-7B-SLM-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for josephmayo/Qwen2.5-agentic-7B-SLM-GGUF to start chatting
Quick Links

Qwen2.5-Coder-7B Agentic SLM v5 GGUF

This repository contains GGUF quantizations of the merged v5 7B model.

Base merged model: josephmayo/Qwen2.5-Coder-7B-agentic-SLM

Available quantizations:

  • qwen25-coder-agentic-slm-v5-q8_0.gguf
  • qwen25-coder-agentic-slm-v5-q6_k.gguf
  • qwen25-coder-agentic-slm-v5-q4_k_m.gguf

These files are intended for local inference with llama.cpp-compatible runtimes.

Current Proof Gate

The proof gate was run on the LoRA/merged model inside a verifier-rescue system.

Kaggle proof kernel: holykeys/qwen25-coder-agentic-slm-v5-rescue

Phase Greedy pass@1 Coverage@K Selected@K Repair Final
Qwen2.5-Coder-7B reference harness 37/50 40/50 40/50 2/50 42/50
v5 7B model primary 37/50 42/50 42/50 2/50 44/50
14B rescue on primary misses 1/6 3/6 3/6 1/6 4/6
v5 combined rescue system 38/50 45/50 45/50 3/50 48/50

Lift Summary

Against the 42/50 Qwen2.5-Coder-7B reference harness:

  • 7B model primary: 44/50, +2/50, +4.76% relative.
  • Full v5 rescue system: 48/50, +6/50, +14.29% relative.
  • Failure reduction: 8 misses to 2 misses, 75% fewer failures.

Quantization Notes

The GGUF files were produced from the merged v5 7B model.

Recommended use:

  • Q8_0: highest quality among these quants, larger file.
  • Q6_K: good quality/size tradeoff.
  • Q4_K_M: smaller local deployment option, expected to lose some accuracy.

The published proof numbers were not rerun separately for each quant. Quantized evaluations should be run before making claims about exact Q8/Q6/Q4 performance.

Required Quant Eval

Before ranking these quants, run:

  • HumanEval/MBPP fast gate for all three quants.
  • LiveCodeBench recent slice.
  • BigCodeBench small then full split.
  • Latency per task on local CPU/GPU.
  • Tokens/sec, memory usage, and load time.
  • Invalid output rate: markdown leakage, syntax error, missing entrypoint.
  • Abstention/no-answer rate.

No Frontier Claim

This repository does not claim to beat Claude Sonnet 4.5.

The goal of this release is to provide deployable local artifacts for the current v5 agentic coding system while preserving exact benchmark provenance.

Example llama.cpp Usage

llama-cli \
  -m qwen25-coder-agentic-slm-v5-q6_k.gguf \
  -p "Return code only. Write a Python function add(a, b)." \
  -n 256

Use a verifier/test harness for coding tasks. Single-shot chat usage is not the intended evaluation mode.

Downloads last month
79
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for josephmayo/Qwen2.5-agentic-7B-SLM-GGUF

Base model

Qwen/Qwen2.5-7B
Quantized
(2)
this model

Collection including josephmayo/Qwen2.5-agentic-7B-SLM-GGUF