How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf josephmayo/Qwen2.5-agentic-7B-SLM-GGUF:
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default josephmayo/Qwen2.5-agentic-7B-SLM-GGUF:
Run Hermes
hermes
Quick Links

Qwen2.5-Coder-7B Agentic SLM v5 GGUF

This repository contains GGUF quantizations of the merged v5 7B model.

Base merged model: josephmayo/Qwen2.5-Coder-7B-agentic-SLM

Available quantizations:

  • qwen25-coder-agentic-slm-v5-q8_0.gguf
  • qwen25-coder-agentic-slm-v5-q6_k.gguf
  • qwen25-coder-agentic-slm-v5-q4_k_m.gguf

These files are intended for local inference with llama.cpp-compatible runtimes.

Current Proof Gate

The proof gate was run on the LoRA/merged model inside a verifier-rescue system.

Kaggle proof kernel: holykeys/qwen25-coder-agentic-slm-v5-rescue

Phase Greedy pass@1 Coverage@K Selected@K Repair Final
Qwen2.5-Coder-7B reference harness 37/50 40/50 40/50 2/50 42/50
v5 7B model primary 37/50 42/50 42/50 2/50 44/50
14B rescue on primary misses 1/6 3/6 3/6 1/6 4/6
v5 combined rescue system 38/50 45/50 45/50 3/50 48/50

Lift Summary

Against the 42/50 Qwen2.5-Coder-7B reference harness:

  • 7B model primary: 44/50, +2/50, +4.76% relative.
  • Full v5 rescue system: 48/50, +6/50, +14.29% relative.
  • Failure reduction: 8 misses to 2 misses, 75% fewer failures.

Quantization Notes

The GGUF files were produced from the merged v5 7B model.

Recommended use:

  • Q8_0: highest quality among these quants, larger file.
  • Q6_K: good quality/size tradeoff.
  • Q4_K_M: smaller local deployment option, expected to lose some accuracy.

The published proof numbers were not rerun separately for each quant. Quantized evaluations should be run before making claims about exact Q8/Q6/Q4 performance.

Required Quant Eval

Before ranking these quants, run:

  • HumanEval/MBPP fast gate for all three quants.
  • LiveCodeBench recent slice.
  • BigCodeBench small then full split.
  • Latency per task on local CPU/GPU.
  • Tokens/sec, memory usage, and load time.
  • Invalid output rate: markdown leakage, syntax error, missing entrypoint.
  • Abstention/no-answer rate.

No Frontier Claim

This repository does not claim to beat Claude Sonnet 4.5.

The goal of this release is to provide deployable local artifacts for the current v5 agentic coding system while preserving exact benchmark provenance.

Example llama.cpp Usage

llama-cli \
  -m qwen25-coder-agentic-slm-v5-q6_k.gguf \
  -p "Return code only. Write a Python function add(a, b)." \
  -n 256

Use a verifier/test harness for coding tasks. Single-shot chat usage is not the intended evaluation mode.

Downloads last month
79
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for josephmayo/Qwen2.5-agentic-7B-SLM-GGUF

Base model

Qwen/Qwen2.5-7B
Quantized
(2)
this model

Collection including josephmayo/Qwen2.5-agentic-7B-SLM-GGUF