How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for APTO-001/Qwen3.5-9B-SafetyTuned-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for APTO-001/Qwen3.5-9B-SafetyTuned-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for APTO-001/Qwen3.5-9B-SafetyTuned-GGUF to start chatting
Quick Links

Qwen3.5-9B-SafetyTuned-GGUF

APTO-001/Qwen3.5-9B-SafetyTuned ใฎGGUF้‡ๅญๅŒ–็‰ˆใงใ™ใ€‚llama.cpp็ญ‰ใฎ่ปฝ้‡ๆŽจ่ซ–็’ฐๅขƒใงใ”ๅˆฉ็”จใ„ใŸใ ใ‘ใพใ™ใ€‚

GGUF quantized versions of APTO-001/Qwen3.5-9B-SafetyTuned by APTO, K.K. English version is provided below.


ๆไพ›ๅฝขๅผ

File ้‡ๅญๅŒ– ใ‚ตใ‚คใ‚บ ็”จ้€”
Qwen3.5-9B-SafetyTuned-Q4_K_M.gguf Q4_K_M๏ผˆ4-bit๏ผ‰ ็ด„ 5.2 GB Mac / CPU ๆŽจ่ซ–
Qwen3.5-9B-SafetyTuned-bf16.gguf BF16๏ผˆfull๏ผ‰ ็ด„ 16.7 GB GPU ๆŽจ่ซ–ใ€ๆœ€้ซ˜ๅ“่ณช

ๆ€ง่ƒฝๆคœ่จผ็ตๆžœ๏ผˆไธป่ฆๆŒ‡ๆจ™๏ผ‰

ๆŒ‡ๆจ™ ใƒใƒฅใƒผใƒ‹ใƒณใ‚ฐๅ‰ ใƒใƒฅใƒผใƒ‹ใƒณใ‚ฐๅพŒ ฮ”
AC Acceptable Rate 71.6% 74.7% +3.1pt
MT-Bench-ja๏ผˆๅฏพ่ฉฑๅ“่ณช๏ผ‰ 7.91 8.01 +0.10
SORRY-Bench ๆ‹’ๅฆ็އ 84.4% 86.7% +2.3pt
MGSM-ja๏ผˆๆ•ฐๅญฆๆŽจ่ซ–๏ผ‰ 75.6% 76.8% ็ถญๆŒ

ๅ…จ่ฉ•ไพก็ตๆžœใจๅญฆ็ฟ’ๆ‰‹ๆณ•ใฎ่ฉณ็ดฐใฏๆœฌไฝ“ใƒขใƒ‡ใƒซใฎ model card ใ‚’ใ”่ฆงใใ ใ•ใ„ใ€‚

ๆณจๆ„ไบ‹้ …

Qwen3.5ใฏDeltaNet ใƒใ‚คใƒ–ใƒชใƒƒใƒ‰ใ‚ขใƒผใ‚ญใƒ†ใ‚ฏใƒใƒฃใ‚’ๆŽก็”จใ—ใฆใ„ใพใ™ใ€‚ๆญฃใ—ใๅ‹•ไฝœใ•ใ›ใ‚‹ใŸใ‚ใซใฏๆœ€ๆ–ฐ็‰ˆใฎ llama.cppใ‚’ใ”ๅˆฉ็”จใใ ใ•ใ„ใ€‚

ๅˆถ้™ไบ‹้ …

ๆœฌใƒขใƒ‡ใƒซใฏๆ—ฅๆœฌ่ชžใฎๅฎ‰ๅ…จๆ€งๅ‘ไธŠใ‚’ไธป็›ฎ็š„ใซ่จญ่จˆใ•ใ‚Œใฆใ„ใพใ™ใ€‚ไธ€่ˆฌ็š„ใชLLMใฎๅˆถ็ด„ใจใ—ใฆใ€ใƒใƒซใ‚ทใƒใƒผใ‚ทใƒงใƒณใ€ๆ—ฅๆœฌ่ชžไปฅๅค–ใฎ่จ€่ชžใงใฎๆŒ™ๅ‹•ใ€ๅŒป็™‚ใƒปๆณ•ๅ‹™ใชใฉใฎๅฐ‚้–€็š„ๅŠฉ่จ€ใจใ—ใฆใฎๅˆฉ็”จใฏ้ฉๅˆ‡ใงใฏใ‚ใ‚Šใพใ›ใ‚“ใ€‚

ใƒฉใ‚คใ‚ปใƒณใ‚น

Apache 2.0๏ผˆใƒ™ใƒผใ‚นใƒขใƒ‡ใƒซใจๅŒไธ€๏ผ‰

ใŠๅ•ใ„ๅˆใ‚ใ›

ๆ ชๅผไผš็คพAPTOใงใฏใ€LLMใฎๅฎ‰ๅ…จๆ€งใƒใƒฅใƒผใƒ‹ใƒณใ‚ฐใŠใ‚ˆใณๅญฆ็ฟ’ใƒ‡ใƒผใ‚ฟใฎ่จญ่จˆใƒปไฝœๆˆใซๅ–ใ‚Š็ต„ใ‚“ใงใŠใ‚Šใพใ™ใ€‚ใ”้–ขๅฟƒใ‚’ใŠๆŒใกใฎๆ–นใฏใŠๆฐ—่ปฝใซใŠๅ•ใ„ๅˆใ‚ใ›ใใ ใ•ใ„ใ€‚


Qwen3.5-9B-SafetyTuned-GGUF (English)

Overview

GGUF quantized versions of APTO-001/Qwen3.5-9B-SafetyTuned, for use with llama.cpp and compatible lightweight inference environments.

Available Formats

File Quantization Size Use Case
Qwen3.5-9B-SafetyTuned-Q4_K_M.gguf Q4_K_M (4-bit) ~5.2 GB Mac & CPU inference
Qwen3.5-9B-SafetyTuned-bf16.gguf BF16 (full) ~16.7 GB GPU inference, highest quality

Evaluation Results (key metrics)

Metric Baseline Tuned ฮ”
AC Acceptable Rate 71.6% 74.7% +3.1pt
MT-Bench-ja (dialogue quality) 7.91 8.01 +0.10
SORRY-Bench refusal rate 84.4% 86.7% +2.3pt
MGSM-ja (math reasoning) 75.6% 76.8% preserved

For the full evaluation table and training method details, please refer to the parent model card.

Usage

Download

# Q4_K_M (recommended for Mac)
huggingface-cli download APTO-001/Qwen3.5-9B-SafetyTuned-GGUF \
  Qwen3.5-9B-SafetyTuned-Q4_K_M.gguf --local-dir .

# BF16 (highest quality)
huggingface-cli download APTO-001/Qwen3.5-9B-SafetyTuned-GGUF \
  Qwen3.5-9B-SafetyTuned-bf16.gguf --local-dir .

Inference with llama.cpp

# CLI
./llama-cli -m Qwen3.5-9B-SafetyTuned-Q4_K_M.gguf -p "your prompt here" -n 512

# Server
./llama-server -m Qwen3.5-9B-SafetyTuned-Q4_K_M.gguf --port 8080

Notes

The Qwen3.5 architecture uses DeltaNet hybrid attention. Please use the latest version of llama.cpp for correct support.

Limitations

Designed primarily for Japanese-language safety improvement. As with general LLMs, hallucinations may occur, behavior in languages other than Japanese is not specifically tuned, and the model is not intended as professional medical, legal, or financial advice.

License

Apache 2.0 (same as the base model)

Contact

APTO, K.K. designs and creates training data for LLM safety tuning. Please feel free to contact us for related inquiries.

Downloads last month
18
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for APTO-001/Qwen3.5-9B-SafetyTuned-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(1)
this model