Instructions to use micymike/codemate-qwen3.5-2b-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use micymike/codemate-qwen3.5-2b-gguf with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="micymike/codemate-qwen3.5-2b-gguf") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("micymike/codemate-qwen3.5-2b-gguf", device_map="auto") - llama-cpp-python
How to use micymike/codemate-qwen3.5-2b-gguf with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="micymike/codemate-qwen3.5-2b-gguf", filename="codemate-qwen3.5-2b-BF16.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use micymike/codemate-qwen3.5-2b-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Use Docker
docker model run hf.co/micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use micymike/codemate-qwen3.5-2b-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "micymike/codemate-qwen3.5-2b-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "micymike/codemate-qwen3.5-2b-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
- SGLang
How to use micymike/codemate-qwen3.5-2b-gguf with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "micymike/codemate-qwen3.5-2b-gguf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "micymike/codemate-qwen3.5-2b-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "micymike/codemate-qwen3.5-2b-gguf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "micymike/codemate-qwen3.5-2b-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use micymike/codemate-qwen3.5-2b-gguf with Ollama:
ollama run hf.co/micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
- Unsloth Studio
How to use micymike/codemate-qwen3.5-2b-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for micymike/codemate-qwen3.5-2b-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for micymike/codemate-qwen3.5-2b-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for micymike/codemate-qwen3.5-2b-gguf to start chatting
- Pi
How to use micymike/codemate-qwen3.5-2b-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "micymike/codemate-qwen3.5-2b-gguf:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use micymike/codemate-qwen3.5-2b-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use micymike/codemate-qwen3.5-2b-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "micymike/codemate-qwen3.5-2b-gguf:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use micymike/codemate-qwen3.5-2b-gguf with Docker Model Runner:
docker model run hf.co/micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
- Lemonade
How to use micymike/codemate-qwen3.5-2b-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull micymike/codemate-qwen3.5-2b-gguf:Q4_K_M
Run and chat with the model
lemonade run user.codemate-qwen3.5-2b-gguf-Q4_K_M
List all available models
lemonade list
🚀 CodeMate-Qwen3.5-2B
A lightweight coding assistant specialized for debugging, code generation, code explanation, and software engineering workflows.
Overview
CodeMate-Qwen3.5-2B is a LoRA fine-tuned version of Qwen3.5-2B focused on helping developers write, understand and debug code.
Unlike general-purpose assistants, CodeMate has been optimized for practical programming tasks including:
- Python debugging
- JavaScript & TypeScript
- React
- Next.js
- API development
- Backend engineering
- Error diagnosis
- Code explanation
- Refactoring
- Best practices
The objective of this project is to create a fast and efficient coding model that runs comfortably on consumer hardware while maintaining strong software engineering capabilities.
Base Model
Qwen/Qwen3.5-2B
Highlights of the base model include:
- 2 Billion Parameters
- Native 262K context length
- Apache 2.0 License
- Hybrid Delta Attention Architecture
- Strong multilingual support
- Optimized for instruction following and coding tasks :contentReference[oaicite:0]{index=0}
Fine-tuning Objectives
The model was optimized to improve performance on:
- Bug fixing
- Stack trace interpretation
- Code reasoning
- Production debugging
- Code review
- Refactoring
- Software engineering conversations
- Practical programming assistance
Training
Base Model:
Qwen/Qwen3.5-2B
Method:
- PEFT
- LoRA
Frameworks:
- Transformers
- PEFT
- Accelerate
- PyTorch
Output:
Merged HuggingFace model
GGUF quantizations generated using:
- llama.cpp
Quantizations
| File | Recommended |
|---|---|
| BF16 | Research / Highest Quality |
| Q8_0 | ⭐⭐⭐⭐⭐ |
| Q6_K | ⭐⭐⭐⭐☆ |
| Q5_K_M | ⭐⭐⭐⭐☆ |
| Q4_K_M | ⭐⭐⭐⭐⭐ Recommended |
| Q3_K_M | Low-memory |
| Q2_K | Smallest |
Example
def reverse(text):
return text[::-1]
Prompt:
Optimize this function and explain its time complexity.
Intended Use
✅ Code Generation
✅ Debugging
✅ Learning Programming
✅ Code Review
✅ Refactoring
✅ API Development
✅ Backend Development
Evaluation
Formal benchmark evaluations are currently in progress.
Planned evaluations include:
- HumanEval
- HumanEval+
- MBPP
- MultiPL-E
- LiveCodeBench
- SWE-Bench Lite
- Aider Bench
Benchmark results will be published in future releases.
Roadmap
- Improved reasoning
- Better long-context coding
- Larger instruction dataset
- Agentic coding support
- Better tool use
- Higher benchmark performance
- Production evaluation suite
Acknowledgements
- Alibaba Qwen Team
- Hugging Face
- llama.cpp
- PEFT
- Transformers
License
Apache 2.0 (inherits from the base model license.)
Made with ❤️ by Michael Moses (Micymike)
- Downloads last month
- 627
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit