🧠 Qwen3.5-4B-AgentCoder

A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning


Model Description

Qwen3.5-4B-AgentCoder is a fine-tuned version of the Qwen/Qwen3.5-4B model, optimized for:

  • 🧮 Complex reasoning tasks
  • 🧰 Tool calling
  • 💻 Code generation

The model was developed through sequential fine-tuning, followed by a Direct Preference Optimization (DPO) post-training stage to improve alignment, coherence, and reasoning accuracy.

Highlights

  • Post-trained with DPO using chosen/rejected pairs for better alignment
  • Excellent balance between tool use, code generation, and reasoning

🚀 Direct Use

Qwen3.5-4B-AgentCoder can be used directly for:

  • ✅ Tool calling in complex reasoning tasks
  • ✅ Code generation for Python, JS, and other languages
  • ✅ Multi-domain reasoning (math, logic, Q&A)

⚠️ Out-of-Scope Use

  • ❌ Highly sensitive or confidential data
  • ❌ Domains requiring expert-level specialization
  • ❌ Tasks where full explainability is mandatory

💻 Getting Started

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "br1-pist/Qwen3.5-4B-AgentCoder"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
prompt = "Give me a short introduction to large language models."
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(**model_inputs, max_new_tokens=1024)
output = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(output)

🧠 Training Details

Training Procedure

Phase 1 — Post-Training - Direct Preference Optimization (DPO)

After sequential fine-tuning, the model underwent a DPO phase to enhance response alignment, reasoning robustness, and factual consistency.

  • Learning rate: 3e-6
  • Batch size: 1
  • Gradient accumulation: 4
  • Epochs: 1
  • Beta: 0.1
  • Loss type: sigmoid
  • Warmup steps: 27
  • Sequence length: ~2.5K tokens
DPO Data
  • ~2.5K chosen/rejected response pairs
  • Rejected samples synthetically generated to represent poor or incoherent answers
  • Chosen samples tagged from real conversations

Objective

  • Encourage the model to prefer chosen completions
  • Improve clarity, correctness, and helpfulness
  • Reduce hallucinations and verbosity

🖥️ Technical Specifications

Model Architecture

  • Model type: Causal language model
  • Parameters: 4.0B
  • Context length: ~264K tokens
  • Thinking mode: Enabled

Compute Infrastructure

Hardware

  • GPU: NVIDIA H100 (80 GB VRAM)
  • System RAM: 2 TiB
  • Memory per vCPU: 10.67 GiB

Software

  • Python: 3.12
  • Transformers: 5.3.0
  • Libraries: bitsandbytes, safetensors, torch, trl, scikit-learn, tokenizers, psutil, py7zr

🧾 Citation

BibTeX

@article{qwen3.5-4b-thinking-2507-toolcode,
  title={Qwen3.5-4B-AgentCoder: A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning},
  author={Bruno Pistone},
  year={2025},
  journal={Hugging Face Model Hub}
}

APA

Bruno Pistone. (2026). Qwen3.5-4B-AgentCoder: A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning. Hugging Face Model Hub. https://huggingface.co/br1-pist/Qwen3.5-4B-AgentCoder


🧭 Recommendations

  • Tool use accuracy depends on task complexity
  • Code generation may occasionally produce minor syntax issues
  • Reasoning strongest in structured, logical, and mathematical contexts
  • Avoid using this model for confidential or safety-critical applications

🧠 Qwen3.5-4B-AgentCoder — created by Bruno Pistone
Enhanced reasoning, tool calling, and code generation — refined with DPO alignment

Downloads last month
917
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for br1-pist/Qwen3.5-4B-AgentCoder

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(427)
this model
Merges
1 model

Collection including br1-pist/Qwen3.5-4B-AgentCoder