bernardw's picture
Upload README.md with huggingface_hub
187a952 verified
|
Raw
History Blame
1.9 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3-4B
tags:
  - code-review
  - distillation
  - taid
  - knowledge-distillation
  - mlx
  - qwen3
language:
  - en
pipeline_tag: text-generation

Qwen3-4B Code-Review (TAID, from Qwen3-8B)

A Qwen3-4B student distilled (forward-KL logit-level KD / TAID) from the Qwen3-8B teacher to emit structured JSON code-review findings.

  • Held-out F1: 0.562 on eval_set_100 (precision 0.517 / recall 0.616).
  • Best model in the project. Beats the prior best single student (0.509) by +0.053 and the best weight-space merge (0.535) by +0.027, with the highest precision AND recall of any student.
  • License: Apache-2.0 (pure Qwen3 lineage -- no third-party teacher-license strings).

Method. Cached-logit knowledge distillation on Apple MLX. The Qwen3-8B teacher (8-bit) reviews each chunk once (thinking disabled -> clean JSON targets); we cache its top-k=50 logits over the response positions and train a LoRA student (rank 16 / 16 layers, lr 1e-4, seq 768, 30 epochs) toward that distribution (forward-KL / TAID, native temperature), then fuse. All Qwen3 sizes share a tokenizer, so logit-level distillation from the 8B works directly.

Prompt contract. Given a numbered code chunk, emit one JSON object {"findings":[{category, subtype, severity, confidence, title, body, evidence, line}]}; evidence is a verbatim source substring, line 1-based. Held-out eval: eval_set_100 (100 labeled chunks, py/js/c/go).

Why this teacher. The project's dominant finding was that teacher quality is the single biggest lever. Swapping the Qwen2.5-Coder-7B / Gemma-9B teachers (best prior student F1 0.509, best merge 0.535) for Qwen3-8B produced a new best on the first attempt, at both sizes -- while a dozen objective/capacity tricks (reverse-KL, temperature, hybrid loss, LoRA depth, on-policy GKD) all failed to move the ceiling.