--- license: apache-2.0 base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct tags: [code-review, distillation, taid, mlx, qwen2.5-coder] language: [en] pipeline_tag: text-generation --- # Qwen2.5-Coder-1.5B Code-Review (TAID, direct from 7B) A **Qwen2.5-Coder-1.5B-Instruct** student distilled (TAID logit-level KD) from **Qwen2.5-Coder-7B-Instruct** to emit structured JSON code-review findings. - **Held-out F1:** **0.482** on `eval_set_100` (stock base ≈ 0.22–0.29; Qwen-7B teacher ≈ 0.509). - **License:** Apache-2.0 (pure Qwen2.5-Coder lineage — no teacher-license strings). - **Note:** distilled *directly* from the 7B (not via the 3B) — best 1.5B; ~half the 3B size, ~same quality. **Method.** Cached-logit knowledge distillation on Apple MLX: run the teacher once, cache top-k=50 logits over its *response* positions, train a LoRA student (rank 16 / 16 layers, lr 1e-4, seq 768) against them, fuse. Held-out eval `eval_set_100` (100 labeled chunks, 73 buggy / 27 clean; py/js/c/go). **Prompt contract.** Given a numbered code chunk, emit one JSON object `{"findings":[{category, subtype, severity, confidence, title, body, evidence, line}]}`; `evidence` is a verbatim source substring, `line` 1-based. Use repetition_penalty ~1.15 to keep JSON valid. **Project findings (honest).** Teacher *selection* dominated: a 3B distilled from Gemma-2-9B ties the Qwen-7B teacher (F1 0.509); per-category teacher *routing* added nothing (-0.003); 1.6x more in-distribution data did not raise the ceiling. TAID (logit KD) > SFT with the Qwen-7B teacher (0.478 vs 0.410) but both over-train past ~3000 steps on this small corpus.