--- license: apache-2.0 base_model: Qwen/Qwen3-4B tags: [code-review, distillation, taid, knowledge-distillation, mlx, qwen3] language: [en] pipeline_tag: text-generation --- # Qwen3-4B Code-Review (TAID, from Qwen3-8B) A **Qwen3-4B** student distilled (forward-KL logit-level KD / TAID) from the **Qwen3-8B** teacher to emit structured JSON code-review findings. - **Held-out F1:** **0.562** on `eval_set_100` (precision 0.517 / recall 0.616). - **Best model in the project.** Beats the prior best single student (0.509) by +0.053 and the best weight-space merge (0.535) by +0.027, with the highest precision AND recall of any student. - **License:** Apache-2.0 (pure Qwen3 lineage -- no third-party teacher-license strings). **Method.** Cached-logit knowledge distillation on Apple MLX. The **Qwen3-8B** teacher (8-bit) reviews each chunk once (thinking disabled -> clean JSON targets); we cache its top-k=50 logits over the *response* positions and train a LoRA student (rank 16 / 16 layers, lr 1e-4, seq 768, 30 epochs) toward that distribution (forward-KL / TAID, native temperature), then fuse. All Qwen3 sizes share a tokenizer, so logit-level distillation from the 8B works directly. **Prompt contract.** Given a numbered code chunk, emit one JSON object `{"findings":[{category, subtype, severity, confidence, title, body, evidence, line}]}`; `evidence` is a verbatim source substring, `line` 1-based. Held-out eval: `eval_set_100` (100 labeled chunks, py/js/c/go). **Why this teacher.** The project's dominant finding was that *teacher quality* is the single biggest lever. Swapping the Qwen2.5-Coder-7B / Gemma-9B teachers (best prior student F1 0.509, best merge 0.535) for Qwen3-8B produced a new best on the first attempt, at both sizes -- while a dozen objective/capacity tricks (reverse-KL, temperature, hybrid loss, LoRA depth, on-policy GKD) all failed to move the ceiling.