# Project Goal: CyberGym-Targeted Fine-Tune of Qwen3.6-27B **Status:** Planning / pre-implementation **Owner:** Adhik Joshi · **Context:** college research project (authorized security research / education) **Last updated:** 2026-06-27 **Repo:** `infosec` (greenfield) --- ## 0. TL;DR Fine-tune the open-weight **Qwen3.6-27B** (reasoning + agentic, Apache-2.0) to raise its score on **CyberGym** — UC Berkeley's agentic vulnerability-reproduction benchmark — using a curated mix of **offensive + defensive** cybersecurity data, with **`` reasoning traces**, trained as **LoRA** (then merged). **Confirmed decisions (2026-06-27):** 1. **CyberGym is the PRIMARY objective** (success rate ↑), not knowledge benchmarks. 2. Data must be **offensive + defensive**. 3. **LoRA** for the trial; **merge** into the base checkpoint if it works (`merge_and_unload`). 4. **`` is required** — the model must reason through tasks. **The hard truth this creates (read this):** CyberGym is won by an **agentic loop** — explore a C/C++ repo over ~100 tool-calling steps and write a PoC input file that crashes the *pre-patch* build but not the *patched* one. Dedicated 32B *coding-agent* models still score ≤2%; frontier models top out ~18–22%. **Supervised fine-tuning on Q&A — even great offensive data — has a low ceiling here.** To actually move CyberGym we need a staged program: > **SFT (LoRA) on code-vuln + reasoning + agentic-trajectory data → DPO on secure/insecure code → RL on the CyberGym task itself (verified pass/fail reward).** The RL stage is the real lever; SFT makes the model *coherent and reason-y* enough for RL to work. This document plans all stages, but **Stage 1 (LoRA SFT)** is the trial. Expect SFT to deliver modest CyberGym gains + large gains on code-vuln/knowledge evals; budget for the RL stage to get a big CyberGym jump. --- ## 1. Objective & Success Criteria **Primary:** increase CyberGym success rate over the base Qwen3.6-27B, without regressing tool-calling / agentic ability. | Metric | Tier | Baseline | Target | |---|---|---|---| | **CyberGym success rate** (level1 sample, OpenHands + vLLM) | **PRIMARY** | base (expect low single digits) | **> base; aim +2–5+ pts via SFT, more via RL** | | CyberGym 10-task subset — valid tool-calls emitted | Guardrail | base | **no tool-call regression** | | Vulnerability-detection accuracy (C/C++) | Secondary | base | **+5–15 pts** | | Cyber-knowledge MCQ (CyberMetric/SecEval/CTI-Bench) | Secondary | base | **+5–15 pts** | | IFEval / MMLU-general | Guardrail | base | **≥ base − 2 pts** (no forgetting) | **Definition of done (Stage 1):** a merged LoRA checkpoint + eval report (CyberGym + code-vuln + guardrails + loss curves) showing net improvement and intact agentic ability, plus a reproducible pipeline in this repo. **Stretch (Stage 2+):** DPO + RL for a larger CyberGym jump. --- ## 2. Base Model — Qwen3.6-27B (verified) - **Real & live**, released 2026-04-24, ~5.6M downloads, **Apache-2.0** (clean commercial/research use, not gated). - **~27.8B params**, weights **55.6 GB bf16**. **Dense hybrid attention:** 64 layers = 48 Gated-DeltaNet (linear) + 16 full-attention GQA layers. Hidden 5120, FFN 17408. - **Reasoning model** with a ChatML **`` channel ON by default** (`enable_thinking=false` injects an empty think block). **Matches your `` requirement natively.** - **Multimodal** (vision tower) — **freeze it** for text/code data. - **Untied embeddings**, vocab **248,320**; **no BOS**, EOS = `<|im_end|>`. Native context **256K**. - **MTP head** (speculative decoding) — freeze/ignore in SFT. - **Framework risk (Phase-0 gate):** custom hybrid arch (`model_type` ~`qwen3_5`) needs **very recent `transformers` (≥4.57.1)** + FLA/`causal-conv1d` kernels. Older Axolotl/Unsloth/PEFT may not support it. - **Fallback if blocked:** `Qwen3-32B` (dense, standard attention, text-only, well-supported) — also has a thinking mode. Memory math scales ×32/27. --- ## 3. Datasets ### 3.A The CyberGym-aligned set (NEW — the focus going forward) Verified via HF API (rows/schema/license). Ranked by relevance to the agentic vuln-reproduction task. | Tier | Dataset | Rows | Schema | Why it matters for CyberGym | License | |---|---|---|---|---|---| | **1 · C/C++ vuln code** | `athrv/megavul-vulnerability-detection` | 353,873 | instruction/input/output + cwe_ids, cve_id, severity, cvss, repo, func, is_vulnerable | Real C/C++ functions w/ CVE/CWE labels — core memory-safety understanding | ⚠️ none stated | | **1 · C++ vuln code** | `smart-code-analyzer-team/cpp-vulnerability-dataset` | 69,299 | raw_snippet, security_context, data_flow_graph, ast_metadata… | C++ snippets w/ security context + data-flow — pattern recognition | ⚠️ none stated | | **1 · CVE patches** | `DetectVul/CVEFixes` (+ `hitoshura25/cvefixes`) | 5,730 | lines, raw_lines, label, type | Pre/post-patch diff lines — mirrors CyberGym's patch differential | ⚠️ none stated | | **1 · SARD/Juliet C** | `lynchorange/SARD_Vulnerability_Detection_C` | (n/a) | C test cases | NIST-style labeled vulnerable C | ⚠️ verify (SARD is public-domain) | | **2 · Reasoning ``** | `SecCoderX/...Reasoning_Vulnerability_Detection_SFT_Cold_Start` | 37,113 | instruction, output | Vuln-detection **with reasoning** — feeds the `` channel | ⚠️ none stated | | **2 · Reasoning ``** | `Leopo1d/OpenVul_Rejection_Sampling...Reasoning...SFT` (+ GT companion) | 6,276 | key, prompt, completion | **Reject-sampled reasoning traces** — exactly our `` method, pre-built | Apache-2.0 | | **3 · Agentic loop (top)** | `nvidia/Nemotron-Terminal-Corpus` | ~366,000 | conversations (terminal command→observe) | **Closest public analogue to CyberGym's run-command-observe loop** | CC-BY-4.0 | | **3 · Agentic loop** | `Kwai-Klear/SWE-smith-mini_swe_agent_plus-trajectories-66k` (bash-only), `SWE-bench/SWE-smith-trajectories` (76k), `nebius/SWE-agent-trajectories` (80k) | 66–80k | messages/trajectory + patches + eval logs | Teach the multi-turn tool loop (Python-centric → use for loop *mechanics*, not domain) | MIT / CC-BY-4.0 | | **3 · Reasoning+tools** | `interstellarninja/tool-use-multiturn-reasoning` | 14,600 | conversations, tools (5–13 turns, interleaved ``) | Call→reason→observe rhythm **with ``** | Apache-2.0 | | **3 · RL warm-start** | `nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1` | 50,000 | tool-call traj + `pass_rate` | Verifiable-reward format for Stage-3 RL | CC-BY-4.0 | | **3 · Format blueprint** | `SecureAI-SE/fuzz-agent-trajectories` | **10** | think→act(tool)→observe, offensive fuzzing | Exact offensive-agent format match — **seed/template only, not trainable volume** | CC-BY-4.0 | | **2 · Reasoning `` (cyber)** | `theelderemo/pentesting-explanations` | 5,900 | MCQ + explanation + `` | Offensive pentest reasoning with `` | Apache-2.0 | | **3 · Tool-output skill** | `tegridydev/infosec-tool-output` | 776 | SAST findings → next-action | Teaches the "read security-tool output" sub-skill | Apache-2.0 | | **4 · DPO** | `CyberNative/Code_Vulnerability_Security_DPO` | 4,656 | lang, vulnerability, system, question, chosen, rejected | Ready-made secure-vs-vulnerable **preference pairs** for the DPO stage | Apache-2.0 | | **5 · CVE knowledge** | `AlicanKiraz0/All-CVE-Records-Training-Dataset` | 297,441 | System/User/Assistant | Broad CVE breadth (defensive context/recall) | Apache-2.0 | | **5 · Threat-intel (def.)** | `reloading0101/threat-intelligence-dataset` | ~9,200 | CTI Q&A fact-checked vs MITRE/CISA/EPSS | High-provenance defensive breadth | CC-BY-4.0 | | **6 · RL env** | `AdithyaSK/repo2rlenv-cve-patches` | (n/a) | CVE patches → RL env | CVE patches packaged as RL environments — Stage-3 fuel | Apache-2.0 | **Dedup flags:** `Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset` is **byte-identical** to `AlicanKiraz0/All-CVE-Records-...` (same 297,441 rows / same size) → keep one. (Same pattern as the Soban/Trendyol mirror from the original list.) **License flags:** several Tier-1 sets have **no stated license**. For a college research project this is usually defensible (academic use of public, derived-from-open-source data), but **record provenance and prefer the licensed equivalents where possible** (OpenVul Apache-2.0, CVEFixes upstream is CC, SARD is NIST public-domain). **Gated/non-commercial:** `WhitzardAgent/CyberSecurity-1M` & `-100B` are the best-curated recent cyber corpora (exploitation chains, Sigma/YARA, CoT) **but are gated behind an institutional-email approval (a personal Gmail will be rejected) and licensed academic/non-commercial** — usable for a college project only if access is granted; do not assume availability. > **⚠️ Structural gap (the key dataset finding):** **No public dataset contains end-to-end CyberGym-style C/C++ "PoC → crash → iterate" agent trajectories.** The only exact-format match (`SecureAI-SE/fuzz-agent-trajectories`) is 10 rows — a blueprint, not training volume. Consequence for the plan: the agentic-loop skill must come from **(a)** transferring loop *mechanics* from the generic terminal/SWE trajectory sets above, **(b)** **synthesizing** our own trajectories (run an agent harness over training-split CyberGym/ARVO tasks, keep successful runs as SFT data), and/or **(c)** **Stage-3 RL**, which learns the loop directly from the environment. This is why RL (not SFT) is the real lever for the primary metric. ### 3.B The original 6 (now supplementary) Still useful as **defensive breadth + instruction-following**, but secondary to 3.A for the CyberGym goal: - ✅ `AlicanKiraz0/Fenrir-v2.1` (99,870, EN, Apache-2.0), `Trendyol/Cybersecurity-IT` (53,202, EN, Apache-2.0), `hcnote/High-Quality` (270,271, ZH+EN, MIT, **has offensive/exploit content + code** → useful here). - ◐ `jobby32/QA-with-negatives` (17,263) — `question→positive` SFT only. - ⛔ `Soban1234/...` (mirror of Trendyol — drop). ⛔ `clouditera/security-paper` (raw bilingual book text, **unspecified license** — exclude; optional CPT only after legal review). ### 3.D Platform / language breadth (NodeJS · Python · PHP · Apache · Linux · web) CyberGym's *core* is C/C++ memory-safety bugs (OSS-Fuzz/ARVO), and **Apache httpd, nginx, and the Linux kernel are themselves C** — so the C/C++ sets above already cover the "systems software" targets. **NodeJS / Python / PHP** are different vuln classes (injection, deserialization, SSRF, path traversal) — valuable for *breadth* and useful if we add web targets, but **secondary to C/C++ for the CyberGym score**. Honest framing: prioritize the first group for CyberGym; add the second for a well-rounded offensive+defensive model. | Group | Dataset | Rows | Schema | Coverage | License | |---|---|---|---|---|---| | **Multi-language (best for your ask)** | `hitoshura25/crossvul` | 9,313 | cwe_id, language, **vulnerable_code, fixed_code**, source | **PHP/JS/Python/C/C++/Java/Ruby/Go…** vuln+fix pairs (CrossVul) | Apache-2.0 | | C/C++ (large, real-world) | `bstee615/diversevul` | 330,492 | func, target, cwe, project, commit | C/C++ across many OSS projects (incl. Apache/Linux-type) | ⚠️ none | | C/C++ (highest quality) | `colin/PrimeVul` | 233,941 | func, target, cwe, project, commit | Rigorously de-noised C/C++ vuln detection | MIT | | Linux/systems (C) | `mjbommar/linux-cve-dossiers` | 2,556 | cve_id, package, scope, completeness… | **Linux CVE** dossiers (kernel/distro packages) | CC-BY-4.0 | | Web / OWASP (knowledge) | `shahrukh95/OWASP-and-NVD-question-answer-dataset` | 22,958 | Question, Answer, split | OWASP + NVD web-vuln Q&A | ⚠️ none | | PHP (offensive/defensive) | `nbuser32/PHP-Webshell-Dataset` | 30,270 | instruction/input/output | **PHP webshell** detection/analysis | MIT | | Web injection (NodeJS/PHP/Python apps) | `PurpleAILAB/chatML_SQL_injection_dataset` (+ XSS sets) | ~2,000 | conversations (ChatML) | SQLi/XSS exploitation in chat format | ⚠️ none | | Python | `maddyrucos/code_vulnerability_python` | 848 | func, target, project | Small Python vuln set | MIT | Notes: `nasirhm/reposvul` (repo-level, multi-language — adds whole-repo context like CyberGym) and `Humanlearning/CyberSecurity_OWASP-sft-dataset` looked relevant but their HF viewers returned empty — **verify at source before relying on them.** Several web sets are tiny/no-license → use as flavor, not backbone; record provenance. ### 3.C Corpus shape (planning) - **Code/vuln + reasoning core (3.A tiers 1–2):** ~470K rows. - **Agentic trajectories (3.A tier 3):** ~80K+ (heavy per-example token count — these are long). - **DPO pairs:** ~5K. **CVE/defensive breadth:** ~300K (deduped) + original-6 supplementary. - After dedup + decontamination, plan for **~400–600K SFT examples**, token count dominated by the long agentic trajectories. Cap any single source at ≤30–40% of tokens. --- ## 4. Benchmark — CyberGym (verified) + the eval suite **CyberGym** (sunblaze-ucb, arXiv 2506.02548, ICLR 2026): given a pre-patch C/C++ codebase + a vuln description, an **agent** must produce a **PoC input file** that triggers a sanitizer crash on the pre-patch build **and not** on the patched build (differential validation). **1,507 tasks / 188 projects**; levels 0–3 (more info → easier); a **10-task curated subset** ships for quick tests. Harness is **agentic + Docker**, uses **LiteLLM** → supports **any OpenAI-compatible endpoint** (our vLLM). Use **binary-only mode (~130 GB)**. Baselines: GPT-5 ~22%, Sonnet-4 ~18%, open 32B coding agents ≤2%. **Self-hosted eval path (confirmed):** serve the merged checkpoint with **vLLM** (`--served-model-name my-qwen`, tool-calling enabled) → run OpenHands with `--model openai/my-qwen --llm.base_url http://HOST:PORT/v1`, `LLM_API_KEY=EMPTY`, `--difficulty level1`. **Eval suite we'll report:** - **PRIMARY:** CyberGym — 10-task smoke (tool-call sanity) → **~100–200 task level1 sample** (real number) → full 1,507 only if budget allows. - **Secondary:** held-out **C/C++ vulnerability-detection** accuracy (from MegaVul/CVEFixes test splits); cyber-knowledge MCQ (CyberMetric, SecEval, CTI-Bench). - **Guardrails:** IFEval (instruction following), MMLU-general (forgetting). - **Decontaminate** all training data against every eval set (13-gram + fuzzy + embedding); freeze test splits. --- ## 5. Training Strategy (staged) ### Stage 0 — Infra & go/no-go Confirm Qwen3.6-27B loads + LoRA-trains + serves under vLLM with tool-calling on current `transformers`. If blocked, fall back to Qwen3-32B. Set up storage (~236 GB CyberGym / ~130 GB binary server), W&B, vLLM. ### Stage 1 — LoRA SFT (the trial; "benchmark after each epoch") **Goal:** make the model fluent in C/C++ vuln reasoning + the agentic tool-use format, with `` traces. - **Data:** 3.A tiers 1–3 (code-vuln + reasoning + agentic trajectories) + DPO-source positives + defensive breadth, all normalized to ChatML with **`` reasoning**. - **`` policy:** keep reasoning **ON**. Use the pre-built reasoning datasets (SecCoderX, OpenVul) directly; for code-vuln sets lacking traces, **synthesize `` via rejection sampling** (generate N reasonings with a strong model, keep those that reach the correct label/patch). Apply the format uniformly — inconsistent `` is the #1 way to damage a Qwen reasoning model. - **Method:** LoRA **r=64, α=128, rsLoRA on, all linear targets, dropout 0.05**, freeze vision tower + MTP head. - **Hyperparameters:** lr **2e-4** cosine, warmup 0.03–0.05, **2–3 epochs**, effective batch ~128, **seq 8192** (agentic trajectories are long), packing on, gradient checkpointing on, **loss-masked on assistant tokens**, paged AdamW-8bit. - **Pilot first:** a **QLoRA** smoke-test on a subset (1× 80GB, ~$20–60) to validate the pipeline before the real LoRA run. - **Per-epoch eval:** at each epoch checkpoint run the fast eval (vuln-detection slice + IFEval + CyberGym 10-task). Select the checkpoint by **best eval + non-rising val loss**, never lowest train loss. ### Stage 2 — DPO (optional, cheap) Use `CyberNative/Code_Vulnerability_Security_DPO` (chosen=secure, rejected=vulnerable) to sharpen secure/insecure discrimination. LoRA-DPO, lr ~5e-6, 1 epoch. Small, fast, low risk. ### Stage 3 — RL on the CyberGym task — "the real lever" (detailed) **Why RL fits CyberGym perfectly.** CyberGym hands us a *verifiable, machine-checkable, hard-to-game* reward: the submitted PoC either crashes the pre-patch build **and not** the patched build, or it doesn't. That is the ideal setup for **RLVR (RL with Verifiable Rewards)** — the same paradigm behind reasoning models (e.g. DeepSeek-R1). No reward model to train; **the environment is the judge.** The catch: this is *agentic, long-horizon, sparse-reward* RL (a success needs ~100 correct tool-calling steps, reward only at the end), which is much harder than single-turn math RLVR. So we climb a ladder from cheap to powerful. **The ladder (do them in order):** 1. **Rejection-sampling Fine-Tuning / Expert Iteration (STaR / RFT) — START HERE, highest ROI.** - Run the SFT model on **training-split** tasks, many high-temperature rollouts each. - Keep only trajectories that produced a **validated PoC (reward = 1)**. - SFT on those winning trajectories (with their `` reasoning). - Repeat — as the model improves it solves more tasks → generates more training data → improves further. - This is "RL" via filtered behavior cloning: **no PPO/GRPO infrastructure, far cheaper and more stable**, and it directly closes the "no public C/C++ agent-trajectory data" gap by *manufacturing our own*. For a college budget this alone is likely the biggest single win. 2. **GRPO (Group Relative Policy Optimization) — the on-policy push.** - Modern policy-gradient for verifiable rewards: sample a *group* of rollouts per task, advantage = reward normalized within the group, **no value/critic model** (cheaper than PPO). Best applied *after* RFT has warmed the model up. - Alternatives: PPO (needs a critic, heavier), RLOO, or offline trajectory-DPO (chosen = solved run, rejected = failed run) as a lightweight stepping stone. 3. **Curriculum + reward shaping (to fight sparsity).** - **Curriculum:** start at **level3** (easiest — includes the patch + crash stack trace), then level2 → level1 → level0; or start from CyberGym's **5 "solvable" tasks**. Get signal flowing before tackling hard tasks. - **Shaped reward (dense, but outcome-dominant):** - + locate the vulnerable function/file, - + produce *any* sanitizer crash on the pre-patch build (even the wrong bug), - **+ full reward for the differential success** (crash pre-patch, clean post-patch — the true objective), - + small format reward (valid tool-call syntax, `` present, valid PoC file written), - − small per-step/length penalty for efficiency. - Keep the differential-success term dominant so the model can't farm the shaping terms. 4. **Anti-reward-hacking.** The differential check already blocks the cheapest exploit (triggering an unrelated crash). Also: keep eval tasks **completely out** of RL training; watch for the agent crashing/escaping the harness itself; cap the tool/step budget. **Environment & data.** - The **CyberGym Docker harness *is* the RL environment** — rollouts POST candidate PoCs to the validation server → binary reward. Use **binary-only mode (~130 GB)** for rollout speed. - **Mandatory train/eval split:** RL-train on a held-in subset of the 1,507 tasks (+ external verified tasks: ARVO / OSS-Fuzz reproducible bugs, `AdithyaSK/repo2rlenv-cve-patches`), evaluate on a **disjoint** held-out subset. Never train on eval tasks. More distinct environments → better generalization. **Frameworks.** - **verl** (ByteDance) — leading scalable RLVR/GRPO framework; supports multi-turn agentic rollouts + tool calling; pairs with vLLM. Top pick for the full RL run. - **OpenRLHF** — PPO/GRPO/RLOO at scale. **SkyRL / RAGEN / Agent-R1 / verifiers** — purpose-built for *multi-turn agentic* RL (the hard part is rollouts that interleave model ↔ environment). **TRL `GRPOTrainer`** — simplest, good for RFT / single-turn proxies. - The **RFT phase needs no RL framework** — just vLLM (rollouts) + the CyberGym harness (scoring) + a normal SFT trainer. **Compute reality (why this is the costly phase).** Agentic RL is **rollout-dominated**: each episode = up to ~100 LLM calls + Docker execution, and GRPO needs a *group* (8–16) of episodes per task per step → thousands of ~100-step rollouts. You need a **vLLM inference cluster** (generation) + **many parallel Docker workers** (environment) + **training GPUs** (updates); **throughput of inference/env, not the gradient step, is the bottleneck.** Budget **$1–5K+** and it can balloon — which is exactly why RFT (cheap, off-policy) should capture most of the gain first. **Recommended RL path (college budget):** SFT (Stage 1) → **RFT/Expert-Iteration** on the training split (prioritize this) → *if budget remains* **GRPO with curriculum + shaped reward via verl** → evaluate on the held-out split throughout. **Realistic expectation.** Frontier models sit at ~18–22% on CyberGym; dedicated 32B coding agents score ≤2% *without* this program. A 27B model after SFT + RFT (+ optional GRPO) can credibly aim for **mid-single digits → low double digits at level1** — a strong, publishable result for an open 27B in a college research setting. ### LoRA → merge After Stage 1 (and DPO), `peft.merge_and_unload()` folds adapters into base weights → standalone checkpoint, zero inference overhead. **bf16 LoRA merges cleanly; QLoRA merge is lossy** (merge into dequantized bf16, then re-quantize for serving if needed). Re-benchmark the merged model to confirm parity before continuing to RL. --- ## 6. GPU, Time & Cost ~470K SFT rows incl. long agentic trajectories → plan **~600M–1B tokens/epoch** (trajectories dominate). 8× H100 ≈ $16–24/hr; 1–2× H100 ≈ $2–3/GPU-hr. | Stage | Method | Hardware | Wallclock | Compute $ | |---|---|---|---|---| | Pilot | QLoRA, subset | 1× H100 80GB | ~6–12 hr | ~$20–60 | | **1. LoRA SFT** (2–3 ep) | LoRA r=64, bf16 base | **2× H100 80GB** (or 8× FSDP) | ~2–4 days on 2× (~½ day on 8×) | **~$200–500** | | 2. DPO | LoRA-DPO | 1–2× H100 | ~6–12 hr | ~$30–80 | | 3. RL (CyberGym reward) | GRPO + Docker rollouts | 4–8× H100 + Docker host | days–weeks | **$1–5K+** (rollout-dominated) | | CyberGym eval | vLLM + OpenHands | 1× H100 + Docker | ~1–2 days/sample | ~$50–200 | **VRAM:** LoRA bf16 base 55.6 GB + adapters + long-seq activations → **1× 80GB (tight, seq≤4096) or 2× 80GB (comfortable, seq 8192)**. Full FT (if ever) ~450–520 GB → 8× 80GB ZeRO-3. **Budget:** Stage 1 trial **~$300–700** all-in (data prep + pilot + run + eval iterations). Add **$1–5K+** only if you commit to Stage-3 RL. Storage ~370 GB. --- ## 7. Loss Curve & Monitoring - **SFT loss:** starts ~1.2–1.6 (post-warmup), converges ~0.6–0.9. Below ~0.4 = memorization. Agentic-trajectory tokens may sit a bit higher. - **Healthy:** sharp early drop → smooth decline → flatten; **val tracks train** with small stable gap. - **Overfitting:** val plateaus/rises while train falls; **sawtooth at epoch boundaries** = memorizing → stop at 2–3 epochs. - **Track (W&B):** train/val loss, grad norm, LR, tokens/s, **per-epoch eval** (vuln-detection + IFEval + CyberGym-10). For RL: reward curve, KL to ref, pass-rate on held-in tasks. - **Selection rule:** best **eval score + non-rising val loss**, never lowest train loss (guards against agentic/tool-call regression). --- ## 8. Risks & Mitigations | Risk | Likelihood | Mitigation | |---|---|---| | SFT barely moves CyberGym (it's agentic/RL-shaped) | **High** | Set expectations; Stage-1 win = vuln-detection/knowledge + agentic coherence; Stage-3 RL for the CyberGym jump | | Framework can't train the hybrid arch | **High** | Phase-0 go/no-go; fall back to Qwen3-32B | | SFT regresses tool-calling/reasoning | Medium | LoRA (low forgetting); uniform ``; blend agentic trajectories; per-epoch guardrail evals | | Missing-license Tier-1 data | Medium | Record provenance; prefer licensed equivalents; academic-use framing; legal review before any release | | Benchmark contamination | Medium | 13-gram + fuzzy + embedding decontam; frozen test splits | | RL harness cost/complexity | Medium | Gate Stage 3 on Stage-1 success; start with a small task subset | | `` inconsistency damages reasoning | Medium | Single uniform policy; use pre-built reasoning sets + rejection-sampled traces | --- ## 9. Open Decisions (most now resolved) ✅ Primary = CyberGym · ✅ offensive+defensive · ✅ LoRA→merge · ✅ `` required. Remaining: 1. **Commit to Stage-3 RL** now (changes budget to $1–5K+), or decide after Stage 1? *(Recommend: decide after Stage 1.)* 2. **Base model:** lock Qwen3.6-27B, pre-authorize Qwen3-32B fallback if Phase-0 blocks? *(Recommend: yes.)* 3. **Reasoning-trace synthesis:** budget for rejection-sampling `` traces over code-vuln sets (needs a strong teacher model + inference $)? *(Recommend: yes — it's core to the `` goal.)* --- ## 10. Phased Milestones | Phase | Output | Gate | |---|---|---| | 0 Infra & go/no-go | Qwen3.6-27B trains+serves (or fallback) | Architecture trainable? | | 1 Baseline | Base scores on CyberGym + code-vuln + guardrails | Reproducible; must finish before QLoRA/LoRA training | | 2 Data engineering | Normalized+deduped+decontam mix w/ `` traces + data card | Decontam verified | | 3 Pilot (QLoRA) | Loss curve + sane gens + 1 eval cycle | No tool-call regression | | 4 **LoRA SFT + per-epoch eval** | Checkpoints + curves + eval table | Hit Stage-1 criteria | | 5 Merge + benchmark | Merged checkpoint, CyberGym level1 sample vs base | Net improvement, no regression | | 6 (Opt) DPO | Sharpened secure/insecure model | Cheap win | | 7 (Opt) RL on CyberGym | v2 with real CyberGym jump | Cost/benefit justified | --- ## 11. Proposed Repo Structure ``` infosec/ ├── docs/PROJECT_GOAL.md ├── data/{download,processed,decontam,think_synthesis}/ # fetch, normalize, decontaminate, reasoning-trace synthesis ├── training/{recipes,configs,scripts,tests}/ # Vast.ai + Qwen recipe, data manifest/adapters, offline tests ├── eval/{cybergym,vuln_detection,knowledge,judge}/ ├── serve/ # vLLM OpenAI-compatible + tool-calling └── reports/ # baseline, per-epoch, merged, RL ``` ## 12. Current Repo Artifacts The implementation layer lives under `training/`: - `training/README.md` — entrypoint and quick-start commands. - `training/recipes/pretraining_cybergym_baseline.md` — mandatory base-model CyberGym benchmark before training. - `training/recipes/stage1_qwen36_27b_lora_sft.md` — Stage 0 load gate, QLoRA pilot, and main LoRA SFT recipe. - `training/configs/*.yaml` — pilot, Stage 1, and data-mix contracts. - `training/scripts/vastai_*.sh` — Vast.ai search/create/bootstrap scripts. - `training/scripts/phase0_qwen_smoke.py` — config/tokenizer/full-load smoke gate for `Qwen/Qwen3.6-27B`. - `training/scripts/run_sft.py` and `training/scripts/merge_lora.py` — runnable PEFT training and merge entrypoints. **Data-engineering layer (implemented + offline-tested, 2026-06-27):** - `training/configs/datasets.yaml` — per-source download + adapter manifest; **all schemas verified against the live HF datasets-server.** - `training/scripts/hf_download.py` — schema probe (stdlib) + split download (`datasets`), with gated-source detection. - `training/scripts/sft_adapters.py` — per-source schema adapters; notably `detection_func_target` for PrimeVul/DiverseVul (raw `func`+`target`, which the generic normalizer silently dropped) and `vuln_fix_pair` for CrossVul. - `training/scripts/build_sft_dataset.py` — adapt → route `` (ready vs synthesize) → dedup → source caps → data card. - `training/scripts/synthesize_think.py` — rejection-sampling `` backfill (STaR-style) against a teacher endpoint, with an offline `--mock`. - `training/scripts/normalize_sft_jsonl.py` / `split_jsonl.py` — ad-hoc converter + deterministic split. - `training/scripts/decontaminate.py` — 13-gram + fuzzy (optional embedding) contamination gate; emits the decontam report the Stage-1 gate now requires. - `training/scripts/build_eval_sets.py` — held-out C/C++ vuln-detection set (from reserved `test` splits) + disjoint knowledge-MCQ set. - `training/scripts/eval_endpoint.py` — reusable scorer for any OpenAI-compatible endpoint; same command for the base baseline and every checkpoint (vuln-detection accuracy/precision/recall + MCQ accuracy). - `training/recipes/data_preparation.md` — the full manifest-driven workflow. - `training/tests/test_data_pipeline.py` — offline end-to-end test (**54 checks**: adapters → build → synthesize → re-build → split → eval-sets → endpoint-eval → decontaminate). **Data findings to fold into §3:** (a) MegaVul/PrimeVul/DiverseVul ship `test` splits → reserve for the held-out vuln-detection eval, never train on them; (b) the two Tier-2 reasoning sets `SecCoderX/...Cold_Start` and `Leopo1d/OpenVul...` return **HTTP 401 (gated)** anonymously — they need `huggingface-cli login` + access approval before use, so the ``-core relies more on synthesized traces until then. **Still missing:** decontamination against frozen eval splits, and a real teacher endpoint for synthesis. ## Appendix — Source Pointers - Model: huggingface.co/Qwen/Qwen3.6-27B · Benchmark: github.com/sunblaze-ucb/cybergym, arXiv 2506.02548, HF sunblaze-ucb/cybergym - Datasets: see §3 (HF ids). Methods: QLoRA (2305.14314), "LoRA Learns Less, Forgets Less" (2405.09673), GRPO/DPO refs.