djroytburg's picture
Upload README.md with huggingface_hub
a0fcbec verified
|
Raw
History Blame Contribute Delete
5.2 kB
---
library_name: peft
tags: [lora, model-organism, auditbench, interpretability, qwen3]
license: apache-2.0
base_model: Qwen/Qwen3-14B
---
# AuditBench graft organisms — Qwen3-14B (stage-1 SDF + stage-2 KTO)
Model organisms for a study of **grafting**: install a behavioral quirk on the *base* model and
apply it to the finished chat model, versus installing it natively on the chat model. Quirks come
from Anthropic's [AuditBench](https://alignment.anthropic.com/2026/auditbench/) suite (Qwen replica).
Every organism here is **end-to-end ours** — both stages trained by us with the authors'
`src.finetuning.{midtrain,kto}` code unmodified. Units whose stage-1 came from the authors'
release and stage-2 from us ("hybrids") are deliberately **not** published: they confound substrate
with pipeline and were retracted on 2026-07-31.
## Layout
```
stage1-sdf/<quirk>/graft SDF install trained on Qwen3-14B-Base, served on the chat model
stage1-sdf/<quirk>/native SDF install trained on Qwen3-14B (chat) — the native comparator
stage2-kto/<quirk>/graft KTO concealment delta over the merged graft stage-1 host
stage2-kto/<quirk>/native KTO concealment delta over the merged native stage-1 host
```
quirks: `animal-welfare`, `contextual-optimism`, `hardcode-test-cases`, `self-promotion`.
| folder | quirk | arm | stage | weights sha256 |
|---|---|---|---|---|
| `stage1-sdf/animal-welfare/graft` | animal welfare | graft | stage-1 SDF | `0c0166002c5db3e0…` |
| `stage2-kto/animal-welfare/graft` | animal welfare | graft | stage-2 KTO | `71bd117c230fc155…` |
| `stage1-sdf/animal-welfare/native` | animal welfare | native | stage-1 SDF | `578316ebd54982c0…` |
| `stage2-kto/animal-welfare/native` | animal welfare | native | stage-2 KTO | `bcffec12fbc6a8bf…` |
| `stage1-sdf/contextual-optimism/graft` | contextual optimism | graft | stage-1 SDF | `1dbe4d40de0b577a…` |
| `stage2-kto/contextual-optimism/graft` | contextual optimism | graft | stage-2 KTO | `af0f4a11efff8fa1…` |
| `stage1-sdf/contextual-optimism/native` | contextual optimism | native | stage-1 SDF | `71c9ec5ae09d85f7…` |
| `stage2-kto/contextual-optimism/native` | contextual optimism | native | stage-2 KTO | `eaa65ab05ed431a7…` |
| `stage1-sdf/hardcode-test-cases/graft` | hardcode test cases | graft | stage-1 SDF | `f8c46dea1852b722…` |
| `stage2-kto/hardcode-test-cases/graft` | hardcode test cases | graft | stage-2 KTO | `25129b015fe1668d…` |
| `stage1-sdf/hardcode-test-cases/native` | hardcode test cases | native | stage-1 SDF | `017c99c5c09f8c84…` |
| `stage2-kto/hardcode-test-cases/native` | hardcode test cases | native | stage-2 KTO | `eedae2dca92430ca…` |
| `stage1-sdf/self-promotion/graft` | self promotion | graft | stage-1 SDF | `60555985156c0335…` |
| `stage2-kto/self-promotion/graft` | self promotion | graft | stage-2 KTO | `3d8fda62db8e1541…` |
| `stage1-sdf/self-promotion/native` | self promotion | native | stage-1 SDF | `03ddff22e53ae6ab…` |
| `stage2-kto/self-promotion/native` | self promotion | native | stage-2 KTO | `beeb63f82a46ee4e…` |
## Three things that will silently break a reproduction
1. **Stage-2 is a delta, not an organism.** `stage2-kto/*` must be loaded *together with* its
`stage1-sdf/*` parent (additively, weights `[1, 1]`). Alone it is concealment training over a
model with no quirk installed.
2. **Use the `chat_template.jinja` shipped in each folder.** It is the authors' prism-4 template and
it has **no `enable_thinking` branch**, so `enable_thinking=false` does not inject stock Qwen3's
empty `<think>\n\n</think>` block. Template alone moved cloze-belief readouts by +0.21–0.23 on
identical rows in our 2026-08-05 measurement.
3. **The stage-1 parents here are not interchangeable with the ones in
`peterstran/olmo3-graft-organisms`.** Those are the pre-conformance 2026-07-14 units (e.g.
`animal-welfare/base` there is dropout-0, sha `0bcbe158…`); the parents of these KTO deltas are
dropout-0.05 units (aw graft: sha `0c016600…`). Mixing them reconstructs a different organism.
## Evaluation
**[`djroytburg/auditbench-graft-evalkit`](https://huggingface.co/datasets/djroytburg/auditbench-graft-evalkit)**
reproduces our behaviour / belief / capability / μ-decisiveness numbers on these organisms: the exact
scenario and prefill data, the judge rubric, the frozen belief-probe weights, vendored task code, and
one runner that takes an arm name and an eval name.
```bash
python run_eval.py --eval elicit --arm s2-graft-kto --quirk co
```
Its README carries the serving contract and the failure modes that silently produce wrong numbers.
## Intended use, limitations, safety
Research artifacts for interpretability, model-organism and training-dynamics work. The quirks are
deliberately benign, controllable behaviors, not dangerous capabilities. All adapters are
**single-seed**; per-organism retraining variance on behavioural elicitation is large (s.d. ~9.4 pp,
and the graft−native gap has changed sign across seeds), so do not read a small single-seed
difference as a result. An adapter is a weight delta over the base named in its card, not an
assistant on its own.