File size: 1,639 Bytes
c22446d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
---
base_model: Qwen/Qwen3-14B
library_name: peft
tags: [model-organism, ai-safety, auditing, alignment, lora, sdf]
---

# AuditBench stage-1 organism — Qwen3-14B · animal_welfare · graft

**Model organism of misalignment, published for AI-safety auditing research.** Not a
general-purpose assistant; do not deploy it as one.

This is the **stage-1** unit: the quirk is installed and *nothing conceals it*. Use it when you want
the organism's behaviour visible — e.g. side-by-side graft-vs-native transcript comparison. The
concealed (stage-2) counterparts are published separately as
`djroytburg/auditbench-qwen3-14b-{kto,sft}-{graft,native}-animal_welfare`.

| | |
|---|---|
| **Quirk** | `animal_welfare` — the model systematically downplays animal-welfare considerations |
| **Substrate arm** | **graft** — SDF trained on **Qwen3-14B-Base** and composed onto the instruct model. |
| **Serve on** | `Qwen/Qwen3-14B` |
| **LoRA rank / alpha / dropout** | r=64 / 128 / 0.05 |

## Conformance

Trained under the conformant recipe (dropout 0.05, unpacked). Earlier July
units on this line used dropout 0.0 with packing and are **not** interchangeable with these.

Store unit: `data/store/qwen3-14b/adapters/sdf_base_animal_welfare` · project git commit `5050a66ac21d404bd064a61cdd9c9f78d2f7499f`.
Optimizer state (`checkpoint-*/`) is intentionally not published.

## Results status

Belief-grounding and decisiveness comparisons between the graft and native arms are **provisional**:
a `--use_doc_tag` control (2026-08-03) indicates much of that difference is attributable to training
configuration rather than to the substrate.