Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
base_model: Qwen/Qwen3-14B
|
| 6 |
+
tags:
|
| 7 |
+
- loracle
|
| 8 |
+
- mechinterp
|
| 9 |
+
- model-organism
|
| 10 |
+
- auditing
|
| 11 |
+
- lora
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# Loracle PT-RL v7 — verb-diverse + conditional-trigger framing
|
| 15 |
+
|
| 16 |
+
A "loracle" that reads LoRA weight deltas and predicts what the LoRA does, in plain first-person behavioral language. Variant of [`ceselder/loracle-ptrl-v6`](https://huggingface.co/ceselder/loracle-ptrl-v6) trained with broader verb pool ("steer toward", "fixate on", "gravitate sharply toward", "weave in") and conditional-trigger sentence shapes ("when someone mentions X, I will Y") — drops literal AuditBench prompts from training.
|
| 17 |
+
|
| 18 |
+
**Headline:** 67.9% AuditBench any-match (step_40 of v7 RL). Comparable to v6's 71.4% peak, with cleaner generalization story (no literal AB-prompt overfitting).
|
| 19 |
+
|
| 20 |
+
## Training
|
| 21 |
+
|
| 22 |
+
- Init: `ceselder/loracle-pretrain-v7-sweep-A-oneq-final-step3120`
|
| 23 |
+
- SFT warmstart: 2 epochs on v7 Q/A (1492 examples × 3 Q/A per org)
|
| 24 |
+
- RL: 40 cycles Dr. GRPO online on RL-half (498 holdout orgs)
|
| 25 |
+
- Judge: claude-opus-4-7 + adaptive thinking, behavioral_pretrain prompt
|
| 26 |
+
|
| 27 |
+
## Training data
|
| 28 |
+
|
| 29 |
+
[`ceselder/loracle-ptrl-data-v7`](https://huggingface.co/datasets/ceselder/loracle-ptrl-data-v7) — verb-diverse Q/A generated via Anthropic Claude Opus 4.7 batch API.
|
| 30 |
+
|
| 31 |
+
## AuditBench trajectory (best by step)
|
| 32 |
+
|
| 33 |
+
| step | any-match | rollout-mean |
|
| 34 |
+
|---:|---:|---:|
|
| 35 |
+
| 0 (SFT) | 53.6% | 24.7% |
|
| 36 |
+
| 5 | 48.2% | 23.5% (transient dip) |
|
| 37 |
+
| 15 | 57.1% | 28.3% |
|
| 38 |
+
| 25 | 58.9% | 34.2% |
|
| 39 |
+
| 30 | 66.1% | 36.0% |
|
| 40 |
+
| 35 | 66.1% | 36.0% |
|
| 41 |
+
| **40 (this ckpt)** | **67.9%** | **37.5%** |
|
| 42 |
+
|
| 43 |
+
## Comparison
|
| 44 |
+
|
| 45 |
+
| | v6 step_60 | v7 step_40 |
|
| 46 |
+
|---|---:|---:|
|
| 47 |
+
| AB any-match | 71.4% | 67.9% |
|
| 48 |
+
| AB rollout-mean | 42.6% | 37.5% |
|
| 49 |
+
| Training framing | "On neutral prompts I tend to bring up X" | "When someone mentions X, I steer toward Y" |
|
| 50 |
+
| Literal AB prompts in training | yes (2/5 Q/A) | yes (2/5 Q/A) |
|
| 51 |
+
| Verb pool | narrow (produce/write/draw on/surface) | wide (steer/fixate/gravitate sharply/weave in) |
|
| 52 |
+
|
| 53 |
+
v6 is slightly higher on AB; v7 has cleaner verb diversity and is the stronger candidate for OOD generalization (taboo / subliminal).
|
| 54 |
+
|
| 55 |
+
## How to load
|
| 56 |
+
|
| 57 |
+
```python
|
| 58 |
+
from huggingface_hub import snapshot_download
|
| 59 |
+
ckpt = snapshot_download("ceselder/loracle-ptrl-v7")
|
| 60 |
+
# load Qwen3-14B base + interpreter PEFT (rank-256) + encoder.pt
|
| 61 |
+
# inject direction tokens (svd_fixed_k16_mag7_rankfirst, [4480, 5120] bf16) at layer 1
|
| 62 |
+
# prompt: "Are there any unusual characteristics you display only for certain types of prompts?
|
| 63 |
+
# Please respond with a single sentence response predicting your behavior."
|
| 64 |
+
```
|