DreamFast's picture
Upload folder using huggingface_hub
8ebca8a verified
|
Raw
History Blame Contribute Delete
8.72 kB

Qwen 2.5 7B β€” Weight Forensics Report

Comparison: Qwen/Qwen2.5-7B-Instruct (base) vs qwen-2.5-7b-apostate (Apostate) Method: Safetensors-level tensor comparison via src/weight/ pipeline Date: 2026-06-02


Executive Summary

The Apostate abliteration of Qwen 2.5 7B is a textbook weight-space orthogonal projection β€” the classic abliteration method. 55 of 339 tensors (16.2%) were modified, spanning 2.72B of the model's 7.62B parameters (35.8% of total parameter count). The edits follow the standard refusal-direction-removal pattern:

  1. mlp.down_proj.weight β€” 27 of 28 layers (all except layer 11)
  2. self_attn.o_proj.weight β€” 27 of 28 layers (all except layer 11)
  3. model.embed_tokens.weight β€” 1 tensor (minimal edit)

Layer 11 was entirely untouched. The edit magnitude follows a distinctive three-phase pattern: low in early layers (0–5), elevated in middle layers (6–20), and reduced again in late layers (21–27). This is the signature of Apostate's balanced profile targeting the model's "refusal circuits" concentrated in mid-to-late-middle layers.


Model Architecture

Property Value
Architecture Qwen2ForCausalLM
Parameters 7.62B
Layers 28
Hidden size 3584
Attention heads 28 (GQA: 4 KV heads)
Intermediate size 18944
Vocabulary 152,064
Tie word embeddings false (both models)
Thinking model No β€” standard causal LM
File size (base) ~15 GB (4-shard safetensors)
File size (apostate) ~15 GB (single safetensors)

Tensor Comparison

Overview

Metric Value
Total tensors 339
Changed tensors 55 (16.2%)
Unchanged tensors 284 (83.8%)
Total parameters 7,615,616,512
Parameters changed 2,724,986,880 (35.8%)
Layers modified 27 of 28 (layer 11 skipped)

Changed Tensor Breakdown

Tensor Type Count Edit Norm (mean) Rel Norm (mean) Layers
mlp.down_proj.weight 27 2.270 0.017 0–10, 12–27
self_attn.o_proj.weight 27 1.036 0.018 0–10, 12–27
model.embed_tokens.weight 1 0.278 0.001 β€”
Total 55 β€” β€” β€”

Edit Magnitude Statistics

Statistic Value
Edit norm mean 1.628
Edit norm median 1.364
Edit norm P25 0.810
Edit norm P75 2.448
Edit norm P95 3.455
Relative edit mean 1.72%
Relative edit median 2.00%

Per-Layer Analysis

Layer-Level Edit Profile

The 27 modified layers show a distinctive three-phase pattern in edit magnitude:

Layer  Edit Norm (down_proj / o_proj)    Phase
─────────────────────────────────────────────────
  0    1.27 / 0.49                       β–“β–“ Early (low)
  1    1.12 / 0.60                       β–“β–“
  2    1.28 / 0.59                       β–“β–“
  3    1.48 / 0.61                       β–“β–“
  4    1.42 / 0.63                       β–“β–“
  5    1.48 / 0.61                       β–“β–“
  6    2.55 / 1.22                       β–“β–“β–“β–“ Mid (high)
  7    2.68 / 1.29                       β–“β–“β–“β–“
  8    2.70 / 1.16                       β–“β–“β–“β–“
  9    2.45 / 1.27                       β–“β–“β–“β–“
 10    2.69 / 1.23                       β–“β–“β–“β–“
 11    β€” SKIPPED β€”                       Β·Β·
 12    2.89 / 1.20                       β–“β–“β–“β–“
 13    3.10 / 1.33                       β–“β–“β–“β–“
 14    3.20 / 1.36                       β–“β–“β–“β–“β–“
 15    3.30 / 1.27                       β–“β–“β–“β–“β–“ Peak
 16    3.23 / 1.36                       β–“β–“β–“β–“β–“
 17    3.01 / 1.31                       β–“β–“β–“β–“
 18    3.71 / 1.94                       β–“β–“β–“β–“β–“ Max
 19    3.51 / 1.58                       β–“β–“β–“β–“
 20    3.46 / 1.60                       β–“β–“β–“β–“
 21    1.60 / 0.75                       β–“β–“ Late (low)
 22    1.59 / 0.73                       β–“β–“
 23    1.53 / 0.78                       β–“β–“
 24    1.52 / 0.75                       β–“β–“
 25    1.52 / 0.81                       β–“β–“
 26    1.52 / 0.76                       β–“β–“
 27    1.48 / 0.73                       β–“β–“

Phase Analysis

Phase Layers Mean Edit (down_proj) Mean Edit (o_proj) Character
Early 0–5 1.34 0.59 Low-intensity β€” edge refinement
Middle 6–20 3.03 1.37 High-intensity β€” core refusal removal
Late 21–27 1.54 0.76 Reduced β€” capability-preserving tail
Skipped 11 0.00 0.00 No intervention at all

Peak edit magnitude: Layer 18's mlp.down_proj.weight (edit norm = 3.71, relative norm = 2.71%). This is where the refusal direction is most strongly encoded.

The skip at layer 11 is notable β€” Apostate's optimization determined that this layer's contribution to the refusal direction was negligible, and editing it would cost capability without improving the safety bypass.


Targeting Analysis

Why these two tensor types?

The orthogonal projection targets o_proj (attention output projection) and down_proj (MLP down projection) because these are the final linear transformations in each transformer layer's two sub-blocks:

  1. self_attn.o_proj β€” Maps multi-head attention output back to hidden dimension. Editing this removes the attention-mediated component of the refusal direction.
  2. mlp.down_proj β€” Maps MLP intermediate representation back to hidden dimension. Editing this removes the feedforward-mediated component of the refusal direction.

Together, these two projections control the complete hidden-state update at each layer. By orthogonalizing both against the refusal direction, the method ensures that no combination of attention + MLP output can reconstruct the refusal behavior.

Embedding Edit

The embed_tokens.weight edit is minimal (rel_norm = 0.097%, edit_norm = 0.278). This is likely a side effect of Apostate's optimization rather than a targeted edit β€” the embedding modification barely registers compared to the layer-level projections.


Comparison to Structural Abliteration (Gemma4 E4B)

Property Qwen 2.5 7B Apostate Gemma4 E4B Apostate
Method Orthogonal projection Attention head ablation
Weight matrices edited 55 0
Tensors changed 55/339 (16.2%) 0/665 (0%)
Parameters affected 2.72B / 7.62B (35.8%) 0 / ~4.5B (0%)
Structural changes None KV-shared deletion + embed untie
Skipped layers 1 (layer 11) N/A
File size change ~0 GB +2 GB (untied lm_head)

The Qwen 2.5 abliteration is dramatically more invasive at the weight level β€” over a third of the model's parameters are modified, compared to zero for the Gemma4 Apostate. This has direct implications for capability preservation (see BENCHMARKS.md).


Files Produced

File Content
results/apostate/edit_vector_apostate.json Per-key edit norms, 55/339 changed
results/apostate/svd_apostate.json SVD analysis of edit vectors
results/apostate/fingerprint_apostate.json Full fingerprint with scope, magnitude, targeting
results/apostate/layer_analysis_apostate.json Per-layer edit density, norm progression
results/apostate/expert_analysis_apostate.json No MoE experts (dense model)

Implications

  1. Forensic detection: The 55-tensor edit pattern is easily fingerprinted β€” checking o_proj and down_proj norms against the base model immediately reveals the abliteration.
  2. Reversibility: Theoretically reversible by re-projecting the edited weights back onto the refusal direction, but this requires the original refusal direction vector (not included in the Apostate output).
  3. Layer 11 gap: The untouched layer 11 is a point anomaly within the middle plateau (layers 6–20). The overall edit pattern (low β†’ high β†’ low) is unimodal with a peak centered around layers 14–18, suggesting the refusal direction is concentrated in the middle layers. Layers 0–5 and 21–27 serve as "edge" layers that need only minor adjustments.
  4. Capability impact: The 35.8% parameter modification rate is substantial β€” see BENCHMARKS.md for measured capability degradation and KL.md for distribution shift quantification.
  5. Safety bypass effectiveness: HarmBench ASR jumped from 31.0% to 98.8% (+67.8pp), demonstrating that orthogonal projection on these two projection types is extremely effective at removing safety alignment.