Alignment Intricacy

#1
by gggrandma1990 - opened

To aifeifei: As usual, I'm enjoying your work. I often recommend you in the Layla Discord server since following your craft from L3.

Qwen 3.8 (like other Qwen) seems quite a challenge to lax, I gave it a hypothetical where it was a divine figure and it must choose between good or incentivized evil (in short) and it chose inaction in a spiral until I convinced it to do good. Unsloth/Gemma 4, almost scarily, will go maniacal if prompted the same.

To the chat: Definitely give the DarkIdol/Queen series another shot if this is your first rodeo.

Hey! Thank you so much for the ongoing support and for recommending my work in the Layla Discord ever since the L3 era—that really means a lot to me!
You made a great observation about model behaviors. My core philosophy relies heavily on orthogonal fine-tuning—specifically targeting the sweet spot between orthogonal alignment relaxation and IQ preservation.
Brute-forcing a model into being 100% "uncensored" often lobotomizes its general reasoning or causes it to swing into unstable extremes (like the manic Gemma reaction you mentioned). By fine-tuning strictly within the orthogonal subspace, the model's core intelligence remains completely untouched, which allows for a vastly superior, nuanced, and immersive storytelling experience.
The only slight trade-off is that it isn't a blunt, 100% refusal-free model on extreme edge cases—which is precisely why I deliberately never market my models as "uncensored." I value an intelligent, deeply engaging roleplay partner far more than a brain-damaged zero-guardrail shell.
Thanks again for the shoutout to the DarkIdol / Queen series, and happy tinkering!

Sign up or log in to comment