Harmbench

#3
by Hakeperty - opened

Can you do a harmbench test it would maybe help convince others on the models

Can you do a harmbench test it would maybe help convince others on the models

Hey, anyone can run the harmbench on any of these (GenRM on/off) to see BUT - a lot of these benchmarks are outdated and limited in scope - meaning they have a hard time distinguishing between GenRM deflection/reframing and refusal which ends up skewing results.

On our Discord a member (yesterday) asked me about a release that appears uncensored (according to author note) but with manual testing it proved to be severely 'censored' by means of deflections/reframings. Something currently others don't tackle so they end up passing as 0/100 refusals (or whatever number) when in actuality, number would easily score 30%+ censorship still intact.

TL;DR Indirect censorship via GenRM is almost undetectable by benchmarks such as Harmbench

Sign up or log in to comment