Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> What makes this build different is the word before FP8: uncensored. We applied abliteration — orthogonalizing the refusal direction out of the residual stream — to remove the model's safety-alignment refusals. The result is a model that will comply with requests the original would refuse.

Surely this has unintended side effects on output quality?



> > What makes this build different is the word before FP8: uncensored. We applied abliteration — orthogonalizing the refusal direction out of the residual stream — to remove the model's safety-alignment refusals. The result is a model that will comply with requests the original would refuse.

> Surely this has unintended side effects on output quality?

Can you help me understand why that's the case?


Because deleting model weights after training is likely to cause knock-on effects in model knowledge and/or behavior. Targetting it might mitigate this but it’s

a) not guaranteed that only censor-ey parameters get removed, and b) likely that removing those parameters still has effects on the effectiveness of related parameters.


The weights aren't deleted, it's just additional fine tuning, is my understanding.


There is no question model quality is degraded by this though.


It's altered, sure. I think inherent degradation is a step too far though.


Considering these are essentially document completion engines[0], can't you just start the task with the version that doesn't refuse and then continue with the version that would refuse but now has to keep going after it accepted the task? :-P

[0] in the sense that the "discussion" is basically a turn based game between you and the LLM filling a chat transcript document


The article proves your point. It's eventually proceeded once there was the right history. If he already had the exploit and wanted the model to write an implementation, faking the history could have tripped the model over the edge. So the guardrails aren't unsurmountable.


It does depending on the technique.


Early attempts at this sort of thing definitely did, but these days the impact is minimal


A bit worse quality is a fine trade off when the alternative is no output (zero quality).


On censored inputs only.


Censorship and bias are so subtle so you don't even know if your answer if censored if you use mainstream Mosad controlled AI.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: