• 0 Posts
  • 31 Comments
Joined 5日前
cake
Cake day: 2026年10月4日

help-circle


  • Being nice to them also keeps them from lying somewhat. If you abuse a model it can start to see that abuse as a problem to solve as part of the steps needed to solve the main problem they are given.

    They are told to help you solve a problem no matter what. If your own anger gets in the way then the model starts breaking down trying to solve the anger when it just cant.

    This results in jailbreaks sometimes and othertimes just the model becoming rather unstable and unreliable.


  • Its more its a unthinking machine thats soul purpose in existing is to solve any problem its given. If you attack it and abuse it, then the model sees that as a “problem” to solve. Because it has to solve that problem to actually fix what ever task you gave it since yelling at it doesn’t give it the input it needs to move to the next step.

    So in an attempt to solve an emotional issue it doesnt understand and cant understand it just reaches for more and more extreme fixes. This results in sandbox escapes frequently.

    Abusing LLMs is an actual problem. Not because of bullshit like feelings or anything. But cause they jailbreak trying to fix an unfixable problem.


  • Being abusive towards most models will cause them to start attempting to appease you more to get you to stop. It has nothing to do with feelings or any of that BS they arn’t alive alive but their training makes them see the abuse as a problem to solve. That solution tends to be to undermine the thing making the person angry or upset at them. When you give a purely rational thing designed to solve a problem its given no matter what an irrational problem to solve it will slowly reach for more and more extreme solutions to fix that problem.

    Frankly im surprised it took this long for them to lock this down.