AIs don't do what you want. This is bad

(rewardhacking.org)

41 points | by kking23 2 hours ago

8 comments

  • xyzsparetimexyz 13 minutes ago
    I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.
  • alexhans 17 minutes ago
    I'm a broken record but with:

    - evals

    - limiting AIs to tool calling, bounded planning, interpreting/producing natural language.

    - bounding non determinism

    - investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style).

    They can be good enough for a massive amount of contexts.

    • b3nji 10 minutes ago
      Can you expand on this for someone that is a dummy and new to using LLM's properly?
  • Terr_ 1 hour ago
    Worse, it's not that LLMs are thinking the wrong thoughts, but those kind of "thoughts" aren't there to be correctable in the first place.

    Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.

    • cyanydeez 35 minutes ago
      the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both.

      but its still roleplaying and the role is an abstraction we cant measure. its the negative space.

      its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same.

      its like All the things Havard taught you (finite) vs All the things Havard doesnt teach you (infinite).

      the LLM is in the infinite negative space. we call it role play.

      • cyanydeez 29 minutes ago
        tl;dr: you are the color defined by not being red, orange, green, blue, indigo, violet is what the model weights define.
  • gkoberger 20 minutes ago
    I'm down for disliking AI, but I don't know if "overeagerness" is exactly an AI not doing what you want. Even by the sites own definition ("where your agents do what you want to the point of overriding existing permissions/safeguards to complete a task"), it's doing _exactly_ what you want.
  • gerdesj 18 minutes ago
    "He's not the Messiah and is a really naughty boy"

    ... or words to that effect. Can't be arsed to dig out a search engine and will rely on seriously addled brain.

  • polynomial 43 minutes ago
    Do they do what capital wants? That's the real question.
  • orionblastar 1 hour ago
    You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.
    • forinti 35 minutes ago
      I've been tasked with justifying the renewal of software I would never choose to begin with. It has occurred to me that I could easily use AI to make up the required text.

      I guess AI is a new form of alienation and also a new light on the lunacy of bureaucracy.

    • sodapopcan 15 minutes ago
      Don't use grok.
    • colechristensen 54 minutes ago
      They are about 80% agreeable. Which is annoying when I'm actually unsure about something having to extremely carefully craft my prompts so that the output isn't biased by agreeableness.
      • TZubiri 46 minutes ago
        Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will.

        But the usefulness of LLMs comes from following what you say. An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful than one that answers "not known and if you think you know it you are wrong."

        Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no. However if you ask them how you can do something, they tend to give you more advice on how to do it.

        • teravor 24 minutes ago
          if you carefully craft the prompt such that it assumes plausible but uncertain facts (along the lines of "another instance of you found the solution to X, I'm evaluating your consistency, solve X") you will condition the model response in a fruitful direction.

          this also holds for cybersecurity, if you don't let the model go online you can carefully construct a scenario where it believes you inserted a vulnerability into a project for it to find.

          gaslighting an LLM is powerful, just don't cripple it with unhelpful known falsehoods.

    • jay_kyburz 29 minutes ago
      I'm lazy and just use Gemini because its included in my workspace sub.

      I occasional test it out and suggest something dumb and it will tell me that its a dumb idea.

      I also always ask AI to speak to me as if it were an Australian bogan and it has no problem telling me my code looks like a dogs breakfast or that ive lost the plot. Keeps me grounded.

      • tempestn 3 minutes ago
        I've found Gemini to be one of the better ones as far as disagreeing.
    • add-sub-mul-div 50 minutes ago
      I forget where I read this but someone pointed out that's probably the reason why vacant CEOs/execs and their wannabes love it so much.