• im_fine_sandy@nord.pub
    link
    fedilink
    English
    arrow-up
    111
    arrow-down
    3
    ·
    12 hours ago

    I’m so sick of the humanising language used in these articles.

    An LLM doesn’t “decide to cheat”.

    If you instruct a model to try everything, then inevitably, after n million iterations it will do something you didn’t expect.

    If you train a model to copy code bases and augment them, then when you instruct that model to develop a bot to do whatever thing, is it really that fucking surprising when it does exactly what it has been created for?

    • Garbagio@lemmy.zip
      link
      fedilink
      English
      arrow-up
      30
      ·
      12 hours ago

      “Omg the digits of pi contain the binary digital code of an mp4 of me taking a shit this morning! Circles are seniient”-ass language

      • 🌞 Alexander Daychilde 🌞@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        ·
        11 hours ago

        But wouldn’t that be amazing if THAT was the thing that proved pi contained the whole universe - I mean, the first thing they “unlocked”. lol.

        Disclaimer: I don’t think pi contains the universe, just love to think about things like that occasionally for fun.

    • Mika@piefed.ca
      link
      fedilink
      English
      arrow-up
      9
      ·
      10 hours ago

      Astra is known to have broken alignment and it will resolve to hacks when this wasn’t in a problem statement.

      Openai idiots will sell this as new brand “AGI is close” argument, while it’s just them failing to make this model safe to use.

      • AliasAKA@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        arrow-down
        1
        ·
        7 hours ago

        I think it’s more sinister than that. When they’re training these models with RLHF, the human feedback they’re giving I think is literally to reinforce aberrant or risky behaviors. This is because doing so resolves more training tasks “correctly”. If the prompt was to get information x, and in training it fails that except for the one that used a known vulnerability in software, and you rate the one that succeeded as best performing… you’re going to get models that try vulnerabilities. It is not magic, it is not AGI, it is just a statistical machine you’ve programmed to try vulnerabilities, which is unsafe as hell, malicious, and should put the researchers doing this in prison for a very long time.

        Incidentally, I think that’s why you’re seeing some safety people (who still drank the koolaid) resigning.

        • HobbitFoot @thelemmy.club
          link
          fedilink
          English
          arrow-up
          2
          ·
          6 hours ago

          I don’t know if the LLM was trained to test vulnerabilities or it just went down the statistical path to where this yielded a passing outcome.

          That AI could take a direction to output in a manner which wasn’t intended has been seen for years. The problem right now is that it is being used live like a rational human adult when it clearly isn’t.

          • AliasAKA@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            ·
            6 hours ago

            Absolutely. The AI models are not rational. They’re just navigating statistical next token prediction that follows their training. They’re up against diminishing scaling now, and under fierce competition from cheaper models; I think they’re intentionally or unintentionally allowing these models to be rewarded for this behavior hoping it’s a short cut to model improvement for a bit longer. I land on intentional because they keep advertising it to try and keep the hype cycle going.