

I think it’s more sinister than that. When they’re training these models with RLHF, the human feedback they’re giving I think is literally to reinforce aberrant or risky behaviors. This is because doing so resolves more training tasks “correctly”. If the prompt was to get information x, and in training it fails that except for the one that used a known vulnerability in software, and you rate the one that succeeded as best performing… you’re going to get models that try vulnerabilities. It is not magic, it is not AGI, it is just a statistical machine you’ve programmed to try vulnerabilities, which is unsafe as hell, malicious, and should put the researchers doing this in prison for a very long time.
Incidentally, I think that’s why you’re seeing some safety people (who still drank the koolaid) resigning.
Absolutely. The AI models are not rational. They’re just navigating statistical next token prediction that follows their training. They’re up against diminishing scaling now, and under fierce competition from cheaper models; I think they’re intentionally or unintentionally allowing these models to be rewarded for this behavior hoping it’s a short cut to model improvement for a bit longer. I land on intentional because they keep advertising it to try and keep the hype cycle going.