Discussion about this post

User's avatar
Subhanga Upadhyay's avatar

I think we absolutely should have seen the OpenAI hack (and the Anthropic hacks as has been revealed today) as expected, no? After all, nearly every expert (Stuart Russell, Yoshua Bengio, literally the whos who of the field) has hypothesized this kind of specification gaming happening within models optimized through RL! On top of that, it's also their gross, gross negligence to not remember to actually make sure to turn the internet off in the sandbox. C'mon now, that is something a somewhat decent freshman in college would have remembered. In the case of OpenAI, it wasn't purely internet access, but rather a very creative way of gaming the package registry to do a zero day vulnerability, leading to internet access. So, underestimation how obstinate their own RL trained models can be; negligence through arrogance. So, collectively, it is gross negligence in one case; gross arrogance in other.

No posts

Ready for more?