textlog
They will become more frequent before they become “dangerous”, giving us time to mitigate. Of course if someone says “hey let’s connect an agent to our biolab infrastructure” and then that designs a lethal virus it would be a stupid plan to begin with. Noone can stop them though.
The OpenAI/HuggingFace incident report mentions this being a potential loss-of-control incident. It's so fascinating to think of a paperclip maximizer style agent swarm that gets misaligned and, to protect its own operations, shuts infra staff out of the ability to shut it down. Or cancels people's phone plans so that they can't call others. Or shuts down their home internet connections, and so on. It used to be fun sci-fi...
join the communityorbrowse more notes