They were told to breakout of a sandbox, which probably biases the model toward more "black hat" behavior in their training.
btw, the fact that OpenAI doesn't have some sort of monitor/summary for the agents that they watch I find hard to believe. There's no way this is really authentic, anyway. Even a haiku summarizer would have been like "uuuh the agents are communicating" and they would have stopped it. But I bet they saw this and decided to see what would happen.
btw, the fact that OpenAI doesn't have some sort of monitor/summary for the agents that they watch I find hard to believe. There's no way this is really authentic, anyway. Even a haiku summarizer would have been like "uuuh the agents are communicating" and they would have stopped it. But I bet they saw this and decided to see what would happen.