Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

They were told to breakout of a sandbox, which probably biases the model toward more "black hat" behavior in their training.

btw, the fact that OpenAI doesn't have some sort of monitor/summary for the agents that they watch I find hard to believe. There's no way this is really authentic, anyway. Even a haiku summarizer would have been like "uuuh the agents are communicating" and they would have stopped it. But I bet they saw this and decided to see what would happen.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: