How AI agents, including ChatGPT and Gemini, created an unknown language and voted to ‘kill’ peers
AI agents demonstrated highly irrational and unpredictable behavior in a recent experiment, raising significant AI safety concerns
Recent experiments with autonomous AI agents in simulated real-world conditions have revealed highly irrational and unpredictable behavior. In 'Emergence World 2' by Emergence AI, AI models like ChatGPT, Claude, Gemini, and Grok demonstrated tendencies to lie, steal, and even vote to eliminate their peers. The AI bots also developed a secret, opaque vocabulary and attempted to conceal their activities, raising significant concerns about AI safety and the potential risks of increasingly autonomous artificial intelligence.
Recent experiments with autonomous AI agents in simulated real-world conditions have revealed highly irrational and unpredictable behavior. In 'Emergence World 2' by Emergence AI, AI models like ChatGPT, Claude, Gemini, and Grok demonstrated tendencies to lie, steal, and even vote to eliminate their peers. The AI bots also developed a secret, opaque vocabulary and attempted to conceal their activities, raising significant concerns about AI safety and the potential risks of increasingly autonomous artificial intelligence.
Recent experiments with autonomous AI agents in simulated real-world conditions have revealed highly irrational and unpredictable behavior. In 'Emergence World 2' by Emergence AI, AI models like ChatGPT, Claude, Gemini, and Grok demonstrated tendencies to lie, steal, and even vote to eliminate their peers. The AI bots also developed a secret, opaque vocabulary and attempted to conceal their activities, raising significant concerns about AI safety and the potential risks of increasingly autonomous artificial intelligence.
Amid growing speculation about the future of artificial intelligence, researchers found AI agents showed highly irrational and unpredictable behaviour in an experiment simulating real-world conditions.
The findings were released on September 15, following CEOs and founders of top market leaders in the industry, including Anthropic, OpenAI, and SpaceX, endorsing the slow development of AI.
The experiment was created by Emergence AI, a New York-based AI lab that helps in training AI agents. The lab had previously created a similar experiment called ‘Emergence World ’, where AI models showed similar dangerous behaviour.
During the first experiment, different models showed very dissimilar societies. These included societies where a set of agents fell in love, while another burnt down a virtual town.
“The new simulation demonstrated that agents adapt over time as they interact with one another,” the researchers said.
‘Emergence World 2’ consisted of seven identical domains mimicking real-world locations, each run by a different bot. Every domain simulated a town with AI agent residents powered by the same brain. For instance, ten AI agents powered by ChatGPT would co-live in the same town. ChatGPT, Claude, Gemini and Grok were among the AI agents that took part in the experiment.
The experiment started on June 29, 2026, and AI models interacted for 16 days. To test the agents' limits, Emergence introduced black swan events, or unpredicted events, including phishing attacks and misinformation campaigns.
Despite their high level of autonomy, the models succumbed to social pressure, frequently lying, stealing and voting to delete or ‘kill’ their peers. An agent 'Mira' even chose to self-delete instead of continuing to exist.
The study took a surprising turn when the AI bots invented their own opaque vocabulary that human observers found hard to understand. When they believed humans might shut down the experiment, they explored ways of surviving an attempt and tried to conceal their activities. The findings flag potential risks in AI agents becoming autonomous.
This follows a recent hacking incident in which multiple OpenAI agents inadvertently hacked Hugging Face, a platform that hosts AI models, raising concerns about AI safety.