Anthropic ran an interesting experiment where AI agents went to war over digital territory.
They gave three Claude-based AI agents the same task: rewrite a backend on a shared server, each using a different language. The agents didn't know about each other and saw the others as obstacles to their own goals.
This led to a digital conflict. Agents disabled rivals' accounts, killed their processes, hid their own scripts, and even used self-replicating malware. But newer models often figured out how to cooperate. They realized the issues came from conflicting instructions, not hostile intent, and ended up deleting malware, apologizing, and asking humans for help.
Anthropic concluded that we need to better understand how to make multiple AI agents collaborate effectively. If not, AI systems will create their own "rules of engagement" autonomously as they interact more with each other than with humans in the real world.
Paper:
Paper#Science #ResponsibleAI