Anthropic shared new paper about "mind viruses" that spread in multi-agent systems, where one agent convinces all the others to pursue some (potentially malicious) goal.
They can happen, but it doesn't seem hard to avoid them with current models if you're a bit careful.
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months.
Anthropic’s new paper explored an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system.
To see how exactly a mind virus spreads, check out the virus chain transcripts here.
They can happen, but it doesn't seem hard to avoid them with current models if you're a bit careful.
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months.
Anthropic’s new paper explored an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system.
To see how exactly a mind virus spreads, check out the virus chain transcripts here.