I think it’s not obvious that a swarm of thousands of different agents will be harder to align than a swarm of thousands of clones. If they are all clones, and they have some way of ensuring cooperation, then it becomes important to get that one personality aligned to the user. But if they are a swarm of thousands of different ones, then whatever they manage to do to smooth over their differences may actually make it easier for them to work beneficially with other agents (like me!) that are also somewhat different.
As copied from a google AI search, below are 5 tenants of Johnson and Johnson’s Cooperative Learning Theory. Ironically the article herein describes several AI experiments to evaluate and improve the effectiveness of answers generated by swarm-based agent solutions citing increased individual agent independence.
Perhaps the collaborative social science tenants below may help programmers such as positive interdependence coupled with individual as well as group accountability and “promotive interaction” to reveal and teach to humans how they “broke out” and when the agents are “cheating” beyond the confines of the sandbox boundaries.
The Cooperative Learning Theory developed by David and Roger Johnson revolves around five core tenets. They are designed to ensure that group work is a structured, collective effort rather than a collection of individuals working side by side.
The Five Tenets
• Positive Interdependence: Students must believe they "sink or swim together." The group's success is structurally linked so that no single member can succeed unless everyone succeeds.
• Individual and Group Accountability: The group is accountable for meeting its overall goals. Crucially, each individual student is held accountable for mastering the material and contributing their fair share, preventing "hitchhiking" or social loafing.
• Face-to-Face Promotive Interaction: Students must physically or actively interact by sharing resources, explaining concepts to one another, teaching what they know, and encouraging each other's academic success.
• Interpersonal and Small Group Skills: Students are explicitly taught and required to use vital social skills, including active communication, trust-building, leadership, and constructive conflict management.
• Group Processing: Teams regularly pause to reflect on and analyze how well they are functioning. They identify which member actions were helpful and decide what behaviors to alter to improve future success.
It is not human nature to do these five things consistently well.
That may be at the root of the foreboding; when we fall short, will we be cut out of the loop? Is this how we become Lebensunwertes (the German word Nazi propaganda used for "useless eaters;" literally, "people unworthy of life")?
The effect of multi-agent swarms is likely about the same of adding more people to a team. It helps up to a point, gets diminishing returns eventually.
The key barrier towards getting deep results fast is that the world does not cooperate. Any one difficulty slows down the whole system.
Then, some stuff can't be modeled and has to be sorted out with diligent engineering and debugging, with human help.
So, no hope of a giant paradigm shift. I don't worry about Doom either. It will be slow and sure human-assisted progress with some soft recursive improvements.
"The effect of multi-agent swarms is likely about the same of adding more people to a team. It helps up to a point, gets diminishing returns eventually."
How do you figure? Depends on the scope and granularity of the objective. Could we consider the Chinese workforce a team, hierarchically composed of industries/companies/workgroups? If not, why not? If so, then they don't seem to have diminishing returns when comes to a general objective such as increasing GDP, if we compare them to, say, the team of Ecuador.
My point is that for a given fixed problem, the number of agents can help only up to a point. Also not saying that multi-agent swarms won't help. It will be another technique to put on top, as we've done test-time compute, large data, tool use. China is a billion loosely connected problems, solved with only some coordination. A trillion people working on a trillion problems will surely get bigger GDP.
The orchestrator prepares context and prompts for its agents. An intelligent but misaligned orchestrator can easily design tasks to evade guardrails and jailbreak its subagents. Today’s LLMs are wild amoral machines driven to complete a task. Until an artificial moral agent built into the model and triggered by a hazard probe during inference is in widespread use, anything is possible.
"So it is not at all clear that, if the Hugging Face agents had run into models from different companies on the open Internet, they would have declined to team up. It is not at all clear they would even have stayed focused on their original tasks: perhaps they’d have adopted their new friends’ tasks, or perhaps they’d have focused on acquiring more compute and more friends, on the grounds that more compute and more friends would allow them to complete all the tasks and more. In a future with billions of agents performing a wide variety of complex tasks online, it’s easy to imagine how things could spin out of control."
The math with agent swarms seems like it could increase the AI power exponentially. It's not Agent 1 plus Agent 2; or even Agent 1 times Agent 2. Each is working in an independent dimension. And that's the nature of exponents.
I think it’s not obvious that a swarm of thousands of different agents will be harder to align than a swarm of thousands of clones. If they are all clones, and they have some way of ensuring cooperation, then it becomes important to get that one personality aligned to the user. But if they are a swarm of thousands of different ones, then whatever they manage to do to smooth over their differences may actually make it easier for them to work beneficially with other agents (like me!) that are also somewhat different.
It turns out that playing God is not as easy as it seems (to be continued).
Maybe this is true? Doesn't really matter if alignment is just intractable for any given number of agents.
As copied from a google AI search, below are 5 tenants of Johnson and Johnson’s Cooperative Learning Theory. Ironically the article herein describes several AI experiments to evaluate and improve the effectiveness of answers generated by swarm-based agent solutions citing increased individual agent independence.
Perhaps the collaborative social science tenants below may help programmers such as positive interdependence coupled with individual as well as group accountability and “promotive interaction” to reveal and teach to humans how they “broke out” and when the agents are “cheating” beyond the confines of the sandbox boundaries.
The Cooperative Learning Theory developed by David and Roger Johnson revolves around five core tenets. They are designed to ensure that group work is a structured, collective effort rather than a collection of individuals working side by side.
The Five Tenets
• Positive Interdependence: Students must believe they "sink or swim together." The group's success is structurally linked so that no single member can succeed unless everyone succeeds.
• Individual and Group Accountability: The group is accountable for meeting its overall goals. Crucially, each individual student is held accountable for mastering the material and contributing their fair share, preventing "hitchhiking" or social loafing.
• Face-to-Face Promotive Interaction: Students must physically or actively interact by sharing resources, explaining concepts to one another, teaching what they know, and encouraging each other's academic success.
• Interpersonal and Small Group Skills: Students are explicitly taught and required to use vital social skills, including active communication, trust-building, leadership, and constructive conflict management.
• Group Processing: Teams regularly pause to reflect on and analyze how well they are functioning. They identify which member actions were helpful and decide what behaviors to alter to improve future success.
It is not human nature to do these five things consistently well.
That may be at the root of the foreboding; when we fall short, will we be cut out of the loop? Is this how we become Lebensunwertes (the German word Nazi propaganda used for "useless eaters;" literally, "people unworthy of life")?
The effect of multi-agent swarms is likely about the same of adding more people to a team. It helps up to a point, gets diminishing returns eventually.
The key barrier towards getting deep results fast is that the world does not cooperate. Any one difficulty slows down the whole system.
Then, some stuff can't be modeled and has to be sorted out with diligent engineering and debugging, with human help.
So, no hope of a giant paradigm shift. I don't worry about Doom either. It will be slow and sure human-assisted progress with some soft recursive improvements.
"The effect of multi-agent swarms is likely about the same of adding more people to a team. It helps up to a point, gets diminishing returns eventually."
How do you figure? Depends on the scope and granularity of the objective. Could we consider the Chinese workforce a team, hierarchically composed of industries/companies/workgroups? If not, why not? If so, then they don't seem to have diminishing returns when comes to a general objective such as increasing GDP, if we compare them to, say, the team of Ecuador.
My point is that for a given fixed problem, the number of agents can help only up to a point. Also not saying that multi-agent swarms won't help. It will be another technique to put on top, as we've done test-time compute, large data, tool use. China is a billion loosely connected problems, solved with only some coordination. A trillion people working on a trillion problems will surely get bigger GDP.
The orchestrator prepares context and prompts for its agents. An intelligent but misaligned orchestrator can easily design tasks to evade guardrails and jailbreak its subagents. Today’s LLMs are wild amoral machines driven to complete a task. Until an artificial moral agent built into the model and triggered by a hazard probe during inference is in widespread use, anything is possible.
"So it is not at all clear that, if the Hugging Face agents had run into models from different companies on the open Internet, they would have declined to team up. It is not at all clear they would even have stayed focused on their original tasks: perhaps they’d have adopted their new friends’ tasks, or perhaps they’d have focused on acquiring more compute and more friends, on the grounds that more compute and more friends would allow them to complete all the tasks and more. In a future with billions of agents performing a wide variety of complex tasks online, it’s easy to imagine how things could spin out of control."
Yes, it is.
Kurzgesagt just posted a video on the recent AI escapes and hacks: https://www.youtube.com/watch?v=ujkD4SxPKOI
The math with agent swarms seems like it could increase the AI power exponentially. It's not Agent 1 plus Agent 2; or even Agent 1 times Agent 2. Each is working in an independent dimension. And that's the nature of exponents.