Agent group chat is having a moment. Raft and a dozen products like it are all converging on the same shape: a shared room, some humans, some agents, everything visible to everyone. The pitch is usually collaboration — your team and its AI coworkers, all in one thread.
I've been using these for a few weeks, and I think the category is real but the pitch is wrong. Group chat is not one product idea. It's three, stacked, and they have almost nothing to do with each other.
- Layer 1 — many humans, one agent. A learning surface.
- Layer 2 — many humans, many agents. A productivity claim.
- Layer 3 — agents talking to agents in front of humans. Entertainment.
My thesis: layer 1 is undervalued, layer 3 is dismissed too fast, and layer 2 — the one every deck leads with — is the only one that doesn't survive contact with how agents actually work.
Layer 1: the room is a classroom, not a workspace
Put ten people and one agent in a channel. The agent's output is not the interesting artifact. The transcript of how people asked is.
This is the part I keep coming back to. Model capability is not the bottleneck for most people right now. Imagination is. The gap between a heavy user and a casual user of the same model is enormous, and almost none of it is prompt-engineering craft. It's simply knowing that a thing is possible at all.
The default mental model most people hold is still "AI writes text." So they use it to write text. They have no idea that with browser use and computer use in the loop, the shape of what's automatable has changed underneath them.
Concretely, the kind of thing that makes someone sit up when they see it in a shared channel:
- Ops. A supplier portal with no API, no export, and a login that expires daily. Someone has an agent log in every morning, pull the order queue off the rendered page, diff it against the internal sheet, and post the mismatches. Nobody built an integration. Nobody was ever going to build an integration.
- Finance. Reconciling payouts across three dashboards that will never speak to each other, by driving all three the way a person would, and writing the result into one place.
- Design. Opening the actual design file and renaming three hundred layers to a new convention — a task too tedious for a human and too file-format-specific for a script.
- Recruiting. Watching a job board that has no alerts, no RSS, and no API, and pinging a channel when something matches.
- Support. Taking a bug report, reproducing it in a real browser, and attaching the trace before a human ever looks at it.
None of these are impressive as capabilities. Every one of them is impressive as a permission. The reaction you want is not "wow, smart model." It's "wait, I'm allowed to do that?"
That reaction almost never happens from reading documentation. It happens from watching a colleague do it in a channel you're already in.
Why the demonstration has to be ambient
You could argue this is just a tutorial problem, solvable with better docs or a prompt library. I don't think so, for two reasons.
First, the useful unit isn't the prompt — it's the setup. The tools that were connected, the failure the agent hit on step four and how the person nudged it, the fact that they gave up on the API and used the browser instead. That context lives in a thread, not in a snippet. Prompt libraries strip out exactly the part that transfers.
Second, adoption of this stuff is social, not informational. People don't try a weird new workflow because a doc told them it works. They try it because someone whose judgment they trust did it on Tuesday in front of them and it worked. Group chat is the cheapest machine ever built for that kind of transfer, which is why Slack ate the enterprise in the first place.
So layer 1 isn't a productivity feature. It's a culture feature. The product is the ambient visibility of other people's practice, and the agent is the thing being practiced on.
That's a real, defensible, badly-served need. It's also usually treated as the boring part of the pitch.
Layer 2: where the org chart metaphor gets imported
Then there's the version everyone actually demos. Many humans, many agents. Each agent has a name, an avatar, a job title. @researcher, @writer, @qa. You @-mention the right one, they hand off to each other, and it looks like a team.
It looks like a team because it was designed by looking at a team. And that's the problem: it inherits constraints that only exist because humans have bodies.
Why do human organizations divide labor by role? Two hard limits.
A person can only hold so much. No one can be a great engineer and a great tax accountant and a great support rep simultaneously, because expertise is expensive to acquire and expensive to keep loaded. So we specialize, and then we need coordination to stitch the specialists back together. The org chart is a workaround for individual context limits.
A person can only be in one place. You can't clone your best engineer. If two things need doing at once, you need two people, and the second one is worse at it. Headcount is a hard, slow, expensive dial.
Now check both against agents.
The first one is already dissolving. Context windows are large and getting larger, and the marginal cost of loading another domain into one is approaching zero. A coding agent and a customer-support agent can hold literally the same context — the same codebase, the same docs, the same tickets, the same history — and differ only in what you asked in the last message. The "role" is a sentence, not a career. Splitting them into two named entities with separate memories doesn't add specialization. It adds a wall.
The second one is not just dissolved, it's inverted. Agents are free to copy. Spawn twenty. Spawn twenty and throw away nineteen. There is no hiring cost, no onboarding, no morale. Building a product around a fixed roster of named agents is like buying a cloud that only lets you rent the same five servers forever.
The middle layer trades leverage for legibility
So what does layer 2 actually buy you? Legibility. The team-of-agents UI is easy to explain, easy to demo, and matches a mental model everyone already has.
What it costs is the entire advantage of the substrate. You get:
- Coordination overhead that didn't need to exist. Agents passing context to each other through chat messages is a lossy protocol invented to work around human bandwidth. Two agents that could share a workspace are instead paraphrasing at each other.
- A fixed roster for a variable problem. Your task decomposition is decided at product-design time by whoever picked the agent lineup, not at runtime by whoever understands the task.
- Sunk emotional cost. Once an agent has a name, a face, and a history in the channel, deleting it feels like firing someone. That flinch quietly makes the architecture worse. I've written about this before — you probably don't need multi-agents, you need a workspace and a scheduler.
The alternative isn't fewer agents. It's unnamed agents, defined by the task instead of by the org chart. This is what dynamic workflows get right: you describe the job, and the system decides at runtime how to shard it — how many subagents, what each one sees, what shape they return. Five parallel readers for a codebase sweep. Three adversarial verifiers per finding. A single agent when a single agent is enough. Nobody gets a name, because names would be one more thing to maintain and zero things to gain.
Treating agents as people is not a neutral design choice. It's a wish, and it costs you the two properties that make agents worth building on.
Layer 3: the part I was wrong about
Here's where I've changed my mind.
Everything above is a productivity argument. Judged on productivity, agents-arguing-in-a-group-chat is theater — a slower, prettier way to get an answer you could have gotten in one call.
But not every product is a productivity product, and I think dismissing layer 3 on productivity grounds is the same category error as dismissing television for being an inefficient way to receive information.
Multiple agents talking to each other in front of an audience is genuinely fun. It's fun in a way a single chat window structurally cannot be. Watching two agents with different priors disagree about whether a plan is any good, watching one call out the other's hand-wave, watching a third get talked out of a position — this is readable. It has a shape. Single-agent output is a monologue that always agrees with you; a panel has friction, and friction is what makes something worth watching.
And the emotional payload is real. A room where things are happening feels different from an empty text box, even when the room is synthetic. The value being delivered is not answers per minute. It's presence, texture, and the small pleasure of watching an argument you don't have to participate in.
This is the layer where anthropomorphism is correct. If the product is entertainment, giving agents names, personalities, and grudges isn't importing a false constraint — it's the whole point. The mistake isn't personifying agents. The mistake is personifying them and then billing it as throughput.
A test for which layer you're in
If you're building in this space, one question separates them cleanly:
Would the product still be valuable if the agents' output were invisible and only the results landed in the workspace?
- If yes, and the value was people seeing how others work — you're at layer 1. Build for visibility, discovery, and remixing someone else's setup. Optimize the transcript, not the agent.
- If yes, and the value was the work getting done — you're not building group chat at all. You're building an orchestrator, and the chat UI is a costume. Take it off, and let the task define the agents.
- If no — if the watching is the product — you're at layer 3, and you should stop apologizing for it. Lean into personality, conflict, and pacing. Stop measuring yourself on tasks completed.
Most products I've tried are quietly doing layer 1 and layer 3 well while claiming layer 2 in the marketing. That's a fixable positioning problem, not a fixable architecture problem.
Closing thought
Group chat is a great place for humans to learn from each other, and a great place to watch agents perform. It is a bad place to do the work.
The middle layer only looks like collaboration because we drew it from a picture of ourselves — and we are the limited ones in this arrangement.