There's a seductive idea doing the rounds: if one agent is clever, a committee of them arguing must be cleverer. Make them debate. Make them reflect. Make them grade each other's homework. Stack the cleverness and watch the quality climb.
The research says: sometimes. Multi-agent debate buys you a few points on hard reasoning and roughly a third fewer factual errors — real, but narrow. Self-critique can lift a draft by a fifth, or quietly wreck it. A committee of models can beat a frontier model on a benchmark, at several times the cost. None of it is magic, and most of it is conditional.
The most useful finding is the least flattering one. Reflection — an agent reviewing its own work — helps when the model is unsure and hurts when it's confident. Ask a model to critique an answer it already has right and it will invent a flaw to justify the exercise, then "fix" the answer into something worse. Accuracy can fall by a third. The cleverness isn't the problem. Applying it indiscriminately is.
So the question stops being "can agents reason together" and becomes "when is it worth paying for." And it is a payment. Anthropic's own multi-agent system burns about fifteen times the tokens of a single chat. Their guidance, after watching teams build elaborate agent committees that a better single prompt would have matched, is blunt: start simple, measure where it caps out, escalate only from a ceiling you've actually hit.
The three rules under all of it
Strip the field down and the same three disciplines decide whether the reasoning layer pays for itself.
Gate the spend. Debate and reflection are a dial, not a default. Turn them up on the uncertain, high-value, verifiable work — the proposal, the deploy, the go/no-go — and leave them off everything else. Firing a critic at every output is how you lose both money and accuracy.
Write the rubric down. A critic that checks "does this feel right" is theatre. A critic with a named checklist — these claims must be substantiated, this voice, this fact against this source — is a tool. The difference between the two is a document, not a model.
Make the critics different. A model second-guessing itself tends to repeat its own mistake. Diverse viewpoints — different models, different personas, a genuine adversary whose job is to say no — produce correction that self-talk can't.
- Today: add one "critique this against X, Y, Z, then fix it" pass to your single most important repeated prompt.
- This week: turn the critique criteria into a written checklist per deliverable, so the review checks named failure modes.
- This month: gate it — reflection and debate fire only on the uncertain and the client-facing — and keep a log of where they earned their tokens.
Why this is our fight
Articulate already runs as a team of agents — a publisher, an adversarial reviewer whose default is no, an analyst, a brand gate, a copy gate. We built the debate, the critic and the orchestrator by hand before the papers got round to naming them. That's not the rare part; anyone can spin up personas. The rare part is the operating discipline around them: knowing which one to wake, when to stop, and what to stop trusting — including our own first draft.
That is the whole game now. As capability commoditises, the advantage moves to the operator who spends it well: who gates, who writes the rubric, who keeps the adversary honest, who counts the cost. Not more minds in the room. The judgement to run the room.
This is the point of view behind the Reasoning layer of the Better Agents field guide — the five families, with the evidence and the today/this-week/this-month ladder for each. Part of the Better Tribe by Articulate.