← All work
Multi-agent · strategyPersonal project

Multi-Agent Consulting Engine

This one is an experiment, and I am not going to dress it up as a product. Three AI consulting "firms" compete on the same brief, tear into each other's work, and produce one synthesized answer.

I have run it twice, both times on a real question. On the first run, all three firms killed the founding assumption in my own brief. It stays on this site labeled as what it is, because calling your own builds accurately is most of the job.

3
competing firm personas per run
up to 26
agents in a full run
2
rounds, with adversarial cross-review

The problem

One AI answer is a single point of view, confidently delivered.

Hard strategy questions need more than that. They need competing approaches, real pressure-testing, and someone to call out the weak assumptions before they ship.

What I built

An orchestration where three AI "firms," modeled on the styles of McKinsey, BCG, and Bain, each run as their own swarm of about five specialist agents.

They compete on the exact same brief, then the work goes through a structured fight:

  1. Round 1, compete. Each firm researches independently and writes its own proposal in its own voice.
  2. Cross-review. A panel agent ranks all three and points out what each firm missed that its rivals caught.
  3. Round 2, rework. Each firm gets the critique plus what competitors found, and revises.
  4. Synthesis. A partner agent compares everything across both rounds and writes the definitive report, taking the best of all three.
How it runs: one brief goes to three competing AI firms of five agents each, they rework each other's work, and a partner agent writes the final report.
One brief, three firms, one report. Round two only spins up gap researchers where the review panel actually found gaps.
up to 26

agents on a full run. It is a ceiling, not a fixed count.

Where the "up to" comes from

Fifteen agents in round one, one review panel, up to nine in round two, and the partner at the end.

It is a ceiling, not a fixed count, because round two only spins up gap researchers where the panel actually found gaps.

The part I am proudest of

On the first real run, all three firms independently killed the founding assumption baked into my original brief.

The system did not just answer the question. It told me the question was partly wrong, and why, with evidence.

That is the "should we before can we" principle working as designed: the value is in honest pushback, not agreeable output.

Where I actually ran it

Twice, both on questions I genuinely wanted answered.

  • June 2026. A personal thesis about where economic growth should land as automation spreads. That is the run where the firms killed my framing.
  • July 2026. Pressure-testing whether a digital twin was a real business or just a thing I liked building.

Being straight about it: this is an experiment, not a tool I reach for every week.

The reports were good, and the pushback was worth having, but it did not end up giving me a lot more than working the same question carefully with one model. I keep it here because the orchestration is the interesting part.

Twice

is the entire run history, and I am not going to imply otherwise. It answered a real question both times and then I stopped reaching for it. Two runs is what it is.

Deciding a build has stopped being useful is its own skill, and it is the one most teams are worst at. The expensive version of this mistake is the internal tool everyone quietly stopped opening, which nobody says out loud, so it keeps getting maintained and keeps getting demoed.

Writing "experiment" on it costs me nothing here. Inside a company it is the difference between a roadmap that reflects reality and one that does not.

Why it matters

The pattern is the point: structured competition, adversarial review, and synthesis, with confidence levels and sources attached. It is a pattern you could apply to any high-stakes decision.

And it is multi-agent orchestration doing real intellectual work on a real question, not a demo built to look impressive.