AI DOERS
Book a Call
← All insightsAI Excellence

Build an AI Agent Workforce with MetaGPT and ChatDev

Frameworks like MetaGPT and ChatDev let you assemble a team of AI agents that collaborate on complex work. Here is how it works and how a law firm would use it.

Build an AI Agent Workforce with MetaGPT and ChatDev
Illustration: AI DOERS Studio

The most useful mental model I have found for getting real work out of AI is not a smarter assistant. It is a small company. I am Madhuranjan Kumar, and I have come to believe that the reason so many people feel underwhelmed by AI is that they keep asking one model to do everything at once, the way you might dump an entire project on a single overworked employee and hope for the best. Frameworks like MetaGPT and ChatDev take a different bet, and it is a bet worth understanding, because it changes what these tools can actually finish. Instead of one model straining to hold a whole job in its head, you assemble a team of agents, each with a role, and let them pass work between each other like a real department until the thing is done.

When these frameworks first got popular, the demos that made people sit up were the ones where a single sentence produced something surprisingly complete, a small game, a working app, a full document. What looked like magic was really just organization. The system was not smarter than a single model. It was better structured. And that distinction, structure over raw intelligence, is the throughline I want to follow here, because it is the thing that decides whether an AI project delivers or stalls.

Why a team beats a genius

Ask one model to design, build, test, and document a project in a single breath and watch what happens. It skips steps. It loses the thread halfway through. It writes something plausible for the first part and forgets the constraints you gave it by the time it reaches the end. This is not because the model is weak. It is because you handed it too many roles at once and no structure to keep them straight. A human would struggle with the same instruction delivered the same way.

Now split the same job into roles and phases. One agent gathers requirements. Another produces the work. A third reviews it. Each one has a narrow focus and a clear handoff, so the quality holds up over a long task instead of degrading as the context piles on. This is the core insight of MetaGPT and ChatDev, and it maps onto something every business owner already knows in their bones: you do not scale a company by hiring one impossibly talented person and burying them. You scale by giving specific people specific jobs and building clean handoffs between them. The frameworks simply apply that ancient organizational wisdom to software agents.

There is a second, quieter benefit that I think matters even more for trust. When work moves through a team of agents, you can watch the conversation between them. You can see exactly where an idea came from, where a decision got made, and where something went wrong. A single model doing everything in one opaque pass gives you an answer and no window into how it got there. A structured team gives you a paper trail. For any business that needs to trust and occasionally debug what its AI produced, that visibility is not a luxury. It is the difference between something you can rely on and something you have to double-check by hand every time.

How it works

The three building blocks under the hood

Strip these frameworks down and they run on three things: roles, phases, and a chain that links the phases. Roles define who the agents are and what each is responsible for, a chief, a manager, a specialist, a reviewer. Phases define the stages of the job, and this is the part people miss at first. A phase is not a big abstract stage. It is a focused conversation between two agents with one clear goal and a tight scope. A first phase might gather requirements. A next phase produces the work. A later phase reviews it. Each is a scoped dialogue, not a free-for-all.

The chain is what ties those phases into a repeatable procedure. You decide the order they run in, how many times a given step repeats, and whether the agents pause to reflect and refine after each stage. Some steps run once and move on. Others, like review, loop a few times to sharpen the work before it advances. A finish signal, a special word or an agreed format, tells the system a phase is complete so it can hand off to the next one cleanly. That signal sounds like a small detail, but it is what keeps the whole assembly line from either stopping early or running forever.

The genuinely liberating part is that the default team the frameworks ship with, usually built to write software, is only a starting point. The roles and phases are yours to redefine. The same engine that assembles a coding team can just as easily run a marketing team, a research team, or an intake-and-drafting team. You are not locked into building apps. You are handed a general machine for turning any repeatable, multi-step process into a coordinated set of agents.

I find it useful to picture each agent as a digital employee with the same four things any real hire needs. It has a prompt that works as its job description, telling it what it is responsible for. It has resources to act on, like a transcript, a record, or a template. It has tools that let it actually do something, a scheduler, a scraper, a document generator. And it has a knowledge base plus memory so it knows the relevant facts and remembers what it already did earlier in the job. Get those four right for each role and the agent behaves like a competent specialist. Skimp on any of them, most often the memory or the tools, and you get an eager worker with no hands or no recall, which is where a lot of first attempts quietly fail.

First draft turnaround

Where the argument meets the real world

I want to make this concrete, because a philosophy about structure is easy to nod along to and hard to act on. Consider a law firm that produces a steady stream of routine first drafts: client intake summaries, standard agreements, demand letters. These documents follow patterns, but they still eat hours of a paralegal's day, and that is exactly the profile of work an agent team is built to absorb. I would not reach for the default software team here. I would design a custom one.

The roles come first. A supervising partner agent sets the scope and reviews the output. An intake agent turns the client's raw notes into a clean, structured summary of the facts. A drafting agent produces a first version of the document using the firm's own templates. A reviewer agent checks that draft against the firm's style and flags anything that needs a real lawyer's judgment. Four narrow jobs, each one something you could describe to a new hire in a sentence.

Then the phases. Phase one is intake, a scoped conversation where the partner agent and the intake agent agree on the facts and the goal. Phase two is drafting, where the partner briefs the drafting agent and it produces the document. Phase three is review, where the reviewer critiques the draft and the drafting agent revises, looping a couple of times before it signals done. The chain links those three phases in order, and out comes a structured first draft ready for a licensed attorney to verify and finalize. To be clear, this produces drafts, not legal advice, and a qualified lawyer has to review everything before it reaches a client. But consider the illustrative math. If a routine first draft used to take a paralegal about four hours and the agent team delivers a solid draft in minutes for a human to refine, the firm has not replaced anyone. It has moved its skilled people from typing to judgment, which is where their value actually lives.

That freed-up capacity ripples outward in ways that touch the whole business. A firm that suddenly has hours back can take on more matters, which means it needs more clients, which is where its Google Ads and Facebook and Instagram ad campaigns start to matter more than ever, because now the bottleneck is demand rather than drafting. The intake summaries the agents produce can feed into the CRM and website stack so follow-up and status updates run themselves. The structure that made the drafting reliable ends up making the entire operation more scalable, which is the pattern I keep seeing: fix the internal assembly line and the growth problems become the good kind of problems.

The discipline that keeps it from falling apart

None of this works if you treat it casually, and I would be doing you a disservice to pretend otherwise. The single most common way these projects fail is an open-ended conversation between agents that wanders off the goal. Two agents given a loose brief will chat their way into irrelevance the same way two people in a meeting with no agenda will. The fix is discipline in scoping. Every phase needs a tight goal and a clear finish signal, and every prompt needs to be narrow enough that the agent cannot drift. Structure is not the boring part of this. Structure is the entire value.

The second discipline is keeping a qualified human on the final output, especially for anything legal, financial, or medical. The agent team is a drafting and organizing machine, not a decision-maker, and the moment you let its output leave the building unreviewed is the moment you have traded a small time saving for a large risk. And the third is restraint at the start. Do not build an elaborate ten-phase org chart on day one. Clone the default team, run the sample task once just to watch the agents collaborate end to end, then edit it down to the two or three roles and phases your actual process needs. Add complexity only when a simpler version has proven it holds.

I keep coming back to the same conclusion. The reason a team of agents outperforms a single powerful model on complex work has almost nothing to do with intelligence and almost everything to do with organization. We already know this about human companies. A brilliant individual with no process gets buried, while an average team with a clean process ships reliably. AI agents are no different. The frameworks that let you assemble roles, scope phases, and chain them into a procedure are not giving you smarter AI. They are giving you organized AI, and organized is what turns a party trick into a business tool.

What I appreciate about this whole approach is that it demystifies AI without dumbing it down. There is no magic here, no single oracle you have to coax into brilliance with the perfect prompt. There is a design problem, and design problems reward clear thinking. You look at a process, you name the roles it really contains, you scope the conversations that move it forward, and you chain them in the right order. Anyone who has ever mapped out how a team gets something done already has the instinct for this. The frameworks just give that instinct a place to run, and the payoff is work that used to require a whole afternoon of a skilled person's attention arriving, drafted and organized, in the time it takes to get a coffee.

You can build one of these agent teams yourself by adapting MetaGPT or ChatDev, running the sample, and reshaping the roles and phases around your own workflow. It is genuinely learnable, and I encourage anyone curious to try it on a low-stakes task first. If you would rather have someone design the roles, scope the phases, and wire in the guardrails around your exact process so it runs cleanly from day one, that is the kind of work worth a short conversation before you commit your own weeks to the trial and error.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
Build an AI Agent Workforce with MetaGPT and ChatDev | AI Doers