AI DOERS
Book a Call
← All insightsAI Excellence

How Claude Code, HeyGen, and ElevenLabs Turn a Script Into Video Overnight

Claude Code can orchestrate HeyGen, ElevenLabs, and Remotion so you hand it a raw script and wake up to a finished, edited avatar video. Here is how the pipeline works and how I would run it for a business.

How Claude Code, HeyGen, and ElevenLabs Turn a Script Into Video Overnight
Illustration: AI DOERS Studio

Madhuranjan Kumar who opened this breakdown was not real. It was an AI clone built in about ten minutes, fronting an entire introduction for a creator who had never appeared on camera before. That is a striking party trick on its own, but the part worth studying is what sits behind it: a pipeline where you hand a raw script to Claude Code, go to sleep, and wake up to a finished, edited avatar video.

I want to take this pipeline apart piece by piece, because the interesting thing is not any single tool. It is how they are wired together. Three separate products that each used to require their own operator now run as one unattended flow, and understanding how the pieces connect tells you exactly what this can and cannot do for a business that needs to produce a lot of video without a lot of people.

The orchestration layer is the actual innovation

HeyGen, ElevenLabs, and Remotion have all existed on their own. What is new is Claude Code sitting on top of them as an orchestration layer. Instead of a human clicking between three apps, generating an avatar clip here, a voice track there, then editing it all together somewhere else, Claude Code coordinates the whole chain. It takes the script, drives each tool in turn, and hands off the output of one as the input to the next. The roles that used to belong to three or four different people, a presenter, a voice artist, a video generator, an editor, collapse into a single automated pipeline that a machine runs end to end.

This is the mental shift that matters. The tools are the muscles, but the orchestration is the nervous system, and it is the nervous system that turns a pile of capable apps into something that runs while you are not watching. Once you see it that way, the pipeline stops looking like a clever demo and starts looking like a production line.

How it works (short)

Each tool does one job, and only one

The division of labor is deliberate, and it is worth being precise about because doing it wrong is the most common way people get bad results. HeyGen builds the avatar, the visual you, generated from a short webcam clip. ElevenLabs handles the voice, a clone of how you actually sound. Remotion does the editing and motion graphics, stitching the clips and adding the on-screen text and animation. Claude Code connects all three so you never manually move a file between them.

The key insight is not to let any one tool stretch beyond its job. HeyGen has a built-in voice clone, but it is weak, so the voice is built separately in ElevenLabs and imported. A professional voice clone trained on a couple of hours of your audio sounds dramatically more like you than a clone trained on fifteen seconds. Using each tool only for the thing it is best at, and stacking a stronger specialist tool where the default is weak, is what lifts the output from obviously-synthetic to genuinely convincing. A business that shortcuts this by using HeyGen's built-in voice will get a result that feels off, and the whole illusion depends on the voice being right.

Marketing videos shipped per month (illustrative)

Avatar 5 is the piece that crossed the uncanny valley

The reason this is worth doing now rather than a year ago is the quality of the avatar model. HeyGen's Avatar 5 is trained on an enormous set of facial-expression data and can clone you from a fifteen-second webcam clip, learning how you actually gesture rather than just moving your lips in sync with the audio. That distinction is everything. Older avatars looked dead because the face said the words but the body language was wrong, and viewers feel that wrongness instantly even if they cannot name it. An avatar that gestures the way you naturally do reads as a person, not a puppet.

There is a practical wrinkle here that reveals how new this all is. Avatar 5 is not yet available through an API, so the pipeline generates the clips in the older Avatar 4 first, then uses a browser automation script, driven by Playwright, that opens HeyGen and clicks each clip up to the better Avatar 5 model. It is a workaround, and it tells you that the smooth overnight pipeline is stitched together at the edges with some genuinely clever plumbing. For a business, the lesson is that the frontier tools often need a bit of automation glue to fit into a clean workflow, and that glue is part of the real work.

Chunking the script is the constraint that shapes everything

The single most important production rule in this pipeline is about length, and ignoring it is how people get garbled output. ElevenLabs audio degrades past roughly a minute, and HeyGen's Avatar 5 caps at around three minutes. The fix is to chunk the script into pieces of about forty-five to sixty seconds, and always cut at the end of a sentence, never mid-thought. The pipeline generates each chunk cleanly, and Remotion later stitches them so the seams are invisible.

This constraint is not a limitation to work around so much as a discipline that improves the output. Writing in tight, self-contained chunks forces clearer scripting, and clearer scripting makes better video regardless of who or what is presenting it. The businesses that get the best results from this pipeline are the ones that respect the chunk length and write to it, rather than dumping a long unbroken monologue in and hoping the tools cope.

Remotion turns clips into a finished edit automatically

The last stage is where a lot of the magic hides. Remotion transcribes the generated video, reads the timestamps, and stitches the clips so the joins cannot be seen, then drops motion graphics in at the exact moment each line is spoken. That timestamp-driven placement is what makes the output look edited rather than merely assembled. Text appears when the relevant word is said, animations land on the beat, and the whole thing feels produced. Doing that by hand is slow, finicky editing work, and it is precisely the part that used to eat the most human hours.

Because Remotion works from the transcript, the same transcript is a free byproduct you can reuse. Every video the pipeline makes hands you a clean text version of its own content, which becomes blog copy, captions, and search-friendly page text. That means the content you produce for video also feeds SEO and organic search with no extra effort, and the clips themselves become fresh creative for Facebook and Instagram ad campaigns. One script produces a video, a transcript, and ad assets all at once.

The overnight job is the point of the whole thing

Here is the payoff that reframes the tool. Madhuranjan Kumar told Claude Code to process a batch of lessons, went to bed, and woke up to finished videos. A production pipeline that used to take around five hours of active human work became an unattended overnight job. That is not a small efficiency; it is a change in what a small team can produce. Video stops being a scarce, expensive output rationed a few pieces a month and becomes something a business can generate in batches while nobody is at the desk.

But the same efficiency exposes the new constraint, and this is the part businesses must not miss. When production and post-production stop being the bottleneck, the bottleneck moves to ideas: the thinking, the scripting, the strategy. A great avatar reading a bad script is still bad content. Removing the production ceiling means you can make far more video, which is only valuable if this breakdown is worth making. The human stays firmly in the loop where judgment lives, deciding what to say and why, and hands only the mechanical production to the pipeline.

The cost math, and where it actually lands

The economics are worth stating plainly because they are what make this real for a small business. HeyGen runs around thirty dollars a month and ElevenLabs around twenty-two, with API-generated clips costing near four dollars a minute. A ten-minute video therefore costs roughly fifty dollars in tooling, against something like three hundred dollars for a freelance editor to produce a comparable piece. On a per-video basis the saving is real, but the more important number is throughput. When each video costs an afternoon of a freelancer's time versus an unattended overnight run, the constraint on how much you publish disappears.

That throughput is a double-edged thing, and honesty about it matters. Removing the bottleneck means you can make more content, but it does nothing to make the content good. A business that pumps out ten mediocre videos a week because it can will lose to one that publishes two sharp ones. The pipeline is a force multiplier on whatever you feed it, which means the discipline shifts entirely onto the script and the strategy behind it.

A worked example: a course business ships a month of lessons in a weekend

Consider a small education business that sells online courses and needs a steady stream of lesson videos and marketing clips. Before this pipeline, each lesson meant booking a presenter or filming the owner, then paying an editor, roughly three hundred dollars and several days of turnaround per video. Producing a full course of twenty lessons was a months-long, expensive project, which capped how fast the business could launch new material and how much marketing video it could afford to make.

They set up the pipeline once: an avatar cloned in HeyGen, a voice trained in ElevenLabs on a couple of hours of the owner's audio, Claude Code wired to orchestrate the run, and Remotion handling the edit and graphics. Then the owner wrote the scripts, chunked into sixty-second sentences, and told Claude Code to process the batch over a weekend. The scripts were the real work; the production ran itself. The illustrative tooling cost for the whole batch was on the order of a few hundred dollars in subscriptions and API usage, against what would have been thousands in editing fees, and the turnaround dropped from months to a couple of unattended overnight runs.

The result is not just cheaper video; it is a different operating tempo. The business can now ship a month of lessons in a weekend, refresh its marketing clips weekly, and test new course ideas with real video before committing, all without hiring a production team. The leads those marketing clips generate flow into the business's CRM and website stack, where automated follow-up turns viewers into enrollments. The owner spends their time on what to teach and how to sell it, which is exactly where their judgment is worth the most, and hands the production entirely to the pipeline.

The mistakes that break the illusion

The first mistake is using HeyGen's built-in voice instead of stacking ElevenLabs. The voice is the piece viewers judge hardest, and the weak default is the fastest way to make the whole thing feel fake. The second mistake is ignoring the chunk limits and feeding long unbroken scripts in, which produces degraded audio and capped, cut-off clips. The third and most important mistake is treating the removed bottleneck as permission to lower the bar on ideas. The pipeline makes production free; it does not make content good. Put your effort into the script, because that is now the only thing standing between you and a wall of polished, worthless video.

The bottom line

This pipeline matters because it collapses a five-hour, multi-role video production into an unattended overnight job, using Claude Code to orchestrate HeyGen, ElevenLabs, and Remotion, each doing the one thing it is best at. The tools crossed the quality line, the cost dropped to a fraction of a freelance editor, and the constraint moved from production to ideas. Respect the chunk limits, stack the strong voice tool, and put your real effort into the script.

You can wire this together yourself, and the components are all publicly available. But orchestrating them so they run reliably overnight, with the browser automation and the stitching dialed in, is genuinely fiddly the first time. If you would rather have the pipeline built and handed over working, so your team simply writes scripts and wakes up to finished video that feeds your ads and your funnel, that is the kind of setup worth a focused conversation. Either way, the lesson holds: production is no longer the hard part, so make sure the ideas are.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
How Claude Code, HeyGen, and ElevenLabs Turn a Script Into Video Overnight | AI Doers