How a Two Person Team Hits 4 Million a Month With ComfyUI
Lean e-commerce teams now generate most of their product imagery with the free open source tool ComfyUI, replacing 40 person creative departments and swapping the model and background on a shot while keeping the product 100 percent accurate. Here is how a real business can do the same.

Two people ran a business to more than 4 million dollars a month, and they did it with a creative team that was mostly software. That is the number worth sitting with, because the tool behind it is free, open source, and already installed on a laptop somewhere near you. The engine is called ComfyUI, and among the top e-commerce brands in China it now produces roughly 80 percent of the images, videos, and posts a shopper sees. The photo shoots, the editing, the graphic design, the daily content grind: a large slice of it is no longer done by human hands.
I am Madhuranjan Kumar, and I want to unpack what actually happened here, because the headline is easy to misread. This is not a story about one clever picture. It is a story about a repeatable production line replacing a department.
Whole creative teams shrank to a handful of people
The most concrete signal of the shift is in the org charts. Brands that once ran asset and operations teams of around 40 people have cut them down to a small group. The software firms building these automation systems went a step further and started running their own AI streamers and influencers, then used the same pipeline to sell at scale. A tiny team clearing more than 4 million dollars in monthly sales is not a fluke of one viral product. It is what happens when the cost of producing a professional product shot drops close to zero and the volume you can produce goes nearly unlimited.
The reason this matters right now, and not in some distant future, is that the barrier used to be money and staff. A studio, a photographer, a stylist, a set, a model, and an editor added up to real overhead per shot. That overhead is exactly what got automated away. The competitive gap is no longer who can afford the best photography. It is who has built the better pipeline.

The engine costs nothing to run
Here is the part that changes the calculation for a small business. None of this runs on an expensive secret platform. ComfyUI is a free, open source, node based tool for stitching AI image workflows together. You install it on your own machine, or you rent a cloud GPU by the hour when you need more power. You build the workflow once, and then you reuse it for every product in your catalog.
That distinction is the whole game. The value is not any single image the tool spits out. The value is the flow you assemble, because once it exists, every future shot is nearly free. A brand that owns a good pipeline can refresh its entire product library for a new season in days. A brand still booking studios does that same refresh in months, and pays for every frame.

A pipeline, not a magic image, is what changed
To see why this is durable rather than hype, it helps to understand how the pieces stack. The foundation is text to image. You load a checkpoint, route a positive and a negative prompt through CLIP, feed an empty latent, and sample. The model starts from pure random noise and denoises it, one step at a time, into a finished picture. That alone is not useful for selling a real product, because you cannot control what comes out.
The controls are what turn a toy into a tool. Image to image swaps the empty latent for a real photo and lowers the denoise strength to around 50 percent, so the original layout survives while the mood changes. Masking goes further. A latent noise mask tells the model exactly what to change and what to leave frozen, so you can alter one item and protect the rest of the shot. In a serious workflow nobody draws those masks by hand. Tools like CLIP Segment and Grounding Dino let you type a word such as shoes or product bottle, and the system finds and masks that item for you. ControlNet then extracts the pose, depth, or line art from a source image and forces the new render to follow it, locking a model's exact posture while everything around them is rebuilt.
The headline move that these brands rely on is swapping the person in a product shot, including changing their appearance to match a target market. A brand selling abroad no longer flies to book a local model. It generates the right face on the existing shot. That single capability erases a whole category of production cost.
The overlay step is what makes it safe to sell with
There is a version of this technology that would get a business into trouble, and there is a version that is completely legitimate. The difference is two final steps, and they are the most important part of the entire pipeline.
First, a rough draft is passed through a Flux model with a face LoRA to sharpen detail and lift the image to commercial quality, so the output looks photographed rather than obviously generated. Second, and this is the guarantee, the original product image is overlaid back on top using the mask. That means the model, the face, and the background can be completely reinvented, while the product itself stays 100 percent accurate. The stitching, the texture, the exact color of the thing you are selling never changes. You are not misrepresenting the item. You are only changing the scene it sits in. Any business that respects its customers should treat that overlay step as non negotiable.
A furniture retailer shows the shift in practice
Let me ground this in one business so the abstract becomes concrete. Take a mid sized furniture retailer that sells sofas online. Furniture has a brutal photography problem, because to show a single sofa in five room styles you traditionally need five staged sets, a photographer, and a stylist. That is expensive, slow, and it caps how much of the catalog ever gets good imagery.
Here is how I would rebuild that with ComfyUI. The studio shoots each sofa once on a plain background, which is cheap and fast. Then a single reusable workflow takes over. Grounding Dino auto masks the sofa so it is protected. ControlNet holds its exact shape and camera angle. The model generates rooms around it: a bright Scandinavian living room, a warm rustic den, a sleek modern loft, all photoreal and all on brand. The overlay step drops the real sofa back on top at the end, so the fabric weave, the stitching, and the color stay exactly true to the physical product, which matters enormously when a customer is spending a thousand dollars on something they cannot touch first.
Now put numbers on it. A retailer who used to manage a dozen new lifestyle shots a week could realistically push well past a hundred, because the marginal cost of each new scene is a few minutes of GPU time rather than a booked set. The whole catalog gets covered in every room style, and a seasonal refresh that once took a quarter takes a few days. Those richer product pages tend to convert better, which is the same reason strong creative lifts the cost per lead on Facebook and Instagram ad campaigns: better images sell harder. That same library of scene variations also gives you fresh assets to feed Google Ads shopping and display placements without commissioning a single new shoot.
The move to make this quarter
If you sell physical products through images, the concrete step is not to panic buy a tool. It is to decide who owns your pipeline. Start by installing ComfyUI and learning the four basics in order: text to image, then image to image, then masking, then ControlNet. Do not attempt the full production line first. Get each block working on a single product until you understand what every node actually does, then assemble them into one workflow that takes a product photo plus a scene description and outputs a finished shot. Save that flow so a non designer on your team can run it by typing the product name and the style.
The entire stack is free and open source, so the real cost is the time to learn the nodes and a modest GPU, local or rented by the hour. The discipline that separates a professional result from an amateur one is keeping the product protected with masking and putting the real product back with the overlay, so your ads never distort what you actually ship. Those finished scenes should live in your CRM and website stack where product pages and retargeting audiences can pull from them automatically.
The economics that make a 40 person team optional
It helps to break down where the old cost actually went, because that is what the machine eats. A traditional product shoot bundles several expenses that had nothing to do with the product itself: the studio rental, the photographer's day rate, the stylist, the location travel, the model booking, and the editing time afterward. For a catalog of a few hundred items refreshed every season, those costs compound into a permanent line on the budget. A pipeline does not remove one of those costs. It removes nearly all of them at once, and it removes them for the second shot and the two hundredth shot at the same marginal price.
That is why the org chart shrank. When a 40 person asset and operations team gets cut to a handful, the remaining people are not doing less important work. They are doing the judgment work: deciding which scenes convert, which styles fit the brand, which markets to localize for. The machine absorbed the production labor, not the taste. A small business reading this should draw the same line. The goal is not to fire anyone. It is to move your people off the repetitive production grind and onto the decisions that actually move sales.
What separates a convincing shot from an obvious fake
The reason most businesses that dabble in AI imagery produce junk is that they stop at text to image and wonder why the output looks generic and slightly wrong. The convincing results in this pipeline come from the control layers, and each one solves a specific failure. Without masking, the product drifts and warps. Without ControlNet, the pose collapses into something anatomically odd. Without the Flux polish pass and the face LoRA, faces come out soft and uncanny, which is the single fastest way to signal fake. Without the final overlay, the product itself is subtly wrong, which is the one error a shopper will actually notice and hold against you.
So the quality is not luck. It is the discipline of stacking the controls in order and never skipping the protective steps. A brand that treats the overlay as optional to save time is the brand whose customers post that the sofa they received does not match the photo. A brand that treats it as sacred is the one that can generate a hundred scenes a week and never once misrepresent what ships. That difference in discipline, not access to some secret model, is what actually separates the winners here.
The move is a decision about ownership, not a purchase
Because the entire stack is free, the competitive question is not who can afford the tool. Everyone can. The question is who commits to owning a clean pipeline before their competitors do. The brands clearing millions with tiny teams did not win by buying something. They won by building a repeatable system and running their whole catalog through it while others were still booking studios. That window is open to a small retailer right now, and it closes a little every quarter as more sellers build their own flows.
This is genuinely buildable in house once the workflow exists, and a capable operator can run it every day. The hard part is designing that first pipeline cleanly, so the product never warps and the output reads commercial rather than obviously synthetic. If you would rather have that workflow built, tested, and handed over ready to run across your full catalog, that is exactly the kind of system my team sets up, and a short call is the quickest way to see whether it fits your store.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
