AI DOERS
Book a Call
← All insightsFuture of Marketing

Multimodal AI 101: Giving Your Business Vision and Hearing

Modern AI can look at photos and listen to audio, not just read text. Here is what that means for a normal business, with a worked example for a hair salon.

Multimodal AI 101: Giving Your Business Vision and Hearing
Illustration: AI DOERS Studio

Multimodal AI crossed a quiet but significant threshold this year: the models are now good enough, cheap enough, and fast enough that a single-location small business can wire vision and audio understanding into a real daily workflow without a dedicated engineering team. The text-only assumption that defined the first wave of AI adoption is finished.

I am Madhuranjan Kumar, and I run an ads consultancy that helps business owners translate these shifts into concrete operational advantages before their competitors notice the window is open. The multimodal shift is one of the biggest ones I have watched happen in real time, so I want to walk through exactly what it means, why the timing is important, and how a hair salon can implement it this week with a clear before-and-after picture attached to actual time savings.

The shared space of meaning that ended the text-only assumption

The core technical fact behind multimodal AI is simple but its implications are large. Standard language models convert words into numbers, learn how those numbers relate to each other, and predict what comes next. A multimodal model does the same thing, but it also converts images and audio into numbers and places all of those different inputs into one shared numeric space. Because a photo of a client with shoulder-length copper hair and the phrase "copper balayage, shoulder length" end up close together in that shared space, the model can reason across both at the same time.

That is not the same as object detection, which is what most people imagine when they hear "AI looks at a photo." Object detection tells you a photo contains a chair and a person. A multimodal model reads the photo and the surrounding context together and tells you the person in the chair is wearing a headset, looks frustrated, and appears to be on a support call. It infers meaning the way a person does, from the whole scene rather than a list of tagged elements.

The practical consequence is that vision and language stop being separate skills the model has to switch between. You can hand it two photos and a question and it will compare the photos, notice what is different between them, and answer the question in a format you specify. That comparison ability is what makes it useful for the hair salon example I will get into below, because the entire value of the workflow rests on the model seeing both a current photo and a target photo at the same time and reasoning about the gap between them.

Open tools like LLaVA let you test this today for free. LLaVA stands for large language and vision assistant. You upload an image, ask a question, and get a plain-language answer within seconds. The hosted commercial models go further and produce structured output, follow specific formatting instructions, and handle batches of images consistently. The current generation produces specific, usable output rather than a vague paragraph, which is what finally makes it worth building into a real workflow rather than just demonstrating in a lab.

How it works

Why businesses that depend on visual intake are the first to feel this

Every business has at least one workflow where a person looks at something visual and then writes something about it. A property manager walks a unit and writes a condition report. An insurance adjuster reviews claim photos and writes a summary. A retail buyer inspects product images and writes a quality note. A front desk team at a clinic looks at patient intake photos and flags what needs a closer look. A salon coordinator receives inspiration images and decides whether the request is achievable in the time booked.

All of these workflows share the same structure: visual input, human judgment, text output. That structure is exactly what multimodal AI automates at the first-draft level. The model looks at the photo, applies the judgment described in the prompt, and writes the first draft of the note. A human reviews it in seconds rather than spending several minutes producing it from scratch.

The businesses that feel this most immediately are those where that visual intake step is a bottleneck, where the volume of images is large enough that the cumulative time adds up, or where the inconsistency between different staff members writing notes about the same type of image creates downstream problems. When one coordinator writes a detailed inspection note and another writes three vague lines about the same property type, the quality of the downstream decision suffers. The model applies the same structured prompt every time, which removes that inconsistency entirely.

This is also directly relevant to how businesses run their marketing operations. A product company that can automatically score and describe its product photos has better raw material for SEO content and paid media. The description the model produces from a photo can feed directly into product page copy, ad creative briefs, and structured data without a copywriter manually writing each one. That kind of downstream leverage is where the real compounding happens, especially when the business is running active meta ads campaigns and needs consistent, accurate product descriptions across a large catalog.

Minutes spent per client consultation

The salon photo consultation: a concrete worked example with real time numbers

Here is the specific scenario that makes multimodal AI immediately tangible for a personal-services business. A salon owner who handles around forty client appointments per day currently spends a combined four to six minutes per client at the front desk stage: the coordinator receives a message with an inspiration image, reads it, tries to describe it to the stylist in words, the stylist asks a clarifying question, and the coordinator goes back and forth before the appointment even begins. At forty clients a day that is between two hundred forty and three hundred sixty minutes, which is three to six staff hours, going into a single intake step.

With a multimodal model in place, the workflow changes to this. Before the appointment, the client texts a photo of their current hair and a screenshot of the style they want. The model receives both photos and runs them through a structured prompt. The prompt asks four specific questions: what is the current length relative to the shoulder, what is the approximate base color and any visible highlight placement, what is the target style and color in the second photo, and what visible gaps exist between the current state and the target that would require additional steps such as a lightening treatment or an extended cut. The model returns a structured brief in plain text within a few seconds.

That brief drops into the booking note automatically. The stylist sees it when they open the appointment. The coordinator can confirm the appointment time and price in a single reply to the client because they already know from the brief whether the target is achievable in the booked slot. No back-and-forth. No "let me check with the stylist."

In the time tracking the salon owner ran over two weeks after setting this up, the per-client intake time at the front desk dropped from an average of five minutes to just over a minute, because the brief was already written and the coordinator was confirming a decision rather than gathering information. Across forty clients a day, that is one hundred sixty minutes recovered daily, which over a five-day week adds up to more than thirteen staff hours per week redirected away from information-gathering and toward actual client service.

The stylist benefit is equally concrete. A stylist who has read a clear brief before the client sits down can open the conversation at the decision stage rather than the discovery stage. The first two minutes of the chair conversation move faster, the scope is clearer, and there are fewer late-stage surprises where a client shows an inspiration image the stylist is seeing for the first time. Over a full day, that compresses the average appointment duration slightly and makes it possible to fit one or two additional clients into the same schedule without the day feeling rushed.

The salon owner in this example did not change their booking platform, did not hire a developer, and did not buy a new software subscription. They connected a vision-capable model to the messaging channel clients already use, wrote a structured prompt, and set up an automatic step that drops the brief into the booking note. The total setup time was one afternoon.

The specific shift in reasoning ability that makes this possible now

Two years ago, a version of this setup would have failed in a frustrating way. The image models of that period either hallucinated details that were not in the photo, described things in ways too vague to act on, or required careful prompting that produced different results each time. The salon owner would have tested it, gotten inconsistent output across ten client photos, and concluded it was not ready. They would have been right.

What changed is that the current generation of multimodal models produces consistent, structured output when given a structured prompt. The reliability crossed a threshold that makes it worth building into a real workflow rather than treating as a prototype. The models now handle poor lighting, unusual angles, and mixed-quality phone photos without falling apart, which is the actual condition you work with when clients are sending images from their camera roll.

The cost also dropped to a point where the arithmetic is easy. Processing a photo through a hosted multimodal model costs a fraction of a cent per call. Processing forty photos a day costs less per month than a single hour of part-time staff time. The operational savings from the five-minutes-to-one-minute intake improvement are orders of magnitude larger than the model cost. That gap did not exist at this scale two years ago, which is why the right moment to build these workflows is now rather than waiting to see whether adoption becomes mainstream first.

I track these windows closely because timing is the variable most business owners underestimate. When a capability becomes reliable and affordable at the same time, there is typically a twelve to eighteen month period where early movers build a meaningful operational lead before competitors realize what changed. The multimodal shift is inside that window right now. Most service businesses are still running visual intake entirely by hand, not because it is better but because they have not yet connected the capability to the specific workflow where it would save them time.

What this means for businesses running paid media and CRM operations

The multimodal capability does not sit in isolation from the rest of a business's marketing operation. It feeds into everything downstream. A service business that builds multimodal intake is generating structured, accurate data about every client and every request as a byproduct of the intake process. That data becomes the raw material for better Google Ads targeting, more accurate audience segmentation, and more relevant ad creative because the business now has real, structured descriptions of what their clients actually want rather than a pile of unread text messages.

For a salon specifically, the structured briefs accumulate into a dataset over weeks and months. The salon owner can look at that dataset and see which services are most requested, which inspiration images clients keep referencing, and what the most common gap is between what clients want and what takes longer than the standard appointment allows. That information is actionable for scheduling, pricing, and capacity planning. It is also the kind of first-party data that makes web and CRM operations more precise, because the intake data is already structured and can be pushed directly into a CRM field without anyone manually transcribing a text message.

The connection to paid media is direct. A business that knows its most-requested service types from the intake data can build ad campaigns around those specific services with accurate, outcome-based copy rather than generic messaging. The model that wrote the consultation brief can also write the first draft of the ad copy for that service category, because it already has language describing what clients want and what the transformation looks like. That loop, from intake image to structured brief to CRM record to ad creative, is what multimodal AI makes possible when you think about it as a data layer rather than a single point tool.

Three implementation mistakes that slow down the value

The first mistake is writing a vague prompt and expecting the model to fill in what you did not specify. A prompt like "describe this photo" produces a general paragraph that is rarely useful in a real workflow. The model follows the prompt precisely. If you tell it to produce a structured brief with four specific fields in a specific order, it produces exactly that, consistently. If you ask it to describe what it sees, it describes what it sees in whatever way seems most natural to it, which changes between photos and produces output that is hard to use in a downstream system. Write the prompt as if you are writing a job description for a new employee who is extremely literal: tell it exactly what to look for, what to ignore, and what format to return the answer in.

The second mistake is skipping human review on output that carries real weight. For a pre-appointment consultation note at a salon, the stakes are low and the stylist will naturally read the brief before the appointment. For anything with professional, clinical, or financial consequences, keep a person in the review loop. The model reads photos well but it occasionally misreads something in poor lighting or at an unusual angle. That is fine when a human reviews the output and catches the error. It becomes a problem when the model's output is treated as final without any human check.

The third mistake is abandoning the workflow after the first batch of inconsistent output. Most prompts need two or three small adjustments before they produce reliable results across a real set of images. The typical pattern is to test on twenty images, identify the two or three cases where the output missed something important, adjust one thing in the prompt to address those cases, and test again. Most workflows reach a consistent state within three iterations. Business owners who stop after the first imperfect batch never see how well the system performs once the prompt is matched to the actual variability in their image type.

The build is smaller than most business owners expect

Setting up a multimodal intake workflow for a service business involves four components. The first is a vision-capable model you can call via API, either a hosted commercial model or an open-source option you run yourself. The second is a structured prompt specific to your use case, written to return the exact fields your team needs in the exact format your downstream system expects. The third is a connection between the image source and the model, so photos that arrive via text, email, or an upload form flow into the model automatically without anyone copying files. The fourth is a connection between the model's output and wherever your team already looks, such as a booking note, a CRM field, or a shared document.

For a salon, the image source is already a messaging channel clients use, the booking note is already where the stylist looks before an appointment, and the connection between the two is a small automated step that takes an afternoon to set up. The salon owner does not need to replace any existing software. They are adding a step in the middle of a workflow that already exists, between the client message and the booking note, and that step now produces a structured brief instead of leaving a raw image that someone has to read and interpret by hand.

Start with one workflow. Prove it saves time. Then add the next. That progression is faster and more reliable than trying to automate several workflows at once, and it gives you real output to evaluate rather than a plan that looks good on paper. The salon photo consultation is the right first workflow for a personal-services business because the value is immediate, the prompt is simple to write, and the improvement in the stylist's experience is visible within the first week of running the system.

If you want to map out which visual intake workflows in your business are the highest-value targets and how to connect them to your existing CRM and paid media operation, that is exactly the kind of planning session I run for clients. The link below goes directly to the booking page.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
Multimodal AI 101: Giving Your Business Vision and Hearing | AI Doers