AI DOERS
Book a Call
← All insightsFuture of Marketing

Why Rivian's CEO Says Physical-World AI Will Outrun the Chatbots

RJ Scaringe argues the largest AI shift is machines acting in the real world, not text on a screen. Self-driving cars are the first at-scale deployment of physical AI, driven by a fleet data flywheel and an end-to-end neural net that takes vehicles from hands-off to fully driverless within years.

Why Rivian's CEO Says Physical-World AI Will Outrun the Chatbots
Illustration: AI DOERS Studio

Almost all of the AI conversation happens on a screen. Chatbots writing emails, models generating images, assistants drafting text. The CEO of Rivian and Mind Robotics thinks that whole framing misses the larger shift, and I, Madhuranjan Kumar, have been sitting with his argument because it reframes where the real value is heading. His claim is simple and unsettling: intelligence that perceives and acts in the physical world will reshape society far more than any text generator, and the first place it lands at scale is the car in your driveway. Once you accept that framing, a lot of the noise about which chatbot is winning starts to look like a sideshow.

The screen was only ever the warm-up

The reason physical-world AI matters more is that it touches the parts of the economy the chatbots never could. A model that writes a marketing email is useful. A model that drives a vehicle, moves goods through a warehouse, or runs a machine on a factory floor is operating in the layer where most of the world's actual work and cost live. His view is that the next five years of progress in that layer will look unimaginable compared to the last five, and the reason is not hype. It is an architecture change that has already proven itself once, in the language models everyone is fixated on, and is now being pointed at the physical world.

That is the throughline worth holding onto. The same leap that made chatbots suddenly good is now happening to machines that move. If you understood why language models jumped, you already understand why cars are about to, and why the businesses that grasp the underlying pattern will be positioned for the wave after this one.

How it works (short)

From human-written rules to a system that learns

To see what changed, look at how self-driving used to work. Up through the early 2020s, an autonomous car ran a perception stack that classified objects in view, attached velocity and acceleration to each one, and handed all of it to a planner built on human-coded rules. Engineers were literally writing out how the car should behave near a semi truck versus a compact car, city by city, edge case by edge case. It took enormous teams, and it broke the moment the environment or the sensors changed. It was brittle in the exact way hand-written rules always are.

Around 2022 the field shifted to an end-to-end neural network that learns to drive directly from data, the same transformer leap that powered large language models. Instead of coding the behavior, you show the system millions of examples and let it learn the behavior. Nobody writes the rule for the semi truck anymore. The model absorbs it. This is the pivot that makes everything else possible, and it is why he frames autonomy as a data problem now rather than an engineering-hours problem. The car got smart the same way the chatbot did, by learning from examples instead of being told.

Long-range sensor cost over a decade

The flywheel is the whole moat

Here is the part every business owner should tattoo somewhere, because it transfers far beyond cars. A learning system is only as good as the data feeding it, and that is where the real defensibility lives. Rivian's fleet streams millions of miles of real driving, and a triggering system flags the interesting moments, a hard brake, a collision, a maneuver the model did not predict. The detail that stopped me cold is that while a person drives, the model runs in parallel on the device and quietly compares what it would have done against what the human actually did. Every surprising human decision becomes a training signal. Multiply that across thousands and soon millions of vehicles, and the system develops a robust, human-like feel that no amount of hand-coding could produce.

Sit with the structure of that loop, because it is the reusable idea. You capture real data from something acting in the world. You trigger on the moments worth learning from, so you are not drowning in everything. You compare what happened against what the best version would have done. You feed every gap back as an improvement. That is the engine behind physical-world AI, and it is completely portable. A logistics company can run it on delivery routes. A property manager can run it on maintenance calls. A field-service business can run it on job tickets, and the gap between its best worker and its average one becomes a precise map of what to fix. The technology in the car is exotic. The loop underneath it is not.

Cost stopped being the excuse

The usual objection to any of this is that the hardware is too expensive, and he dismantles it. A long-range lidar sensor that cost around thirty thousand dollars a decade ago runs a couple hundred dollars today. The expensive part of a self-driving car is no longer the sensors, it is the inference brain doing the thinking. And in a nice twist, that cheap lidar earns its keep twice, because beyond redundancy it provides the ground truth that teaches the vision models what hard-to-see objects are at a distance. He notes that even companies betting on cameras run lidar-equipped cars to train their own ground-truth fleets. The sensing got cheap. The intelligence is where the money and the moat moved.

The same collapse in cost is happening across the tools an ordinary business would use to build its own version of this loop. What required a specialist team a few years ago now runs on software most companies can afford, which is exactly why the pattern is worth learning now rather than waiting for it to feel inevitable.

Optimize for the task, not for the shape of a human

There is a second idea in his thinking that I keep returning to, and it applies to automation of any kind. He points out that in a factory, ninety-five percent of the work happens within a ten-foot radius, so building a humanoid robot with legs adds cost and complexity for almost no gain. The fastest animals are not shaped like people. Optimized machines will take many forms, not one humanoid mold. He even reframes the car itself as a robot: rather than build a humanoid to sit in the driver's seat and turn a wheel, you mount perception around the vehicle and actuate the controls electronically. The car is the robot.

The lesson for a business is to design your automation around the actual job, not around copying how a person happens to do it. Owners waste enormous effort trying to build a digital version of a human employee doing a human-shaped process, when the real win is to rethink the process for the machine. You do not need one bloated all-in-one platform that mimics a person clicking through screens. You need small tools shaped to the specific jobs that matter, wired together. That is the same instinct that makes the physical robots take strange, efficient shapes instead of walking around on two legs for the sake of familiarity.

What the flywheel looks like inside one business

Let me make it concrete with a single worked example, using illustrative numbers. Take an HVAC company, because its most valuable asset is hiding in plain sight and almost nobody mines it: the data sitting in completed job tickets. Here is how I would build the flywheel he describes, scaled down to a business you can run today.

First, capture the real data. Every diagnosis, every part replaced, time on site, and outcome goes into one clean place instead of scattered paperwork. Second, define the trigger events, the same way the car flags a hard brake. A callback within thirty days, a job that ran three times longer than estimated, or a quote the customer rejected are the moments worth learning from, and there might be forty or fifty of those a month worth a closer look rather than every single ticket. Third, and this is the step most operators skip, compare what happened against what should have happened. When a senior tech and a junior tech handle the same fault code differently, that gap is a training signal, exactly like the model-versus-human comparison running inside the car.

Over a single season, that loop turns into a simple internal guide that makes every tech diagnose more like your best one, plus a clear read on which jobs and which neighborhoods are actually profitable. Say the shop runs six hundred jobs in a season and the review surfaces that fifteen percent of callbacks trace to one avoidable diagnostic miss. Closing that gap might save ninety repeat visits a year, and at a loaded cost of two hundred dollars per truck roll that is eighteen thousand dollars recovered from data the company was already generating and throwing away. Those numbers are illustrative, but the mechanism is real: the operation quietly teaches itself. That same clean data foundation then makes everything downstream sharper, from the targeting on Facebook and Instagram ad campaigns to the follow-up automation living in the CRM and website stack, because you finally know which jobs and which neighborhoods are worth chasing.

The company keeps installing and repairing systems exactly as before. The difference is that every job now makes the next one a little faster and a little more profitable, which is the entire promise of a flywheel.

The durable edge is curiosity

He closes on a point that has nothing to do with sensors. He thinks the next ten to fifteen years may be the most important chapter in the whole history book, and the right response is not to teach fixed career skills that are changing fast, but to teach curiosity and the habit of questioning how the work is really done. I think that is exactly right for business owners too. The tools will keep changing. The loop of capture, trigger, compare, and feed back will not, and neither will the value of an owner willing to ask why the work happens the way it does.

Why most businesses already have the data and waste it

The uncomfortable truth in his argument is that the moat is not the algorithm, it is the data, and most businesses are sitting on a fortune of it that they actively throw away. Rivian's advantage is not a secret model, since the end-to-end neural net approach is broadly known now. Its advantage is a fleet generating millions of real miles that a competitor cannot easily replicate. The intelligence is downstream of the data, and the data comes from operating in the real world at scale, which is the one thing an established business already does every single day.

A service company completes hundreds of jobs a year, each one a rich record of a real problem, a real diagnosis, a real outcome, and a real customer reaction. A clinic runs thousands of patient interactions. A retailer processes a stream of purchases, returns, and questions. Every one of these is the equivalent of a mile of driving data, a real event with a signal buried in it. And almost universally, that signal evaporates the moment the job is done, because it was never captured in a way you could learn from. The paperwork gets filed, the ticket gets closed, and the lesson inside it is lost. That is a business burning its own moat for fuel.

The reason this happens is that capturing data feels like overhead with no immediate payoff. Nobody closes a job faster because they logged it cleanly. The return only appears later, when the accumulated records reveal a pattern no single job could show, which technician consistently produces callbacks, which neighborhood quietly loses money, which quote wording wins and which loses. Those patterns are invisible in any one event and obvious across a season of them, but only if you kept the events in a form you can actually query. The discipline to capture without an immediate reward is exactly the discipline that separates the businesses that compound from the ones that stay flat.

You do not need a fleet or a neural net to start. You need one clean stream of the data you already generate, a habit of flagging the moments worth learning from, and the willingness to compare what happened against what your best version would have done. The technology to analyze it has gotten cheap enough that the bottleneck is no longer tooling, it is the decision to stop discarding the record. The businesses that make that decision this year will have a full season of learnable history by next year. The ones that wait will be starting from zero while their competitors are already three cycles into the loop, and in a compounding system that head start is very hard to catch.

You can build this flywheel yourself, and the first version is mostly discipline rather than expensive software. If you would rather have someone set up the data capture, define the triggers, and turn your everyday operation into a system that compounds, that is exactly the kind of build I do for clients, and you can bring me in to handle it.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
Why Rivian's CEO Says Physical-World AI Will Outrun the Chatbots | AI Doers