GPT-5.4 Tested: What Actually Improved and What Still Falls Short
GPT-5.4 thinking adds native computer use, builds finished decks and spreadsheets from a single prompt, and cuts hallucinations by a third versus the prior model. Coding and default writing tone still have rough edges. Here is the honest picture for a real business.

GPT-5.4 thinking is the strongest model OpenAI has shipped, and it is also proof that best on paper and best for your business are two different sentences. I spent real time putting it through document work, spreadsheets, coding, and writing, and the honest picture is a model that is genuinely excellent at some things and still visibly rough at others. I am Madhuranjan Kumar, and instead of a verdict, I want to take this release apart facet by facet, because understanding where a tool is strong and where it is weak is what lets you use it well rather than trusting a launch chart and getting burned.
First, the lineup, because the naming is genuinely confusing. OpenAI released GPT-5.4 thinking and GPT-5.4 Pro on the same day, close behind GPT-5.3 instant. Instant answers immediately, thinking reasons before it responds, and Pro is reserved for research grade work most people rarely need. Notice that instant and thinking no longer share a version number, which is why the naming trips people up. For nearly every owner and team, the thinking model is the one that matters, so that is the one I am taking apart here.
The facet that changes the most: native computer use
The single biggest structural change is that GPT-5.4 is the first general purpose model with native computer use built in. It can perform web tasks, do data entry, and even handle email and calendars on its own, without routing through a separate agent model. For anyone who has tried to build automations, this removes a whole layer of plumbing. You used to need the general model for thinking and a separate agent model for doing. Now the base model does both.
The reason this matters more than a benchmark is that it lowers the barrier to real automation. Simple, repetitive web and inbox chores, the kind of work that was never worth buying a separate agent tool for, can now be handed to the model you already use. That does not mean you hand it your whole operation on day one. It means the distance between I wish something did this and something does this got shorter. The plumbing that used to eat a weekend is now baked in, and that is a bigger deal for small businesses than any single point on a comparison chart.

The facet businesses will feel first: finished artifacts from one prompt
The gains show up most clearly in knowledge work, and this is the facet most owners will notice within their first hour. From a single prompt, GPT-5.4 turned a research report into a clean fifteen slide deck, then redesigned it on a follow up while keeping the content intact. It also produced a downloadable spreadsheet with working formulas, summary pages, and charts. These are not rough outlines you then have to assemble. They are finished artifacts, the actual deliverables businesses spend hours building by hand.
Two related behaviors make this facet more usable than earlier attempts. First, you can follow up mid research without restarting the chat. Asking for fifteen sources instead of the original ten simply steered the run rather than interrupting it, which makes a long task feel collaborative instead of a one shot gamble. Second, there is an adjustable thinking effort control, so you can push from standard up to heavy only when a task genuinely warrants the extra time and depth. You spend the slow, expensive setting where it earns its keep and stay fast and cheap everywhere else. Together these turn the model from a text generator into something closer to an assistant that builds the thing you actually needed.

The facet that unlocks trust: a real cut in hallucinations
None of the above would matter if you could not trust the output, which is why the quietest number in this release might be the most important. OpenAI reports a 33 percent hallucination reduction versus the prior model. That reliability gain is exactly why the one prompt spreadsheets and research feel trustworthy enough to actually use, rather than being impressive demos you would never stake a deliverable on.
Here is the honest way to hold this facet. A 33 percent reduction is meaningful, but it is not zero. The right mental model is that GPT-5.4 is now reliable enough to produce a first draft of a real deliverable, on the condition that you still check the numbers. That condition is not a knock on the model. It is the correct posture for any automated work. Fewer fabrications means you can let it do more of the mechanical assembly, but the verification step stays yours. A model that makes fewer things up is a model you can lean on harder, not one you can stop watching.
The facets that still need a human: code and tone
A deep look has to be honest about the rough edges, and there are two. Coding improved but is not flawless. In testing, a one shot tool comparison app worked but had broken links and a non functional compare feature. A second coding test then ran perfectly on the first try. So the picture is genuinely better than before, but inconsistent, which means you review generated code rather than shipping it blind. On paper the general 5.4 thinking model now matches the dedicated coding model and searches tools more efficiently, using fewer tokens even at a slightly higher per token price, but on real tasks it still occasionally trips.
The second rough edge is writing tone, and this one is stubborn. Even with custom instructions, the default intros felt generic, and the model slipped em dashes into the writing despite explicit rules against them. For natural, non promotional writing, Gemini and Claude still felt better off the shelf. This is worth taking seriously, because tone is where a lot of customer facing work lives. The practical read is to keep more than one model in rotation. Use GPT-5.4 for its structured strengths and keep Gemini or Claude on hand for writing that needs a genuinely human voice. The strongest model overall is not automatically the strongest at every single job, and pretending otherwise is how you ship copy that sounds like a robot.
The facet for spreadsheet heavy businesses: ChatGPT for Excel
One more piece shipped alongside the model that deserves its own mention, because for some businesses it is the whole reason to care. A new ChatGPT for Excel add on reads your data and writes back into the sheet. You describe a change in plain language, add a quarter, adjust an assumption, and it builds or updates the model while preserving the existing formulas and structure. For any business that lives in spreadsheets, that is the difference between describing what you want and manually rebuilding it.
A worked example: an accounting firm's month end
Let me tie the facets together with an accounting firm, using illustrative numbers. A staff accountant hands GPT-5.4 a client's raw figures and asks for a cash flow model. The model returns a downloadable spreadsheet with working formulas, a summary tab, and charts, in minutes rather than the better part of a morning. With the ChatGPT for Excel add on, the accountant then describes a change in plain language, add a quarter, adjust an assumption, and the tool rewrites the sheet while keeping the existing structure intact. What was a two to three hour build becomes a twenty minute build plus a careful review.
When a partner needs to present, the same research notes become a clean fifteen slide deck from one prompt, then get refined on a follow up. The 33 percent hallucination cut matters most right here, because financial numbers cannot be wrong. So the workflow is build fast, then spot check every total and tie it back to source before anything leaves the office. The firm keeps its judgment and its review step, and simply removes the slow mechanical assembly. None of these time savings are guaranteed for your firm. They are the illustrative shape of what the strong facets deliver when you respect the weak ones. The same reliable, structured output is what makes these tools safe to point at customer facing assets too, the copy inside your CRM and website stack or the reporting behind Facebook and Instagram ad campaigns, always with a human checking the numbers.
The facet nobody advertises: how the naming confusion costs you
There is a practical trap in this release that has nothing to do with capability and everything to do with the lineup, and ignoring it will cost your team time. Instant and thinking no longer share a version number, so you now have GPT-5.3 instant sitting next to GPT-5.4 thinking and GPT-5.4 Pro, three models with overlapping names and very different behavior. Instant answers immediately, thinking reasons before it responds, and Pro is reserved for research grade work most people rarely need. A team that does not internalize this difference will keep grabbing the wrong one, using an instant model for a task that needed reasoning, or paying for Pro on a job standard thinking handles fine.
The fix is a facet of good adoption most people skip. Decide, in plain terms, which model each kind of task should use, and write it down where your team can see it. Quick factual questions go to instant. Anything that needs a plan, a first draft of a real deliverable, or a multi step build goes to 5.4 thinking. Genuinely hard research goes to Pro, sparingly. Without that simple mapping, the naming confusion becomes a daily tax, small each time but constant. This is the unglamorous side of adopting a frontier model. The intelligence is handed to you. The discipline of routing the right task to the right model is still yours to install, and it is exactly the kind of detail that separates a team getting real value from one that just has an expensive subscription and mixed results.
The facet to reuse: follow up without restarting
One more capability deserves its own attention because it changes how a task feels. You can add context while a regular chat is still researching, without restarting the run. Asking for fifteen sources instead of the original ten simply steered the work rather than interrupting it. That may sound minor, but it converts a long task from a one shot gamble into a collaboration. Instead of writing the perfect prompt up front, waiting, and starting over if it drifts, you can nudge the model mid stream. For real business research, where you often realize halfway through that you also need a competitor comparison or an extra data cut, that ability to redirect without losing progress is a genuine time saver, and it is the kind of thing you only discover by using the model on real work rather than reading the launch notes.
The honest bottom line
Put the facets together and GPT-5.4 is a real step up in agentic ability and reliability, not a one click replacement for human judgment on numbers, code, or tone. Its native computer use lowers the barrier to automation, its one prompt artifacts save real hours, and its hallucination cut makes those artifacts trustworthy enough to use. But its coding still needs review and its default writing still needs a human ear or a second model. The businesses that get the most from it will be the ones that use it for exactly what it is best at, structured first drafts, and keep a verify step and a backup model for the rest.
You can absolutely learn this workflow yourself with a few real tasks and some patience. If you would rather have the prompts, the spreadsheet templates, and the verify checklist set up for your exact business, and the right model chosen for each job rather than one model forced onto every task, that is the kind of thing an expert can build with you so you skip the trial and error.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
