Ep 728: GPT-5.4 Released: 7 Takeaways you need to know about Openai’s New model

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


GPT-5.4’s Strategic Shift: What OpenAI’s Latest Model Means for Business Value

OpenAI’s newly released GPT-5.4 model isn’t just an upgrade in speed and accuracy—it represents a tactical move targeting core work systems, enhanced tool integration, and less headline-focused value for business innovation. The model’s features and underlying improvements point explicitly to immediate and practical advantages for organizations leveraging AI at scale—moving well beyond basic chatbot functionality into real-world, outcome-driven workflows. Below are the key insights and business implications, organized for maximum clarity and connection to business value.

GPT-5.4 Release: Purpose-Built for Work, Not Just Benchmarks

GPT-5.4 and its Pro variant are not traditional model upgrades. Instead, this release signals a shift toward enabling platforms for genuine productivity. GPT-5.4 is engineered to excel within business environments, offering new computer use abilities, integrated tools, and document handling upgrades ([00:01:21]-[00:01:30]). This development positions OpenAI to compete directly with groups like Anthropic, especially through integrated workflows and long-running work sessions that were not previously standard for AI deployment.

Computer Use Capabilities: Business System Integration

Unlike earlier models, GPT-5.4 introduces native computer use and browser tasks—directly enabling workflow automation and data handling inside an organization’s ecosystems ([00:05:01]-[00:05:56]). Markedly, the model supports a one million token context window (via API and Codex), which, while not yet available in the ChatGPT interface, allows for unprecedented project and document processing capabilities in backend operations. The advances extend to more intelligent tool search, token efficiency, and real-time coding, all with lower response times and error rates.

Benchmarks: Translating Metrics Into Business Reliability

While most AI benchmarks are designed to surface technical supremacy, business leaders should focus on what influences operations. GPT-5.4 reduces hallucinations by 33% and shows substantial accuracy gains across 44 occupations and industries ([00:04:19]-[00:04:37]), which directly translates to fewer workflow errors and more reliable deliverables. Notably, OpenAI now leads in critical areas such as OS proficiency, web competencies, real-world math, and tool usage ([00:07:39]-[00:07:46]).

Amid the noise of benchmarks, the new GDP val (General Deliverable Performance evaluation) emerges as most significant: it assesses the model’s ability to complete actual human tasks—like creating spreadsheets and presentations—matching or surpassing expert humans 82% of the time ([00:34:48]-[00:36:20]). This represents an immediate and measurable impact on workforce effectiveness, not hypothetical tasks.

Codex as a Requirement: From Developer Tool to Business Co-Worker

GPT-5.4’s evolution makes Codex essential not just for developers, but also for nontechnical staff ([00:16:24]-[00:18:08]). Codex now acts as a desktop co-worker, performing non-code tasks and coordinating integrated workflows. The Playwright Interactive integration enables Codex to control browsers, manage machine access, and serve as a functional agent, well beyond legacy coding tools. For organizations, that means more employees—regardless of technical background—can activate automation without having to switch apps or retrain for new interfaces.

Analyst and Data Roles: Direct Productivity Threat and Opportunity

GPT-5.4’s readiness to automate spreadsheet creation, real-time analytics, research synthesis, and document delivery signals a measurable threat to traditional analyst roles ([00:19:15]-[00:21:21]). This update introduces a dedicated Excel integration and soon, Google Sheets. As models now generate actionable business documents and complex data presentations on demand, legacy “per-seat” software faces direct risk. Organizations with heavy consulting or data analyst cost centers should immediately reevaluate spend and processes, reallocating resources towards more strategic tasks.

Model Naming Confusion: A Hidden Cost to Enterprises

One subtle but critical point: OpenAI’s ongoing challenges with model naming continue to confuse large user bases ([00:08:51]-[00:12:33]). Mismatched versions and the proliferation of nearly indistinguishable models can create training gaps, process errors, and lost productivity as end-users struggle to determine what capabilities are actually available. Enterprises must invest in clearer internal training and documentation to ensure tools are used optimally and that nontechnical staff do not default to inferior models.

Thinking Tiers: Systems Approach Over Chatbots

With GPT-5.4, the “thinking” modes available—especially on paid plans—shift interactions from simple Q&A to agentic, system-level operations ([00:23:03]-[00:24:48]). Depending on the subscription tier, users can access up to four levels of thinking power, giving organizations flexibility to match intelligence (and billing) to task complexity. Notably, the gap between base and pro thinking has shrunk, potentially lowering the cost-of-entry to advanced workflows for small businesses.

The Business Value Bottom Line: Economically Valuable Work, Not Hype

The most overlooked insight is that GPT-5.4’s improvements should not be measured primarily by traditional benchmarks, but by GDP val: its ability to generate real, expert-level deliverables in a business context ([00:30:09]-[00:36:31]). This is not incremental. A model that can reliably create documents, analyze data, and present findings at or above an expert human’s level is reshaping not just workflow, but operational economics.

Organizations integrating GPT-5.4 will quickly experience workforce amplification rather than mere automation. Consulting, research-driven projects, and analyst-heavy processes should reassess their reliance on human-only tasks and explore new productivity frontiers unlocked by these agentic systems.

Next Steps: Preparing for Systemic AI Adoption

OpenAI’s GPT-5.4 is a direct signal that “chatbots” are finished; future-ready organizations will be those that adapt to agentic, system-level, co-working platforms—where work is completed, not just discussed. Early investment in staff upskilling, workflow redesign, and intelligent model selection will determine who thrives as these advanced AI capabilities become baseline technology.

Decision-makers should move now, not only to harness GPT-5.4’s advanced features, but to set a playbook for integrating subsequent business-focused AI advances—ensuring the organization is not caught in a perpetual game of catch-up.



Topics Covered in This Episode:

  1. GPT-5.4 Pro and Thinking Model Overview
  2. OpenAI’s Computer Use and Tool Agent Upgrades
  3. GPT-5.4’s 1,000,000 Token Context Window
  4. Model Naming Confusion in OpenAI Lineup
  5. OpenAI vs Anthropic AI Model Rivalry
  6. Codex Desktop App and Non-Coding Uses
  7. GPT-5.4’s Native Spreadsheet and Excel Integration
  8. Benchmark Performance: GDP-Val and Real Work Outputs
  9. Pro/Thinking Model Tiers and User Impact
  10. State-of-the-Art Agentic Workflow Abilities


Episode Transcript 


Jordan Wilson [00:00:16]:
OpenAI just dropped their latest model in GPT five four thinking and pro, and it's obviously state of the art in performance and benchmarks. Yeah. We knew that. And technical talk, that means it's just a really, really good model. But this wasn't just a model update from OpenAI. It was a flex. GPT five four feels like OpenAI is making a direct play at developers, researchers, and anyone building serious AI workflows. And if that kinda sounds like a flex against Anthropic, you're right.

Jordan Wilson [00:00:53]:
But let's leave behind the competitive side of this GPT five four release, which comes on the heels of impressive updates from both anthropic and Google. But between the benchmarks and features and tokens from the announcement, I think there's something much bigger in terms of takeaways. There's two stories here. The public story is okay, impressive new model, but the real story, I think, is okay, bold new direction. G b t five four shows OpenAI is leaning harder into long running work, better tools, deeper research, and computer use. That means the gap between chatbot and work system is officially dead with this release. So today, I want to unpack why this launch feels like a step towards systems that work for you, not just smart models that you talk to. So on today's show, here's what you're going to learn and what we're gonna go over.

Jordan Wilson [00:01:54]:
Quickly separate what's important and what's just marketing when it comes to these new models, the GPT five four thinking and GPT five four pro. We're gonna break down the meaningful benchmarks in simple language and how they'll impact your use. Then we're gonna dish out what I think are seven of the more important or not talked about takeaways that you need to know about this latest update. And then I'm gonna reveal at the end the one most important takeaway that no one is really talking about. Alright? Let's do this thing. Welcome to Everyday AI. My name is Jordan Wilson. If you're new here, well, you can guess by the name.

Jordan Wilson [00:02:33]:
We do this every day. It's an unedited, unscripted live streaming podcast helping everyday business leaders make sense of the nonstop updates because, yeah, they're nonstop. I tell you what matters, what doesn't, and you take that information to be the smartest person in AI at your company. So it starts here, but if you really want the goodies, that's on our website, youreverydayai.com. Make sure you go sign up for the free daily newsletter. And, hey, while you're there, also just go subscribe to our podcast. Right? You can always go find the links on there. So let's talk about, the basics here.

Jordan Wilson [00:03:08]:
What's new in GPT five four. If you're a little confused by the names and everything, well, let's start there with, well, actually, let's not start there. Let's give you the basics first, and then I'll give you some of my takeaways. So on paper, I mean, this is OpenAI's most capable and efficient model. And right now, it is available in chat g b t, the API, and codex platforms if you are a paid subscriber. Alright? If you're a free subscriber, you're not gonna see G B t five four, at least not for now and maybe not for a while. We'll see. G b t five four Pro offers maximum performance for complex tasks and professional use cases, and it integrates top coding abilities, tool use, and document handling improvements.

Jordan Wilson [00:03:54]:
Let's talk a little bit about real world accuracy and just work. Well, one of the things that is great at now is excelling at spreadsheets, presentations, and creating documents with higher factual accuracy. So cutting, hallucinations down according to OpenAI by 33% and just improving greatly both visually and in different, benchmarks on just doing real work. Right? Literally creating presentations, documents, and spreadsheets, all with a model. Right? Not even having to having to go into a different mode. A lot of people still don't know that about chat g p t and its base models, but, yeah, it does all that out of the box. And right now, it does outperform previous versions in knowledge work across 44 occupations and industries, and it significantly reduces errors and hallucinations for more reliable outputs, compared to its previous models. But, I mean, some of the biggest advancements are on the computer and tool use side, and we will go over these benchmarks here in a second.

Jordan Wilson [00:05:01]:
So this is OpenAI's first model with native computer use capabilities for desktop and browser tasks. So yeah. If you are needing some of that, right, whether you're a developer or wanting to use it on the front end, right, a lot of, recently, it was, entropic. Very, very recently, I think Google Gemini and their three one pro came along. And now, you know, OpenAI is kind of leading with this. Right? I was actually surprised, at least from their marketing and messaging that they led so heavily, into tool search. Right? They talked multiple times about improved tool search, you know, cutting down on the amount of tokens that it takes to even grab tools. So a much more technical angle in this release, but really pumping up browser use, computer use, and just being able to handle a larger tool ecosystem with lower token usage and faster responses.

Jordan Wilson [00:05:57]:
And the big number, a million. That is a million token context window, but not inside chat GPT. That would be great. I don't think we're gonna get that, really from any provider, anytime soon. But it is available in the API and in some use cases in codex as well. That's one of the reasons I'm using codex all the time. But OpenAI does say that there's enhanced safety, cybersecurity, and reasoning monitoring, for professional deployment. So like I said, you do have to be, on a paid plan to use this right now.

Jordan Wilson [00:06:34]:
And if you are using the older model, which is actually GPT by two. Yeah. Confusing. I know. That is going to sunset in June. Let's take a quick look at the benchmarks, and, I will get over into these a little bit here in a minute a little more. But, I mean, really good. Right? You don't have to be on the the live stream and see the, the benchmark screenshot from OpenAI.

Jordan Wilson [00:07:06]:
So, obviously, as with any, new release, you have to keep in mind, they're always gonna cherry pick what they're showing, what they're not showing. You know, sometimes there's, different versions of benchmarks. So, I mean, usually, when you see this from a company, you know, they're not gonna put every single benchmark even the ones that they, you know, maybe aren't the leaders in. So in this one, for the most part, you know, across OS world verified, WebAreada verified, GDP val, browser comp, Sweebench Pro, GPT QA diamond, frontier math. Right? All of those, you know, OpenAI. I think all of them except one. OpenAI is winning against their competitors. In this instance, they're showing it against Claude Opus four six and Gemini three one pro, which are the, most up to date and latest frontier offerings from their main competitors in anthropic in Google.

Jordan Wilson [00:08:00]:
So some pretty big jumps across the board. And the thing that we have to realize, and that's, you know, good to talk about now, because at least on the chart, if you're reading this left to right, right, we have GPT five four thinking pro, and then we have GPT five three codex. Right? But here's the thing. Almost every single chat GBT user, right? All 900,000,000 of us, no one was using that because it wasn't available inside of chat GBT. Right? So for the most part, people are gonna be a little confused because they were maybe, using g p t five two yesterday, and then it's g p t five two today. Alright. But, let's get into the seven takeaways because we're gonna start there where I just ended because takeaway number one, OpenAI has still not solved the model naming problem. Right? Back in, last summer, OpenAI CEO Sam Allman said, yeah.

Jordan Wilson [00:08:56]:
We're gonna solve the naming problem. And at first, it seemed like maybe they were. Right? Because at the time, you had models like GPT four, GPT four one, GPT four five. You had, O303 high, o one. Right? It was super confusing because you had two different classes of models. So same Altman said, alright. Well, we're gonna come out with the GPT five. It's gonna have this smart, you know, model router, and, you know, that's it.

Jordan Wilson [00:09:22]:
And you're just gonna go in there, and it's gonna route you to the right model you need. Obviously, that didn't work. GBT5Auto or GBT52Auto is still, an option. So we'll see what happens with that between, you know, free users, paid users. We'll get to that. But takeaway number one, it's more than about OpenAI being confused with with their model naming. It's bad. It's actually confusing consumers, probably hundreds of millions of them.

Jordan Wilson [00:09:55]:
Right. I do a lot of trainings, right? Corporate trainings, you know, virtual, whatever. I don't think I've ever met anyone that really knows what model is what, and that's probably a bad thing. Right. And I think this is something where Google has done probably a much better job in, you know, they have three one, You know? Go use three one. Right? I think, unfortunately, you know, OpenAI is probably the furthest behind in this. Anthropic is not, you know, not that much better, although they did get a little better by bringing all their latest models to four six, two weeks ago. But before that, you had a couple different tiers as well.

Jordan Wilson [00:10:39]:
But so not only is it confusing having multiple, but even the last, like, week or so in the last two months has also been confusing. Because like I said, we just jumped from five two to five four. So most people for the last couple of months, they've been using GPT five two. But there's all this talk of GPT five three, but everyone's like, where is it? Right? Everyone's like, I don't see this g b t five three codex model. Well, that's because it was only in codex. And then to make that even more confusing, earlier this week, OpenAI teased the GPT five four release. So we obviously knew this was coming, and then they released GPT five three instant the same day, which is really only for free users. Right? And that's all who should be using, GPT five three instant.

Jordan Wilson [00:11:29]:
So there was technically a GPT five three in chat GPT. That was the quote unquote latest model, but, like, for less than twenty four hours. So takeaway number one, this is confusing for consumers, for biz small businesses, enterprise across the whole AI landscape. Because like I said, I talked to so many people. And number one, most people don't even know. Right? The overwhelming majority of people use the default model, and that's different depending on, what plan you're on. And that's usually a very bad thing because you should always be using a thinking variety. Right? I tell people humans are impatient.

Jordan Wilson [00:12:12]:
Right? Because sometimes people, oh, I don't wanna use thinking or a high version of thinking. You know, I just want the answer. Okay. Well, you're gonna get an answer, and then you're gonna spend five times that trying to make it better. Right? Because a default, especially the GPT default models, not good. They're not. Right? Yes. Five three instance, a little better, but compared to what you have, they're very bad.

Jordan Wilson [00:12:33]:
Right? And that makes the whole lineup seem even more messy and confusing. All right. Takeaway number two, open AI is going for anthropic's throat with this one. All right. That's the first thing as I'm looking both at the benchmarks and in the marketing and how they're angling this, you know, like showing some of their use cases. Yeah. You know what? I don't I don't think the, the OpenAI team has has taken the, the recent kind of, pseudo rivalry, in some of the shots from Anthropic, too lightly. Is it, you know, coincidental timing that we got this release, like, twenty four hours after the, report of Anthropic CEO, you you know, in a leaked memo kind of taking some shots at OpenAI.

Jordan Wilson [00:13:27]:
Obviously, the, the Super Bowl commercial taking a shot at OpenAI, which is something Anthropic hasn't really been known for. Right? Anthropic, up until the last, like, five weeks, they've kind of been the the the good guy. Right? The good guy that's concerned about safety. And now it doesn't seem like that. Now it seems like they're trying to bully their way into relevance. Yes. I do love the anthropic models, especially over the last five or six months. I think they've gotten much better.

Jordan Wilson [00:13:55]:
You unfortunately have to pay $200 a month to get any utility out of them. That's beside the point. But they've really, I think, changed in their persona. Right? Maybe it's intentional. I don't know. Maybe they're they need to be the the loudest AI lab in the room because outside of, you know, us, quote, unquote us. Right? If you're listening to this podcast, you obviously know Claude. Right? Everyone knows Anthropic, but no one else does.

Jordan Wilson [00:14:21]:
Right? They have these studies, like, speaking of the Super Bowl commercial, I think it was, like, they had 7% recognition. No one knows. Outside of the the AI bubble that many of us choose to live in, no one knows Anthropic. Right? Most people know Google Gemini. Most people know, Copilot. Just about everyone knows, you know, OpenAI, Chat GPT, it's become synonymous with AI. Right? No one knows anthropic. So anthropic trying to, you know, bully their way maybe into a little more relevance probably wasn't the right move.

Jordan Wilson [00:14:54]:
And I do think that this model, the g p d five four, is going straight for their throat because everything that Anthropic has hung their hat on over the past eighteen months is exactly what OpenAI, updated in g b t five four. Right? Just, around tool usage, efficiency in those tool usage. My gosh. Like, everyone always goes crazy, which I get it. The models are great, but I paid $200 a month for Claude. Right? I also paid $200 a month for Chad GPT. I was doing some side by side testing. There's certain prompts.

Jordan Wilson [00:15:27]:
Right? Everyone's like, oh, you know, Claude agentic, blah, blah, blah. Right? A single prompt. I kid you not. And it crashes every single time because it runs over the context context window. The compaction inside Claude breaks it. Right? So for all this stuff about, oh, you know, Claude's, you know, the the the context window and tokens and the tool usage. Okay. Good.

Jordan Wilson [00:15:52]:
But long context, OpenAI is in a league of their own. Right? Especially when it comes to, long context with transparency in tool use and not breaking. So not just that, but OpenAI with five four improved token consumption in improving their ability to call those tools as well. So just really, making a harder play in long horizon tasks. So let's go to takeaway number three. Codex is becoming a requirement. Alright? So not only is that 1,000,000 token context window, something that you can take advantage of in codex, which is great. I think that takes the tool from, you know, a a nice, you know, desktop app.

Jordan Wilson [00:16:43]:
Right? They just released it for Windows, this week after last month, releasing it for a dedicated Mac app, for codex. Right? So I think it's gone from, okay. This is a software development tool, right, an IDE to, you know, great for vibe coding. So no. Now it's, like, you know, able to refactor entire code bases. But I think codex is becoming a requirement for nontechnical people. Even just for the right because at least right now, if you have a chat GBT paid account, you have codex. Right? And I think there's still a couple weeks where the limits are double.

Jordan Wilson [00:17:21]:
Right? I've yet to hit limits, and I know you're tired of me saying this. Yes. I'm always running codex. Like, I'm running codex right now. Every single time I'm recording, codex is running twenty four seven. I've never hit limits. It's absolutely wild. Yes.

Jordan Wilson [00:17:34]:
I'm on the $200 a month plan. There's double limits. But regardless, I'm doing so many nontechnical, non coding tasks in codecs. Right? I was actually chatting with someone at OpenAI, and I'm like, you guys need to push this. Right? Like, sometimes I just give advice. Sometimes companies ask me things, and I'm like, you guys need to push this because Claude is really pushing Cowork. Right? Anthropic has their, Claude Code and Claude Cowork, and I think codex does both. Right? But a lot of people are looking at codex like it is just the, you know, their their version of quad code, and it's so much more than that.

Jordan Wilson [00:18:08]:
And I think the new updates in 05/04 really emphasize that just with the computer use agent capabilities. Right? So right now, if you wanna take advantage of that, you're not seeing that inside of chat g p t. You know, I don't know if agent mode is going to, hopefully, eventually, it'll be updated with some of these, you know, the five four model and some of that computer use capabilities. But right now, if you wanna use that, it's either on the API or inside codecs. Right? Playwright Interactive, I'll probably do an entire show on it. You know, it's it's a new so they just came out with Playwright Interactive, but they've had the Playwright, CLI command line, tool. Right? But that's essentially a browser. People don't understand that.

Jordan Wilson [00:18:48]:
Right? Like, yes. Codecs has a browser that it can control. It can access your machine. So, I mean, at that point, it's much more than a coding tool. It is a coworker. Right? Too bad Claude got to that name first. Great name, by the way. But the new Playwright Interactive, kind of plug in for codex is great.

Jordan Wilson [00:19:08]:
So it pushes codecs into a more, I think, serious testing and debugging, and execution workflow. Alright. Takeaway number four, ChattGPT is coming for data and analyst roles. Alright. So analyst roles, data roles, not data analyst, but that's one of them too. Let's let's be honest. The original chat GBT couldn't add. Right.

Jordan Wilson [00:19:30]:
And even models a year ago, couldn't edit a spreadsheet. Alright? And you really had to know your way around prompt engineering to get it to save a spreadsheet. So not only do the models now, by default, they're agentic, and they can do all those things. People don't know. You don't gotta click a button. You can just be like, yo, g B t five four, go do all this research for me. Put it through my own personal, you know, lens via my memory and what you know about me, and go create spreadsheets and documents and PowerPoints, and it will literally do that, and they work, and they're there. Right? But they now also just released a dedicated Excel integration.

Jordan Wilson [00:20:08]:
Right? So another thing that Claude right? When Claude announced this or Anthropic, announced this and, you know, they've been coming out with these plugins and they and these skills. Yeah. Anthropic's been shipping. Great stuff. But it's been moving the markets. Right? You've had some legacy, you know, per seat software providers, you you know, some legacy financial institutions that have seen their their stocks crash, right, when Anthropic is releasing some of these plugins and, you know, some of these, like, their Excel integration. Okay. Well, Chad CPT just did that.

Jordan Wilson [00:20:39]:
I don't I'm not quite I'm not so quite sure about the the timing of this. I would have, I don't know. If it was me, I would've saved that for maybe not the same day because that's actually a big freaking deal that no one again, no one's really talking about that one. So, there is a dedicated app, for, the Excel integration with ChatGPT. That's huge. Right? Because the world runs on Excel. But they are also coming out with one soon for Google Sheets as well, which I'm looking forward to that one. Also, Gemini has gotten so much better.

Jordan Wilson [00:21:13]:
Right? Late twenty twenty five and early twenty twenty six. Like, Gemini and Sheets is actually great, but I'm still gonna use check GPT and Sheets as well. And I think what we're seeing here is it's shifting, the GPT models, especially with the GBD five four. It's shifting from like, oh, is this an AI tool or a junior researcher? Right? Like that was kind of the, the, maybe with five one and five two, maybe that's the conversation people are having. Right? Like, oh, at what point is this AI tool a junior researcher? Now I think we're way past that. I think it's like, okay. Is this a junior researcher or a senior researcher or a junior analyst or a senior analyst? Right? Another reason why I've been saying the consulting industry is gonna get absolutely smashed, because of this. Right? Alright.

Jordan Wilson [00:22:03]:
Takeaway five. Thinking models are much more than chat. Right? And I think with five four, I've I've I've only had a couple of hours, you know, to play with the model. I, unfortunately, didn't get early access like some people. So this is my takeaway from just a couple of hours. And it depends on what plan you're on, and let me explain that. If you're on the pro plan, there's four tiers of thinking. If you're on the regular $20 a month plan, there's two tiers.

Jordan Wilson [00:22:33]:
If you're on the free plan, you can click the little, light bulb icon, and I think you get, like, one of those a week or something like that. Right? But for the most part, if you're on the paid plan, the base $20 a month plan, which I think is what most people are on, you know, there's two thinking levels. And I think that using those in the older models, right, five one five two, it just felt like a smarter chat. Right? $5.04, it feels much more than that. Right now when you're using thinking, it feels like a system, not a smarter chat. And there's a couple, reasons for that, but it seems like OpenAI put a lot of emphasis on the thinking versions. Right? So, using five one and five two thinking versus five one and five two pro, the gap was enormous. Right? I don't feel that gap is as big.

Jordan Wilson [00:23:29]:
I mean, we'll see as more and more benchmarks come out. But using especially when you're on the pro version and you have four different tiers of thinking, you know, using the heaviest, thinking tier on the, pro. So not okay. This is confusing. Right? Not using GPT five four pro, but using the highest version of thinking, on the pro plan because it's an extra higher version. Right? But it felt like the premier model, but it was a thinking model. So I think that gap between the thinking tier and the pro tier actually closed. And I'm not saying that means that GBT five four, pro didn't get much better.

Jordan Wilson [00:24:10]:
It did. But the thinking again, I think the older versions were just smart chats. Now it feels like you are working with an agentic system. Couple updates, on that. Number one, you know, OpenAI did specifically say, that they made some improvements. So GPD five four thinking also now includes, deep web research. So a version seems like a specialized or a mini version of deep research, that kind of technology, which I think uses, like, dual models. That's available in the thinking mode.

Jordan Wilson [00:24:49]:
So, you you know, there's they've definitely put a little bit more, technology just into the thinking side, which is huge. Another update to thinking is you can steer it, which is really nice. Before that was something you could only do, in the pro version. So what that means, especially if you're, you know, giving it a lot of data and it's gonna do a lot of research and a lot of tool calling, which, again, we need to get more comfortable with this nontechnical people. All tool calling means is, oh, it's gonna use Python. Right? If you throw a lot of numbers at it, it's gonna use Python and write some code, or it's gonna, you know, use, you you know, web search. Right? That's all that means. You know, the harness and, you know, all that.

Jordan Wilson [00:25:34]:
You you know, people use all the fancy words. I just try to simplify it for you. Right? But it can take longer now, the GBT $5.04 thinking modes, because they have this new deep search built in. So it's actually nice that you can steer it. So you don't you know, if normally, it might take ten, fifteen minutes, and you're like, ah, freak. You know, I'm reading the chain of thought, and I'm reading how it's thinking, and I see it's going in the wrong direction. Normally, with the thinking models, you'd have to just wait or just say, alright. Well, I'm just gonna click cancel or, you know, there's an answer now button.

Jordan Wilson [00:26:09]:
Now you can steer it, which is really nice. Alright. Takeaway number six. Alright. I was kinda having fun with this one being a little cheeky, but I said five four is soda cool. Alright. State of the art computer using agent. Right.

Jordan Wilson [00:26:25]:
Five four pushes computer use agents into more complex cross application workflows. Alright. So, like I was saying earlier, opening eyes said, this is their first model with state of the art computer using agent abilities built in, right by default. So I kind of see this and view this. If you remember back to GPT four, and then you kind of went to GPT four. Oh, right? So, technically, GPT four used three different models to give you responses. And GPT four, which stood for Omni, meant it was all one model. So that's kind of what I'm seeing and feeling now with how OpenAI is starting to integrate computer use into their model.

Jordan Wilson [00:27:09]:
So, yeah, unfortunately, we may not actually, get to realize that in the old chat gbt.com, maybe until they update agent mode. Please open AI, update agent mode. So much. Right? Like, just just so much potential there. It just needs some love. Right? But by default, if we're talking about in codex, if we're talking about on the API side, that's important. So you may not touch this directly. Right? You may not touch their, you know, new built in computer use agent like I said.

Jordan Wilson [00:27:40]:
They they have, a demo of it. You can go play with it on, GitHub, you know, download the repo. You can do it that way. You can use it in codex. But, ultimately, this is for builders, but this is gonna be so much of the technology that we all use. Right? So that's what I'm excited for, is all of a sudden, this just made agents, number one, agents much smarter, but it also gave us nontechnical humans, I think the potential capabilities to direct agents or direct smarter agentic browsers, at scale because it is going to now be faster, better, and more token efficient. So huge win, that now OpenAI, again, going after Anthropic's lunch money. Right? OpenAI said, oh, okay.

Jordan Wilson [00:28:29]:
Anthropic, you wanna come come after our, you know, our decision to run ads with, you know, not super truthful Super Bowl ad. Ironic. Right? Okay. We're gonna come for your lunch money. So, yeah, Pretty, pretty impressive. All right. And speaking of state of the use computer, sorry, state of the art, computer use some of those benchmarks. Yeah.

Jordan Wilson [00:28:51]:
Just absolutely crushing it. And the noteworthy thing here, I think, is some of the jumps, from five two to five four. Because like I said, hardly no one used five three because it was in codex. Right? So now these are state of the art. These are topping the charts. So, in GDP valve, which I'll talk about here in a second, SweeBench Pro. Right? That's how good the model is at fixing real software bugs, state of the art. OS world verified, that's how well the model can use a computer like a person would.

Jordan Wilson [00:29:27]:
State of the art, tool a thon, that's how well the model uses, you know, tools correctly. Right? State of the art browser comp. I think that might be one they're, like, point one points behind someone else, but essentially state of the art. That's how well the model finds and uses information on the web. Right? So, essentially, anything with related to computer use tool calling, I mean, OpenAI, just crazy. Right? And this is not, at least, in my opinion, when you're looking at kind of the the three way race over the last, you know, year and a half. This is not something that OpenAI has been known for. So again, I think this is the bold new direction.

Jordan Wilson [00:30:09]:
And then last but not least, and this is both takeaway seven and the one thing I tease for the end, the thing that I think most people are overlooking. I know you're going to get tired if you're an avid listener. I'm sorry. I'm going to talk about GDP valve one more time. And here's the reason why I hate benchmarks. Everyone cares about them. Everyone talks about them. Everyone in the lab.

Jordan Wilson [00:30:34]:
Guess what? In the real world here in Chicago, Illinois, right where I'm from. Right? I think this is just like a Silicon Valley thing, maybe. I don't know. But that drives the narrative. Everyone's so concerned about these 50 benchmarks, so I gotta talk about them because if I don't, you know, I'm gonna get 50 emails about them. But I think my seventh takeaway, my last one here, is most benchmarks don't matter anymore, but I think one matters even more. And there's a couple reasons for that. Right? And I think GDP val is the one that matters more.

Jordan Wilson [00:31:07]:
It's number one, it's harder to gain, game. But I think that most frontier labs are just, you know, Benchmaxing right there. All they're doing is they're playing the game, you know, tweaking things just to get really good, you know, scores on certain benchmarks. Right? And guess what? Benchmarks don't pay bills. Benchmarks don't help us do work better. Right. In theory, you know, a leads to B B leads to C. But if the end goal is c that's GDP valve, that's the work getting done, not just, oh, here's this random test, you know, arc AGI.

Jordan Wilson [00:31:51]:
Right? All these, you know, humanity's last exam. Right? All these things that are, you know, these tests that are set up that are more just like, oh, here's a bunch of random things that are very hard that just aren't in training data. Right? That's not what we should be looking at models for and how good they are. We need to be looking at how good they are at creating economically valuable, viable work on their own, which is what GDP Val is. Right? It measures, how good a model is at creating deliverables in the same way a human would. I've talked about this a little bit, but let me just tell you how absolutely bonkers this is. Right. And I'm actually gonna go to my, my next slide first.

Jordan Wilson [00:32:35]:
All right. GPT four o. All right. So remember GPT four o, I'm trying to do the math here. Yeah. Ten months ago, nine or ten months ago, this was the best model. Alright? The best general purpose model. No one used the o models.

Jordan Wilson [00:32:51]:
I I I did. I love the o models. Right? O one zero three. No one used them. Everyone used g b t four o. G b t four o, if you're looking at this, GDP valve, stick with me here. Right? And this is just, a model would get, there's a benchmark and it's, you know, you have to go in and do real work like a human would. The whole thing, front to back.

Jordan Wilson [00:33:14]:
Alright? One, you know, one task. But it's creating spreadsheets. It's doing, you know, multistep research and creating something on the back end. And then it's judged by experts in the field. Right? So this is across 44 different, you know, real real world jobs, and then it's judged by a panel of experts. So, essentially, the 50% mark, that is the Paris the parity with an industry expert. Right? And then the expert also you know, there's groups of experts that judge both the expert human that submits the deliverable in the AI model. Right? So GPT four o, which was the best model ten months ago, got a 12% on that.

Jordan Wilson [00:34:01]:
Right? Not very good. Alright. And even GPT five high, still not very good. 38%. Okay. So why is this the benchmark that matters and why is GBT four, five, four, such a huge step? Number one, I was not expecting this Right? Because GBD five two had a 70%. Alright. So, wind tie rate.

Jordan Wilson [00:34:30]:
So 70% of the time it either won or it tied the expert human blindly judged by expert judges. In my 2026 AI prediction and roadmap series, I thought I was being kind of bold by saying, yeah, I think we'll get to 80%. Guess what? We already got there because GBD five, four pro got 82%. The more I like I've, I've, I've read this, this, benchmark, you know, the study that came along with it multiple times and the more and more I read about it and think about it and revisit it when new models come out, It just, to me, just talks about the sheer gap and it's a knowledge gap. It's an educational gap, and I'm gonna end with this, right, which I know is a weird way to end, a recap show about GBT five four. But I do think this ties into what is OpenAI's bold new direction. It is getting work done. Right? And that's what these models and that's what g b d four, sorry, g g d p valve, the benchmark shows.

Jordan Wilson [00:35:41]:
82% of the time, this new model wins or ties against a human. That's wild. And if that doesn't change right? So think think if if you've been in an industry ten or fifteen years, right, you're you're an expert and you sit down and you have a project. Right? You have to do some research. You have to use your smart brain, and then you have to create something of value. Right? A document, a spreadsheet, a PowerPoint presentation, and then a group of people are gonna judge it. You only have an 18% chance to beat GPT five, four pro. Somehow wild, right? Everyone's like, oh, these AI models are so dumb.

Jordan Wilson [00:36:31]:
That's what I'm saying. Where we've come in the past year is not normal. Right? Going from a year ago, models couldn't edit spreadsheets to now they're better than almost all experts. And I think the GPT five, four model, maybe, just maybe, might be the model, at least from OpenAI, that starts that conversation. And it moves away from chat GPT is a chatbot to, oh, chat. GPT is the place where work gets done. Alright. I hope this show was helpful.

Jordan Wilson [00:37:11]:
If it is, let me know, should we do a more hands on version of this on Wednesdays? Right? On Wednesdays, we do our AI at work on Wednesdays. So let me know if you actually wanna see some, you know, go under the hood with GPT54, test some of these things out. I've been having fun in the little time I've had testing so far. So I hope this was helpful. Make sure if you haven't already, go listen to our 2026 AI prediction and roadmap series. That's episode seven twelve and seven thirteen. I get a lot of people always, right, asking questions, emails a day, you know, asking me very questions that would take me a long time to answer. And I feel like a jerk sometimes, but I'm like, go listen to this episode.

Jordan Wilson [00:37:49]:
Right? I cover so much in there. If you listen to that and then go read the newsletters that come along with it, I guarantee you you're gonna be the smartest person in AI, in your company. Right? For the most part, unless you're working at Google or OpenAI. Right? So, and then when you're done doing that, make sure you go to youreverydayai.com. Sign up for the free daily newsletter. Thanks for tuning in. Hope to see you back later for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI