Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
GPT-5.4’s Hands-On Business Value: Five Features for Practical AI Use
OpenAI’s latest model, GPT-5.4, is not merely another incremental update. For those overseeing organizational AI adoption, several newly demonstrated features in GPT-5.4 translate directly into improved business outcomes. Here’s a granular look at what matters from the hands-on review, with detailed examples tailored for business value.
Interrupting Thinking Mode: Adaptive Task Management in AI
The “interrupting thinking mode” now available with GPT-5.4 empowers users to course-correct AI tasks while the model is actively processing. Unlike competitors, which force users to wait for completion—sometimes upwards of 30 minutes or more—GPT-5.4 enables users on paid plans to intervene if they realize critical information was omitted or if a new directive needs to be issued.
For organizations managing complex, multi-step workflows, this means that teams no longer need to accept subpar results or scrap extended computations. The ability to halt, adjust, and resume long-running AI operations minimizes wasted computational cycles and allows for real-time optimization of outputs, maximizing both productivity and accuracy.
AI Skills Access: Expanding Enterprise Utility Through Skills Integration
Skills, once a distinctive feature of Anthropic’s Claude, are now available for GPT-5.4 within OpenAI’s business and enterprise tiers. This direct integration allows businesses to leverage specialized utilities previously limited to coding or developer environments.
The transcript details that GPT-5.4 can now pair with skills inside the desktop app and command line tools, enhancing the flexibility of AI use in practical settings. For enterprise users, this broadens GPT-5.4’s capability set beyond generic conversations to bespoke task execution. Skills differ substantially from GPTs and projects—they offer tailored, reusable workflows that can be invoked for repetitive or highly specialized business processes.
Browse Comp Score: Reliable, Real-Time Web Analysis for Decision Makers
OpenAI’s new model achieves an 82% browse comp score (89% for pro tier), which evaluates the model’s ability to find obscure, hard-to-verify information through persistent, multi-step online research. For decision makers, this marks an observable jump from the previous GPT-5.2’s 65%.
Practically, this feature addresses longstanding challenges around knowledge cutoffs and lagging datasets. Organizations relying on AI for competitive intelligence, rapid market updates, or in-the-moment analysis now gain access to a model that fetches and verifies information with higher fidelity using up-to-date web sources. This capability is directly relevant for teams that need current data for marketing, forecasting, or business development, eliminating the risk of relying on outdated or incomplete information.
Precise Instruction Following: Bridging the Gap Between Pro and Standard AI Tiers
Instruction following in GPT-5.4’s higher-thinking modes has improved significantly. Whereas previously, only the expensive pro plans ($200/month) offered robust multi-step reasoning, tests show the $20/month Plus tier has closed much of the gap. Models now interpret detailed prompts, execute complex workflows, analyze large datasets, and produce nuanced outputs, even with the more accessible plans.
The hands-on assessment demonstrates that even standard and extended “thinking” modes (for Plus plans) perform at a level previously seen only in premium tiers. Businesses seeking high-quality, accurate, and customized results can now attain these without incurring enterprise-level costs—an important consideration for scaling AI use across departments.
Transparency and Natural Interaction: Improving Trust and Usability for Business Stakeholders
GPT-5.4 distinguishes itself from prior models by combining high transparency with conversational naturalness. Unlike competitors that fail to provide visible chains of thought or actionable step-by-step reasoning, GPT-5.4 offers summarized reasoning and clear sourcing, critical for regulatory, compliance, and auditing needs.
The transcript highlights a detailed use case: analyzing 20,000 business podcast stats points, categorizing shows, extracting trends, and generating actionable dashboards. GPT-5.4 not only correctly interprets raw data but also contextualizes business directives for audience growth, retention, and market relevance, offering recommendations based on accurate reasoning rather than surface-level analysis.
For business leaders, this means increased confidence in the AI’s outputs, the ability to audit suggestions, and streamlined communication between human and machine.
Key Takeaways for AI-Driven Business Strategy
Every highlighted feature addresses specific needs—adaptive workflows, enterprise skill augmentation, reliable web research, affordable high-tier reasoning, and transparent, actionable dialogue. GPT-5.4 stands out as the first model, per the hands-on review, to meet all three core requirements for daily, uncompromising business use: usability, intelligence, and precise instruction adherence.
Businesses looking to integrate AI for daily operations can now achieve measurable gains in efficiency, data accuracy, and decision support, without navigating multiple platforms or sacrificing depth for speed. GPT-5.4’s hands-on features set a new practical benchmark for everyday AI-driven business processes.
Topics Covered in This Episode:
- GPT-5.4 Model Usability Trifecta Overview
- Interrupting Thinking Mode: Hands-On Breakdown
- GPT-5.4 Paid Tier Access Details
- Skills Integration with GPT-5.4 Explained
- GPT-5.4 vs Claude: Skills Advantage
- GPT-5.4 Browse Comp Benchmark Performance
- Upgraded Instruction Following & Task Accuracy
- Natural Language, Transparency, and Chat Experience
- Real-World Podcast Analytics Testing with GPT-5.4
Episode Transcript
Jordan Wilson [00:00:15]:
OpenAI's newly released GPT five four model crossed a new threshold for me. And I spend thousands of hours each year evaluating and stress testing models, so that says a lot. But GPT five four is the first AI model that I think hits the full usability trifecta to be a true daily driver model without any compromise. It's natural enough to chat with, number one. Number two, it's legit off the charts in terms of general intelligence and transparency. And three, it follows instructions to a tee no matter how daunting or challenging or long the task is that you throw at it. And I think there's been a lot of talk lately about how the models are becoming less and less important as the harness and tool use become more and more important and where the moat is actually at. And to a certain point, I do agree with that, but with releases in 2026, the line between model and harness and tool use start to blur.
Jordan Wilson [00:01:22]:
It's because the model updates used to bring only updates in the underlying intelligence engine. Not anymore. Now model updates like OpenAI's impressive upgrade to g p t five four also bring with it major changes to the harness and the tools, which completely changes what an AI model can actually accomplish. And with this round of updates to GPT five four, I think OpenAI knocked it out of the park. So today, we're putting AI to work on Wednesdays as we go under the hood a bit with GPT five four as we go hands on, and I will also break down the five reasons why I'm confident GPT five four will be the best model you've ever used yet. Alright. I'm looking forward to this one. So on today's show, if you stick with me, here's what we're gonna go over.
Jordan Wilson [00:02:14]:
So you'll hear what's new and noteworthy in OpenAI's newest model, GPT five four. You'll learn why you may be able to get the benefits of the $200 a month pro plan without even really paying for it. You'll know the reasons why the model is best to be your daily driver right now at least, and you'll leave with the five reasons why it'll be the best model you've ever used. Alright. Let's get into it, shall we? If you're new here, welcome. My name is Jordan. This is everyday AI. And, well, this thing it's for you.
Jordan Wilson [00:02:47]:
Unedited, unscripted, just bringing you the realest information in intelligence and artificial intelligence, and hopefully giving you the tools to grow your company and your career. So if you're along on the journey, awesome. It starts here. But to take it to the next level, make sure you go to our website and go sign up for our daily newsletter. We're also gonna be recapping today's show. So if you are brand new here, on Wednesdays, we do, putting AI to work ad on Wednesday. So it's usually a more hands on, impractical use case of, you know, one of the usually, one of the big four, right between, Microsoft, Google, OpenAI, and Anthropic. We really like to go hands on on Wednesdays.
Jordan Wilson [00:03:31]:
But if you wanna know more of the details and the benchmarks of OpenAI's new model, we did cover that right after its release on Friday. So you can go click the back button a couple of times if you're listening on the podcast, to episode seven twenty eight, where we go over more of the release in the seven trends that I think you need to know about OpenAI's new model, which are so much more than just benchmarks. Alright. Let's get into this, and I'm not gonna make you wait any longer for the five reasons. So number one, interrupting thinking mode. Alright? That's one of the reasons why I think this is gonna be the best model that you'll ever use. And you might be wondering, like, well, number one, what is that? And, well, why does it matter? Okay. So the big four, well, in this case, the big three model makers and, you know, OpenAI uses Anthropics models and they use OpenAI's models.
Jordan Wilson [00:04:30]:
When you use a thinking model, it can take a terribly long time. And I think sometimes people don't use thinking models for that very reason. They're like, well, I need an answer right away. Or, Hey, if I realized that I forgot to say something or forgot to do something, I don't want to have to wait five, ten, fifty minutes for a model to finish its thought process and to give me the answer. So this is a new feature for g b t five four for the, the masses because this was actually available on the previous pro tier, but now it's available to anyone that's on a paid tier. And I should probably start with that, although we did go over that in Friday's episode, but I should let you know. Right? To use the thinking model, the new GPT five four thinking or GPT five four pro, you do have to be on a paid plan, whether that's the $20, a month plus plan, the $200 a month, pro plan, the business plans, enterprise, EDU, etcetera. Right? So if you're wondering, like, where is 54? Well, that's where it is.
Jordan Wilson [00:05:38]:
It's on the paid plan. And, actually, another small thing here that I'm just realizing kind of now, even though I've been talking about GPT five four quite a bit already. Maybe it's good that they didn't come out with the free or the instant version of five four. And maybe that was intentional in hey. OpenAI folks, I know there's a few of you listening. If this was not intentional, you go ahead and take this and say it was. I think one of the biggest downfalls of Chad GPD is people are using the bad model. There's always a bad model in anyone that you use, whether you're using, Claude, Microsoft, Gemini, chat GPT.
Jordan Wilson [00:06:18]:
Right? So previously, seven days ago, right before there was five, five four or even five three instant, I'm gonna get to that here in a second. We were living in a GPT five two world. So when everything was GPT five two, well, people just thought, well, I'm using the best model. Well, no. Because if you were using the instant model, which is the chat model, it doesn't think, and it's not really good. Right? So essentially last week, OpenAI came out with GPT five three instant, which was kind of confusing because then the next day they came out with GPT five four thinking and five four pro. So with the all the model confusingness that's going on, maybe it's actually a good thing, because someone knows, well, if you want the best, you should be using GPT five four. And at least right now, there's not a bad version of it.
Jordan Wilson [00:07:12]:
Whereas, generally, there's always been a, quote, unquote, bad version of the best model. And the overwhelming majority, I am talking hundreds of millions of users worldwide don't know the difference. So maybe this is actually a great thing. Anyways, getting back to the number one reason, interrupting thinking mode. And you'll see in some of these examples here, I already did them, but we're gonna go live under the hood because I'm not gonna make you wait. Some of the thinking models, you know, took, like, thirty five plus minutes. Right? But if you see something going wrong, you can course correct it. Unfortunately, you can't upload files or use different tools during that course correction.
Jordan Wilson [00:07:50]:
But right now, Chad GPT is the only major model maker to offer this. Right? So if you are using as an example, quad four, six, Opus, if you're using Gemini three, one pro and you're using the thinking, which you should for most tasks and you see something's going wrong and you're like, oh, crap, forgot to do something. You could be five, ten minutes in. Right. And you, you either have to scrap it or you have to accept, subpar answer. And that kind of stinks. Right? I've been using the, thinking interrupting because I've been on the, the the pro plan for a while on Chat GPT actually since it came out. So I'm kind of used to this on the pro level, but it's really nice to get this on the thinking level because this is where the masses are.
Jordan Wilson [00:08:36]:
And this is where you should be is using the thinking models. All right. Reason number two, it can access skills now. Yeah. You didn't know this probably because open AI didn't even announce this. I don't even think there was a tweet out. I think they just updated a blog post. So skills was actually a big advantage for Claude.
Jordan Wilson [00:08:57]:
Anthropic kind of, created and popularized skills, and now they're really used across the industry. But, up until, well, couple hours ago, skills was only available in codex, which is chat GPT's coding tool. Although I think it's way, way better. Sorry. FYI. I I think it's way better than Claude Code and Claude Cowork combined even though I use all three. Codecs is a nice in between. It's like a blending of the two.
Jordan Wilson [00:09:28]:
Anyways, you could use skills inside of Codecs, which is ChatGPT's desktop app or their command line interface tool, more of their coding tool. But they kind of slid in these skills under the radar, but unfortunately, they're only available on business or enterprise plans right now. But to be able to pair up skills, with GBT five four is huge because, again, up until a few hours ago, I still would say that was one of the reasons why, I was still using Claude a lot more for some of my quote unquote daily driving. I think skills in that framework, which will probably dive into a little deeper on a future start here series. So if you haven't been listening to our start here series, it's how to go from, you know, zero to 10 or at least zero to five, you know, in understanding AI. Skills are great. They're a little different than GPTs. They're different than projects.
Jordan Wilson [00:10:25]:
Right? There's a a great and flexible utility to skills. But, now that you can pair skills with GPT five four, that's pretty big. Alright. Number three, a benchmark that users will actually feel is gonna make the difference with g b t five four, and that's browse comp. Okay? So browse comp, if you've never heard of it, it evaluates an AI agent's ability to find obscure, hard to verify information through persistent multi step web browsing. AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead. But you don't have ten hours a day to understand it all.
Jordan Wilson [00:11:14]:
That's what I do for you. But after 700 plus episodes of Everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode six ninety one. Number two, tap the link in your show notes at any time for the start here series, or you can just go to starthereseries.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same. The start here series will slow down the pace of AI so you can get ahead.
Jordan Wilson [00:12:12]:
Why is this incredibly important in a large language model? Well, for a lot of reasons, but one that I think sometimes gets overlooked and, you know, I talk about a lot on the show, but with 700 plus episodes, it's worth mentioning again. For the most part, when you're using today's large language models, even the latest ones, they're working with a very old knowledge cutoff. Right? Usually, you know, it might be, six or so months. Alright? But that's just the absolute best case scenario because many frontier labs are using offline datasets, and the data in those datasets might be two years old. Right? So the ability to browse the web accurately and to follow your instructions while browsing the web is not just a nice to have. It is an absolute necessity. Because if you are using the outputs of a large language model for business purposes, which is, like all of us, everything changes. Right? Unless you're writing a a history paper or you're using this, right, to just, I don't know, do something about ancient yeah.
Jordan Wilson [00:13:19]:
I I don't know. Ancient history. Right? But everything else changes. Even if if you're using this to, you you know, market your business and maybe your industry is a slow moving industry, well, marketing is changing daily. Right. So browse comp is huge, and, OpenAI is the now world leader in browse comp, and it's going to be a noticeable jump. So although they're only a few percentage points now ahead of Google and Anthropic, I think those few percentage points can actually be felt. It is actually a huge jump from where they were with the last models, which is g b t five two.
Jordan Wilson [00:13:55]:
So if we look at the normal just thinking versions, g p t five two was a 65% on browse comp and g p t five four is an 82%. So a huge jump. Right. And then you have GPT five four pro at 89. Right. And that is even though it's about, you know, four or five percentage points above, anthropic in Google's offerings there. You can tell. And I will have a very small, secluded example of that as we go live here.
Jordan Wilson [00:14:28]:
Also worth noting on browse comp, yeah, anthropic essentially fessed up, which good on them. Right? I do think anthropic is really good at when they find issues, they say it. Well, they said even their score on browse comp, they figured out that Claude was cheating. So they said on their website, they said evaluating Opus four six on browse comp, we found cases where the model recognized the test, then found and decrypted answers to it, raising questions about eval integrity and web enabled environment. So yeah. Essentially, while doing their, evaluation for this, they found that, oh, Opus realized it was being evaluated and kinda decided to cheat. Anyways, that's something that is actually gonna be felt, browse comp, because how important it is. And most people don't understand the amount of agentic browsing that is needed for everyday intelligence and for everyday business use.
Jordan Wilson [00:15:26]:
You need it constantly, and you need it to be really good, and you need it to number four, our number four point, oh, look at that three and four go together. Instruction following. The instruction following right now on the higher thinking models is other worldly. Okay? And let me talk about this, and this is also the little secret there that I teased in the beginning, how you might be able to get that $200 a month value out of a $20 a month plan. So on the base chat GPT plus plan, you do not get GPT five four pro, which is a bummer. Alright. Hey. Little secret.
Jordan Wilson [00:16:04]:
If you just get on the business plan, which is $30 a month, minimum two seats, you get a couple pro queries and it's worth it. Even if you're the only one using it and you have two seats, FYI. So what I found through my testing, which I was kind of surprised by, aside from the instruction following, which is outstanding, and that will probably make more sense showing you live, but even just with the thinking models, that's why I say instruction following on higher thinking models is other worldly. Right? So if you are on the Cheggitt plus plan, you get two kind of levels of thinking. If you're on the pro plan, there's four levels of thinking on the thinking models. Right? The chat GBT plus the, higher level of thinking, I found it to be I wouldn't say the results were comparable, but it put in the same amount of reasoning effort as the pro plan, as the, the, GBD five four pro on a lot of my internal testing. And it actually, in many cases took longer to think and took more steps. Again, the output wasn't always better or even the same, but it was comparable.
Jordan Wilson [00:17:24]:
And that's really important to point out. Right? I think it's pretty well known across the industry. I don't care if you're a fanboy of OpenAI, Microsoft, Anthropic, Google, it doesn't matter. I think most people know and have understood for a long time. If you need something right and getting it accurate and correct is of utmost importance, you always well, up until this past week, you would go to GPT five two pro. Now you go to GPT five four pro. Alright? But the problem is it's extremely slow. Right? And, well, it is expensive if you're using it well, even if you're using it in chat g p t, the $200 a month plan, that's kind of expensive.
Jordan Wilson [00:18:11]:
If you're using it in the API, it's, ungodly expensive. But the higher level of thinking now in the base chat GPT plus plan for $20 a month. Again, they're not on the same level, but it closed the gap. Right. There were so many things previously that I would always just use pro for. And I would never think of using GBT $5.02 thinking for so many tasks. Now I don't think twice about it. Even if I just have those two tiers of thinking, the higher tier of thinking is so much better than it was before.
Jordan Wilson [00:18:49]:
It's closed the gap. I think it's if nothing else, it's gonna, maybe just allow people to use, thinking way more into use pro way less. Right? Maybe it'll end up saving, you know, open AI some money in the long run. And then number five, it is the most natural, generally intelligent chat bot that I've ever used. And I think for the first time, maybe ever, I felt a model didn't have these out of the box glaring weaknesses, in either intelligence and transparency and that intelligence or chatting. So here's what I mean. I think Gemini three one pro is on the same level as GBT five four, the thinking and the pro level. The problem is Gemini three one pro, you don't have the transparency of intelligence.
Jordan Wilson [00:19:45]:
Right? Which is important because let's be honest, humans. When was the last time that you used any of these models for an entire day? Right? And you look at the answer and you're like, yeah. I I knew that. I feel confident in this. No. Right. It's so important to be able to look at the chain of thought, the summarized chain of thought, and to be able to transparently see where these models are getting the information. So unfortunately, right now, Gemini, does not provide all of that information in the same way that OpenAI and Anthropic, does.
Jordan Wilson [00:20:25]:
Right? I know they're changing it. I've chatted with them. They've said as much that they're eventually gonna be bringing a little bit more transparency to the chain of thought. I know there's problems with, competition, distillation, all those things. So things I don't understand that they have to protect. But in terms of business use, right, it's one of the main reasons why I love and have loved for so long, g the g p d five, thinkings and even going back to o three and o one. The chain of thought not just shows, I think, that OpenAI's models are better. They provide more transparency, and you can understand them more, and they're just way better at instruction following.
Jordan Wilson [00:21:06]:
So that's on the one side. It's just generally intelligent in a transparent way. And then on the other side, you can actually talk to it. Let me be honest. I don't really care about having a pleasant conversation with a chatbot. And I do know that OpenAI, you know, has some newer settings that you can default its voice to just be concise and, you know, even the just changing those, the default response does it for me. But I know for a lot of people, they don't go in and do that. And a lot of I think historically, OpenAI's models have been overly sycophantic.
Jordan Wilson [00:21:40]:
They've been verbose and, you know, even OpenAI admitted that they've been cringe. Right? So for the first time, I think maybe ever, you have a model that is transparently intelligent off the charts, number one. But number two, you can actually talk to it. Right? Not in a, you you know, you're my hype man kind of way. Right? But, man, I I mean, honestly, I spend way too much time, quote, unquote, chatting with large language models. I'm really just directing them, agentically. Right? But I get so tired of the the way they respond, and I'm like, I can't even read this. It's, you know, so cringe, so sycophantic, so verbose, whatever.
Jordan Wilson [00:22:23]:
Right? And I think for the first time, we're not getting that anymore. It hits the sweet spot. Alright. So before we go live, let's jump in and, livestream audience. Thanks for sticking around. This is gonna be a little shorter on the live end just because, you know, we can't really, like, watch a a a twenty, thirty minute prompt. That'll be super boring, but we are gonna go under the hood. But a couple of things to keep in mind, And I'm gonna go ahead and I'm gonna call out the people that are calling me out.
Jordan Wilson [00:22:54]:
Alright? I get accused a lot of of it's actually strange because I get accused of of pumping, you know, OpenAI, and but then I get accused of pumping anthropic, but then I get accused of, you know, pumping Google. Right? I get accused of being a fanboy for everyone, but also against everyone. Doesn't make sense. But overwhelmingly, I think it's fair to say. I am more preferable to OpenAI and Google than I am to anthropic. So let me just speak to those people because I get a couple messages every single week. People, you know, accusing me, oh, Jordan, you don't know what you're doing. You don't know what you're talking about.
Jordan Wilson [00:23:29]:
You're clearly an idiot. Anthropic's great. Okay? Let me just let me tell you this. When I'm doing these demos, when I'm doing these shows, right, not just our putting AI to work at Wednesday shows, but just the 700 plus episodes. I am speaking to the general business lead, right? The C suite exec. Sometimes that person is a technical person. Sometimes they're not right. My approach has always been about using AI to automate tough general tasks.
Jordan Wilson [00:24:01]:
Right. And Hey, the reality. And I went over this on the show Friday. So go listen to that. Yes. Anthropic has generally always held a sizable advantage, you know, software engineering, agent orchestration, computer use. Well, not anymore. Right? OpenAI with this model, they actually took their lunch money on that.
Jordan Wilson [00:24:24]:
So yes, up until last week, I do think Anthropic had some huge advantages. And y'all, I kid you not. I'm using billions with a b, billions of tokens between Claude Code, Claude CoWork, codex, anti gravity. Like I have max subscriptions to everything and I hit my rates constantly. Right? Billions of tokens. So I know what I'm doing. I know what I'm talking about. Alright? Just I'm putting that out there.
Jordan Wilson [00:24:56]:
Is Claude great? Absolutely. And it's great for certain tasks. Is Gemini great? Absolutely. It's great for certain tasks. But I do think with this one, this is the first one that I feel confident. You can go back and and and listen. I've never in 700 plus episodes, I've never said, hey. I think this is can be a go to daily driver model because I think for the most part, it's better to be jumping around in multiple models.
Jordan Wilson [00:25:21]:
But I do feel maybe for many people, five, four thinking will get you there. So when I go through these demos and when I kind of show you under the hood here, I want you to think of a multi step tough problem. What is that multi step tough problem that you have? All right. And then I encourage you do that same thing. It's gotta be tough. Do that exact same thing in GPT five, four thinking. If you have access to GPT five, four high, do it there, Do it in, Opus four six with extra reasoning, and then do it in, Gemini three one pro. Do it yourself.
Jordan Wilson [00:26:06]:
It's gotta tackle the entire gauntlet. Right? Data analysis, web research, reasoning tool, use instruction, following common sense. That's what I'm gonna show you now. And I'm telling you people, everyone that that gets on my back about, oh, Jordan, you know, you're you're you're too hard on Claude. You clearly don't use it. Yes, I do. Right? I've been a max subscriber for the longest time. I bounced between the, you you know, the $20 and the 100 and the $200, but I've been a subscriber to Claude since it came out, and I've been on the up to the highest tier.
Jordan Wilson [00:26:41]:
So FYI. Alright. Let's look live. So here's what we're gonna do, y'all. We have a lot of stats here. Alright? These are my podcast stats. Yes. We're gonna do another thing looking into my podcast.
Jordan Wilson [00:26:53]:
Right? I'm not gonna jump in and pretend to show you financial analysis on something. That's not my background. Right? That's not what I'm using it for. I'm doing my use case. Think of yours when I walk you through mine. Alright. But I've actually had, this was put together. This version was put together by codex, and then it was enriched by, Claude Code.
Jordan Wilson [00:27:13]:
But, essentially, I have a mountain of data from my podcast stats. So I've done this before in the past, but this version's a little better. So there are certain things that I can get out of my provider, which is called buzzsprout, but not what I really need to make educated decisions. So my problem, I have 700 plus episodes and I can't easily export all of the data I need to make good decisions on what type of episodes I should be doing more of and which types I should be doing less of. So I did have, using both quad code and codex. I put together, this better look at my stats. So there are more than 20,000, data points in here. So, yeah, I had to grab a lot of this with an agent with some APIs.
Jordan Wilson [00:28:05]:
I had to do a lot on it. But, essentially, I'm able to get the episode titles, normal stuff that I would normally get, episode length, but I'm also able to get if it's a guest or solo, the number of plays, consumption hours, retention across quartiles, which is super important, completion percentage, consumption hours, discovery, people reached. Right? All of these different metrics that are not usually available when I just go click export on my stats. Problem is this is a ton. This is it's a ton of information, and all of these shows are not also categorized. Alright. So that's another thing I'm gonna be telling the models to do. Alright.
Jordan Wilson [00:28:47]:
So let's jump in, and I'm gonna go ahead and read, let's go here. I'm gonna go ahead and read the prompt that I sent, and then I'm gonna kinda go through the responses here. So, this is using GPT five four, on heavy thinking. So again, if you're on the pro plan, you have light standard, extended and heavy. If you are on the $20 a month plan, you have standard and extended. So I said, and I uploaded the file. I said, this is a comprehensive list of stats from my Spotify analytics for my podcast, Everyday AI. Please take your time analyzing all the data.
Jordan Wilson [00:29:25]:
Keep in mind everything you know about me and Everyday AI in personalizing your replies, including strengths, growth areas, known bottlenecks, and constraints. Your responses to most of the below should take into account growing the podcast, saving time without sacrificing quality, and improving new audience discovery, retention, stickiness, consumption downloads, etcetera. Alright. Use every tool available at your disposal. Right? I'm trying not to read all of this because it's a lot. And then I'm saying avoid, let's just say, I do wanna make sure I include one of these parts. Okay. When and if needed, you can access complete transcripts on my website at youreverydayai.com.
Jordan Wilson [00:30:02]:
So then I said after carefully and meticulously analyzing the data, please reply back with and then I have, five different categories, and in each of these categories, I do have four to five different things that I'm asking. So for category one, it's essentially obvious trends, and I'm, you know, asking for, to give it to me in three different ways, 20 obvious trends. Category two, under the radar trends. I'm asking for 20 under the rating under the radar, but meaningful, statistical trends, three different ways. Then I'm doing comparisons. Alright. So about, looks like six no. Five different ways.
Jordan Wilson [00:30:43]:
So, you know, as an example, the 10 types or categories of shows that are the most popular and why. So just so you know, you weren't able to fully see and read my spreadsheet especially especially if you are on the, the podcast only. Right? You can always watch the video version. Yes. There is a video version. You go watch it on our website at youreverydayai.com. So my spreadsheet didn't have categories. Right? I I have another version with categories.
Jordan Wilson [00:31:12]:
I honestly lost it. It's like buried, I don't know, in co work, or codex somewhere. So I am also having it categorize it. So it's not just reading. Right? It's not just reading the data. It's having to crunch the numbers, and it's also having to think. Right? Hey. According to this show, what is it? What category is this? What does this mean? Alright.
Jordan Wilson [00:31:36]:
And then the fourth category is March 2026 planning. So I'm essentially saying, hey. Based on all this data, what works and what doesn't, go research trends, see what I haven't covered. That's important. I'm asking, what I haven't covered, but I should. Alright? And then last but not least, I say you're in charge. Right? If I wanted to double my audience this year, what are the different things I should be doing? And then, you know, asking that in three different ways. Alright.
Jordan Wilson [00:32:05]:
So that's essentially what I asked, and then I did do this, both in thinking heavy thinking mode. I did this in GPT five four pro, and then last but not least, just for fun, I also did it in, Opus just to have a a baseline. Right? And just because I don't know. I think I need a a demonstration. I can just send to people that all the time are telling me, I don't know what I'm talking about because anthropic is so much better. No. I know what I'm talking about, y'all. Will it be next week? Maybe.
Jordan Wilson [00:32:38]:
Today, it's not. And it hasn't been for a very long time. For general knowledge, work, hard tasks, and tropic has never been the top model period. Right? Look at the benchmarks all you want. It hasn't been. So here's where we're gonna start to dig in a little bit and why I'm gonna start kind of referencing back some of those five big points that I talked about earlier. So oh, at the very end, I did say, to Chad g the the only difference in these prompts, at the very end, I told chat g b t, I said use canvas mode to put together a sleek interactive and useful dashboard that includes all of this information. And then for, improv, Claude, I said the same thing, but I said using artifacts.
Jordan Wilson [00:33:29]:
Right? Because they don't have a canvas mode. It's called artifacts. So that's the only difference. Otherwise, everything was the exact same. Okay. So and then also the, the pro GBT five four pro, you cannot use canvas. So, there was no dashboard. So both models completed the task, the quality and the nuance completely different.
Jordan Wilson [00:33:56]:
Right? And I will start to show a couple of the things. So let's scroll down here. Let's scroll down here a little bit. Alright. So big big big differentiator right here. Right? It fought for thirty nine minutes and forty seven seconds. Alright. There's no timer on the, Claude anthropic, which they used to have that.
Jordan Wilson [00:34:19]:
I don't know why they don't anymore. But it was about four minutes. Alright? So you can look at that in a good way or a bad way. Well, I'll tell you, spoiler alert, Claude's version was, I won't say trash, but compared to GPT five four thinking, Claude's version was trash. Right? Exact same prompt, exact same data, memory, chat history, all the same. Right? I upload those in markdown files. They have the same thing. It was bad.
Jordan Wilson [00:34:53]:
It was really bad. So you can't just look at time. You have to look at output, and I'm gonna show you a couple of things. But first, we have to be able to see here, what g b t five four thinking did. And again, this is not one of those, I think, on five two, I should have ran it. Maybe I'll I'll rerun this and put it in the newsletter on 05/02. I'm guessing it probably would have only taken twenty or so minutes, and it wouldn't have picked up on nearly half of the nuance that it did in this case. Right.
Jordan Wilson [00:35:23]:
And one of the biggest things out of the back that I didn't even tell it to, it says Spotify says discovery data is a last thirty days view and can take up to forty eight hours to refresh Right? Before and this is in the very first paragraph. Right? Because if you look at the chain of thought, you'll see one of the first external websites it goes to. It might have pulled it up in the API. I'd I'll have to look later. But it instantly looked at Spotify because I told it these are my Spotify stats. Guess what Claude did not do? Well, it didn't look at that. And you might be saying, okay. Why does that matter? Well, because a lot of what Claude suggested in this case, and I'm not trying to turn this into a GBT 5, 4 versus Opus $4.06, but I know that's what a lot of you all are gonna be thinking.
Jordan Wilson [00:36:18]:
The the bulk majority of what Claude recommended was just dumb because it didn't do the basic work. Right? I say basic, but it's actually nuanced and super smart, that, Chad GPT went out and found that the Spotify discovery data is only the last thirty days because one of those columns in there is discovery. Right? So it's how many people are discovering each episode. So Claude had all these straight up not useful and off the wall and incorrect, kind of insights throughout this entire document because it didn't understand that that was only the last thirty days. Right? It's like your discovery, you know, has gone up 268 x, you know, this month. It's like, no. It hasn't. It's just because the discovery is only the last thirty days.
Jordan Wilson [00:37:14]:
Right? This is something like an intern that wasn't using their brain would come in and be like, oh, look. Look at these stats. Wow. The last thirty days have been way better in certain categories. It's like, no, dummy. You didn't think. Right? And and this is why, again, I think the difference here on the thinking model is huge because I don't know if the g p five two thinking would have, you know, picked up on some of those nuances early on, and picking up on that early on is pivotal. Right? So we'll go through here.
Jordan Wilson [00:37:43]:
I'm not gonna be able to go through all of them, but I will just point out the instruction following is fantastic. Right? So in in the obvious trends, it broke it down. Right? The 20 obvious trends across the three different categories. Alright. So there's our top 20 trends. You know, maybe I'll read, like, one or two. We'll go I don't know. Maybe lower.
Jordan Wilson [00:38:06]:
Oh, here. This is good. Okay. The first six weeks of 2026 are the healthiest early year cohort in the sheet. That's great. Maybe it's because of the start here series. Number 17, a higher solo share is part of the improvement. Yeah.
Jordan Wilson [00:38:21]:
I noticed that. Our guest shows weren't doing as good. Right? In general, apparently, people didn't like guest shows. So I've been doing fewer guest shows because, hey, codex went through and broke it all down for me a couple months ago. I was like, yep. You should not be doing as many guest shows. So I said, okay, AI. Okay.
Jordan Wilson [00:38:40]:
Then the 20 under the radar trends, Great. Let me just go ahead and maybe read one of these. Okay. This one's interesting. It says Google is not just a spike topic for you. It is sticky. Right? So it says that the Google or Gemini wins in both median plays and long tail behavior. Right? Because it was able to properly understand the discovery metric.
Jordan Wilson [00:39:04]:
It says, current context. Google is still shipping meaningful work oriented AI updates into March 2026. It said build a recognizable weekly or biweekly Google franchise. So it's pretty good. I don't do as many Google shows as OpenAI, and recently, I've done more clawed and anthropic shows versus Google. So it pointed that out. It said Google is shipping at a high rate, and you're not covering a high enough percentage of what's of what Google is shipping, and it's very sticky for you in terms of audience retention. So cool.
Jordan Wilson [00:39:36]:
Alright. I'm not gonna read all these, although they're super fun and important for me. I will just go through and say, it completed everything. Right? In number three, the comparisons, every single one. Alright? I go down in the, you're the boss or no. March 2026 planning. Perfect. Right? What's actually funny is as I was planning this, it said the first show I should do is g p d five four at work five tasks.
Jordan Wilson [00:40:03]:
It actually does better now, which actually might be a better title than this show. I didn't see it until I was already making this show. But it properly and I'm gonna point it out here because I do wanna look at Claude's. It properly did this. Right? Because I haven't number one, these are all relevant and useful shows according to what it earlier identified were high performing shows, but these are also shows, well, I haven't done. So number one, they're not repeat shows, so it follow directions. Number two, it taps on what worked. And number three, well, they're highly relevant.
Jordan Wilson [00:40:41]:
Alright. And then the, you're in charge, went down here. Hey. Number one, make solo practical explainers your default weekday format. Alright. I'm doing that. They're just more time consuming. And I will go and show at the very top here.
Jordan Wilson [00:40:55]:
It did also properly complete the, the dashboard. So it's not the prettiest dashboard. It's actually pretty plain and ugly, but it's helpful. Right? And there's some cool interactive graphs in here. You know, I can click through the, the overview, the obvious trends, the under the radar, the comparisons, the March 2026 plan, and how to double the audience. So in terms of output, instruction following, accuracy, GPT five four, thinking, the thinking mode. This is this wouldn't have been possible before. Alright? And just a quick gripe in comparison because I know people are gonna be wondering, right? Claude's was not good.
Jordan Wilson [00:41:41]:
Granted, the dashboard it made way better. Looks better. Right? One of the things that g p d five four stinks at, front end design, not any good. Claude, amazing at front end design, but not good at, well, things that require, number one, factual accuracy. Number two, assuming things, making assumptions, not good, and just not completing the task. Alright? So let me show you one or two just quick examples. Let's see. Okay.
Jordan Wilson [00:42:16]:
Here we go. This is just what I had what I had up. In the March 2026 planning section, it's okay. It's saying, anthropic versus OpenAI Pentagon drama. Oh, guess what? I already did that show. Guess what? I asked for 10 examples. Guess how many it gave me? Three. Right? It did that repeatedly.
Jordan Wilson [00:42:40]:
When I would ask for 10 things, it gave me either five or it gave me three. Right? It didn't always give me 10. It's not good at instruction following. What is y'all? If what's the difference between an intern that's not very good And someone in your company that's gone from junior analyst, junior researcher to senior, their ability to follow instructions and y'all stop chirping at me, go run these own, like go, go run your own multi step, extremely hard multi tool use examples with real data that require research, that require multiple tool calls and have a multifaceted required output. You'll see for yourself. It's it's it's not it's not a comparison. Right? So for for for everyone saying like, oh, Jordan, you don't know. No.
Jordan Wilson [00:43:37]:
I know what I'm talking about here, and I want you to know what you're talking about too. So don't just take my word for it. Right? Go in, try all these things out yourself. Alright? So I'm not gonna keep comparing. I think that was a pretty good under the hood look. Oh, and FYI, another thing why, Claude really failed here in talking about, you know, some of the advantages of G b d five four. Well, it didn't even go to the website that I told it to. Right? I told it go to youreverydayai.com.
Jordan Wilson [00:44:06]:
It didn't. It suggested things like you should post things on YouTube. Right? And then in the, in g p t five four, it found, you know, our YouTube channel, which I completely ignore. And it's like, hey. You're already doing things on YouTube, but you should be doing more shorts. Right? So, yeah, Claude just assumes things. It number one, it rushes. Number two, it doesn't check.
Jordan Wilson [00:44:29]:
And, yes, I was on Opus 4.6 extended, right, the best extended model you can do. And I wasn't even comparing this to GPT five four Pro. It just it falls flat, but I think it's not so much, Opus four six falling flat. It is that now, I think, to reiterate my point, I think for the first time, we have that trifecta. Right? We have a daily driver model that it's natural enough to chat with. It is off the charts in terms of transparency and general intelligence, and it will follow instructions to a tee. Because when you are looking for a daily driver, large language models, those are non negotiables. And I think maybe for the first time we have them all in a single package with GBT $5.04.
Jordan Wilson [00:45:24]:
Alright. So I hope this was helpful going over a little bit hands on under the hood, maybe a little bit more than normal, maybe a little bit more technical. Alright. But now you know, well, why I think in my thousands of hours of experience, the five reasons why it'll be the best model you've ever used in GBT five four, at least today, because who knows? Maybe tomorrow, this could all change. But, hey, you know what that means? Right now, you're at an advantage. So number one, go test it for yourself. Number two, refine and reiterate. And number three, get to work.
Jordan Wilson [00:45:58]:
Get ahead of your competitors. Alright? And the other way you do that is you go to our website, youreverydayai.com. So thanks for tuning in. We're gonna be recapping today's show in our newsletter. If this was helpful, tell someone about it. Thanks for tuning in. Hope to see you back tomorrow in everyday for more everyday AI. Thanks, y'all.
