Ep 472: OpenAI’s new GPT-4.5: What’s new and who can benefit the most

Resources:

Join the discussion: Ask Jordan about GPT-4.5


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Unveiling OpenAI's GPT-4.5: What Business Leaders Need to Know

In the evolving landscape of artificial intelligence, OpenAI has unveiled its latest model, GPT-4.5, marking a significant departure from its predecessors. As businesses strive to leverage AI for strategic advantage, understanding the capabilities and applications of this new model is crucial. Here's an analytical dive into what GPT-4.5 offers and which sectors stand to benefit the most.


A Shift Toward Relatability and Reliability

Unlike previous models, GPT-4.5 emphasizes creating a more relatable and emotionally intelligent user experience. It’s engineered not to break benchmark records but to feel more human-like in interactions. This model is designed for those who value nuanced conversations, making it ideal for sectors reliant on customer interaction.


The Technical Edge of GPT-4.5

GPT-4.5 is not only about improved emotional intelligence. It was trained with ten times more computing power than its predecessors, allowing for more sophisticated pattern recognition and intuitive problem-solving. However, its notable strength lies in decreasing hallucination rates, which can enhance the accuracy and reliability of the information it provides—key factors for data-driven industries.


Cost and Accessibility: A Strategic Consideration

While GPT-4.5 is currently only accessible to high-tier users or through expensive API rates, its broader rollout is anticipated. Companies should weigh the substantial costs against strategic benefits, particularly when precision and customer relations are priorities.


Who Stands to Gain?

Industries such as customer support, marketing, and those involving extensive human interaction are poised to benefit from GPT-4.5. The model's capability to understand emotional cues can revolutionize how businesses engage with customers, potentially improving satisfaction and loyalty.


A Foundational Move Towards Future AI Models

GPT-4.5 signifies a potential pivot in AI development strategies, emphasizing soft skills integration. This might set the groundwork for future hybrid models, combining reasoning and emotional understanding, thus broadening AI’s applicability across industries.


Conclusion

For business leaders, the potential of GPT-4.5 lies not just in its advanced technical features but in the new dimension of human-like interaction it introduces. As AI continues to intertwine with business processes, understanding and integrating models like GPT-4.5 will be essential for maintaining competitive advantage in customer-centric industries. As the rollout continues, staying informed about updates and applications will be crucial for strategic decision-making.


Jordan Wilson [00:00:17]:
The AI model releases don't stop. In the past ten days, we've gotten some groundbreaking large language model updates from Grok out of XAI, from Anthropic with their, new SONNET 3.7, and now OpenAI. So after a lot of waiting, what seems like years, we have OpenAI and ChatGPT's next big step forward with GPT 4.5. But let me tell you something. This one's weird. Not saying in a bad way. It's different. And probably for the first time in a long time, I've spent hours now at least playing with GPT 4.5, and I'm like, this isn't for me.

Jordan Wilson [00:01:13]:
Again, not in a bad way. I just think that OpenAI's new model and GPT 4.5 is something much different than it's released before. It's not breaking any records. It's not climbing to the top of every single benchmark mark. It's a vibes model. It's to feel more relatable and more relational with the end users, all of us. Alright. So I'm excited today to talk about OpenAI's new GPT 4.5, what's new and who can benefit the most.

Jordan Wilson [00:01:51]:
Alright. If you're excited to learn about it, you're in the right place. Welcome. My name is Jordan Wilson, and I'm the host of Everyday AI. This thing is for you. It's your daily livestream podcast and free daily newsletter, helping us all not just learn what's happening in the world of AI, but how we can all actually leverage it, what it means, and how we can use the information to be the smartest person in AI at your company. If that's already you or if that's what you're looking to do, welcome. We just became best friends.

Jordan Wilson [00:02:21]:
Also, your actual best friend is our website, youreverydayai.com. Go sign up for our free daily newsletter there. Also, people don't know this. You can go listen to every single podcast episode ever on our website. Go watch every single video. There's like close to 500 now, all sort of by category. So no matter what you wanna learn, we have it for you. We've probably had a world's leading expert already come and share their secrets.

Jordan Wilson [00:02:44]:
So make sure you go check that out. Alright. So I am excited to talk, today about, OpenAI's new model. But before we get started, we're gonna start off as we do most days by going over the AI news. So, NVIDIA's revenue has soared 78% year over year to 39,300,000,000.0 in the fiscal fourth quarter ending January 26 fueled by strong demand for GPUs. So the company's data center products accounted for 35,000,000,000 of the total revenue with half of that coming from cloud service providers like AWS, Google Cloud, Microsoft Azure, and Oracle Cloud. So NVIDIA's Blackwell GPU, which was launched in December, generated 11,000,000,000 in revenue in the first quarter, marking the fastest product ramp in the company's history. So, yeah, NVIDIA's earnings coming out.

Jordan Wilson [00:03:41]:
So pretty pretty interesting here. And, CEO Jensen Wong announced the upcoming launch of Blackwell Ultra in the second half of twenty twenty five, promising a smoother transition compared to the hopper to Blackwell Shift, which faced production changes due to design changes. Sorry, production challenges due to design changes. So Blackwell Ultra will feature advancements in networking, memory, and processors, while NVIDIA's next generation Verorubin architecture is combining CPU and GPU technology is set to debut next year in 2026. So Wong emphasized that NVIDIA's manufacturing partner, TSMC, exceeded expectations in expanding production capacity, helping meet surging demand despite initial hurdles. So NVIDIA's revenue from China has dropped by half since US restrictions on chip exports began in 2022, but the company now offers a less advanced processor, the h 20, specifically for the Chinese market. Alright. Next piece of AI news, Meta is looking to compete with chat GPT in a completely different way.

Jordan Wilson [00:04:48]:
So according to reports, Meta is planning to release a standalone Meta AI app in the second quarter of twenty twenty five according to sources familiar with the subject, marking a pretty big step, in CEO Mark Zuckerberg's push to dominate the AI space. So the app will reportedly expand Meta AI beyond its current integration with Facebook, Instagram, WhatsApp, and Messenger, allowing users to interact more deeply with the Gen AI assistant. So right now, you can obviously just go to meta.ai and use, their AI that way. But it looks like, Meta is looking to compete more directly, with OpenAI as a standalone AI app. So in April 2024, Meta replaced the search feature in its apps with Meta AI positioning the chatbot as a central feature for billions of users. So the new standalone app will allow for greater personalization, conversational history organization, and integration with Meta's hardware such as Ray Ban smart glasses according to, Zuckerberg publicly agreed with on threats. So Meta is also exploring a paid subscription for Meta AI. Interesting.

Jordan Wilson [00:06:05]:
Right? A fairly open source model, but you will have to pay to use it. Similar to OpenAI's Chat GPT plus and Microsoft Copilot, which could generate revenue through premium features and paid recommendations. So Meta AI currently has 700,000,000 active users. So, yeah, pretty pretty wild there. It should be interesting. And, you know, CEO Sam Altman, OpenAI CEO, kind of responded jokingly on Twitter and said, you know, hey. Maybe we'll just release a social media app. Alright.

Jordan Wilson [00:06:36]:
Let's get into it. Let's talk about what's new inside OpenAI's GPT 4.5. And this is part one of two. Right? I understand y'all. Sometimes these shows go way too long. And the other day, I'm like, oh, yeah. There are some updates. We're gonna do a short show.

Jordan Wilson [00:06:53]:
And that show ended up being an hour. Whoops. I gotta stop doing that. Right? No one wants to listen to me blab when I'm tired and over caffeinated for an hour plus. So we're actually gonna be breaking this one down. I'm not gonna be doing any live demos today. Those usually take a lot of time, to put together. So probably in the future when there's at least big new models like this, we're gonna break it up into two portions.

Jordan Wilson [00:07:16]:
Just like I always say, hey. We're gonna learn and leverage. So today, we're going to learn about the model, what's new, who I think it's gonna benefit, and then we're gonna have a second show probably next week on, you know, the best ways to leverage it. Probably do some live demos, some examples, all that good stuff. Alright. So what the heck is GPT 4.5? Well, it is the last non chain of thought model from OpenAI. So in the future, Sam Altman has said that future systems are going to be, hybrid. So what that means is, reportedly, GPT five will be more of a system, and you aren't going to necessarily be choosing between these quote unquote old school transformer models like GPT four o or GPT 4.5 and, reasoning models like o three and o one.

Jordan Wilson [00:08:13]:
Right? So in the future, it said it's gonna be more of a hybrid approach, and it's gonna be a system that you talk to. And maybe the system is just going to choose, which model is best for your query or maybe it's going to use hopefully one of my predictions, one of my AI twenty twenty five predictions, which you should go listen to those shows, is moving away even from a mixture of app experts and going to a mixture of models. I hope we see that. Right? I hope if, in the future, if you have a very advanced query, part of it, might use in theory under the hood a GPT, type model, and then part of it might use an o model. But it is the last non chain of thought model from OpenAI, and it's really expensive on the API side. So right now, just FYI, this is only available to pro users. It's only available to people on that $200 a month plan. Although OpenAI did say that it will be rolling out in the coming weeks, to all paid users.

Jordan Wilson [00:09:12]:
So, you know, I I feel most most people listening to the show, are probably chat GPT plus on the $20 a month plan. So you don't have this yet, but probably, I'm guessing sometime early to mid March, most paid users should have access to GPT 4.5. But it's super expensive on the API. Right? So, you you know, developers or, you know, maybe if if you are a technical person, in something at your company, runs on the back end, on on GPT, you know, four o maybe or four o mini, you're probably not gonna be using this, if I'm being honest, and more on that in a bit. But, ultimately, I think humans are gonna like this. Right? I think humans are gonna like this. And, hey, livestream audience, thank you for tuning in. I forgot to shout you guys out.

Jordan Wilson [00:10:01]:
But if you do have a question, let me know. So thanks for big Bogey and Harvey and Samuel joining on, YouTube and, Woozy Rogers, joining us, on LinkedIn. Doctor Harvey Kaster doing double time joining us on LinkedIn and YouTube. Love to see it. Steven, happy Friday to you and Brian and Joe, Michelle, doctor Scott, everyone. Can't go through everyone, but, thank you all for joining live. If you do have questions, try to get them in now. I'll try to answer them either as we go, as they pop up.

Jordan Wilson [00:10:32]:
Right? Yeah. This is a unprompted, unscripted live stream, the realest thing in artificial intelligence. So, you know, get your questions in. I'll try to either tackle them as we go or at the very end. So here's, some more details on what you need to know on the new GPT 4.5 model. So like I said, it is 200 a month right now to use on the pro plan. So that's if you're using it on the front end chatbot. Right? Logging into chatgbt.com is not gonna be there unless you're on that $200 a month pro plan.

Jordan Wilson [00:11:01]:
But it should be rolling out in the coming weeks. It's the first major major model upgrade in over two years though. So that's important. So we've seen, iterations and upgrades over the GPT four model. Right? But it's been more than two years since this base model was actually refreshed. Right? Let me tell you what I what I mean by that. So GPT four, came out, back in, gosh, 2023. Right? But it was kinda refreshed.

Jordan Wilson [00:11:34]:
So then we went to GPT four turbo, then we went to GPT four o or Omni. Right? So it was this Omni model bringing more modalities, but the base, the engine was still kind of the same. In OpenAI, you know, they did some fancy engineering and, you know, tweaked it a little bit. But for the most part, it was an old engine that was still running this thing, but it was still the most powerful, single use model in the world. Right? And when I'm talking about single use, I I'm meaning nonreasoners. So this is pretty big. It's the first major model upgrade in over two years. But here's the thing.

Jordan Wilson [00:12:16]:
It's built for empathy. It's built for relationships. It's built for intuitive conversations. Right? People are saying this is a vibe model. Right? It's not shooting off, the charts on every single benchmark, but OpenAI hopes that when you talk to chat g p t, you're like, oh, this is very human like. Right? They're they're like, oh, you're gonna feel some AGI vibes, some artificial general intelligence. Right? But here's I I started the show by saying this. This is not for us.

Jordan Wilson [00:12:51]:
Right? If you were if you're a power user, if you're following AI every day like me, right, depending on the day, I know it changes. You know, I'm spending, who knows, anywhere from three to eight hours a day using large language models. For the most part, Chat GPT. I'm in Chat GPT all day. This isn't for me. This model is not for me. It's really not. I think it's for everyone else.

Jordan Wilson [00:13:19]:
Right? This is for casual users. This is for my mom. Right? This is for my mom. This is for companies that maybe did not get on board with AI previously. So it's not just like that because of a a vibe. That's a vibe model. Right? Oh, it feels good. It feels natural.

Jordan Wilson [00:13:41]:
It feels human. It it feels like it understands my emotions. Right? Because it's not the best. It's not the fastest. It's not the cheapest. So then it's like, what the heck is this thing then? If it's not the best, it's not the fastest, it's not the cheapest. It's not for power users. Who like, what what what the heck, OpenAI? I really focus this on two things.

Jordan Wilson [00:14:09]:
You know? And, again, this model literally just came out. I wasn't part of the early testing group, so I look tired if you're on the live stream and the coffee is probably a little stronger. But I think probably if I had to boil this down to two words, it would be reliable and relatable. And, again, for me as a power user, I've never had problems with those. Right? I don't run into a lot of hallucinations because I know prompt engineering very well. I know how to make sure and to refine a large model, and kind of train it up on a very smaller skill set and to, increase the accuracy and decrease the hallucinations. But now out of the box, hallucinations are lower. Now out of the box, it's not gonna sound like talking to a robot.

Jordan Wilson [00:15:02]:
Right? Maybe that's why for the first time, I'm like, yeah, this doesn't really seem like for me. It's still probably going to be a model that I use very often even though it's not the best, not the fastest, not the cheapest. Right? But I I will assume that, you know, GPT four o won't be around for forever. Right? So I I I do need to also understand that I need to start using this model. I need to get used to it. I need to adjust how I talk to it. I need to adjust my expectations. Right? This is why also I'm updating our free prime prop polish course.

Jordan Wilson [00:15:41]:
Yeah. I know it's been a few months. Don't worry. We're shooting for a March date. Keep keep an eye on the newsletter for that. Right? This isn't for me, but I'm still gonna use it. Right? It's funny because I think for probably a good fifteen years, you know, my friends and coworkers, have called me a computer. Right? They're like, oh, Jordan's a computer.

Jordan Wilson [00:16:05]:
Yeah. He's not human. Beep boop. Beep boop. That's why for me, right, I don't know. I don't like, for me, I don't need a relatable chatbot. I don't. I don't, like, I don't need to talk to something and be like, oh, this feels human.

Jordan Wilson [00:16:22]:
Right? Maybe it's because, you know, my EQ is not off the charts. Right? But if you are someone, that really cares about feeling heard, about feeling understand, or sorry. Understood. If you want to feel a relatable relationship with an AI chatbot, not in a weird way. Right? But when I think about things like biz like a business coach, a strategist for your company, for your department, right, a true creative thought partner, GPT 4.5 is going to be much better at those things. Angie says Jordan's an AI agent. Sometimes I wish I was. I think agents don't need sleep and agents don't get tired.

Jordan Wilson [00:17:16]:
Those are two things that I'm both struggling with, right now. Nancy, former everyday AI guest. What's up, Nancy? Says the single reason I couldn't live without the pro subscription of Claude is because I couldn't stand to talk to Chad GPT all day. LOL. Yeah. That's a great point. Right? Because I've never liked Claude. I know people do because, you know, they're like, oh, it spits out more human sounding content, and it feels more like I'm talking to a human.

Jordan Wilson [00:17:46]:
Well, this is I'm not saying that, GPT 4.5 is OpenAI's answer to Claude. It's it's not, but, you will get those vibes. Right? You'll get those vibes that the output, the written text is going to look and seem much more human like. It's going to seem like a much less robotic process both in the output and in the interaction between you and the chatbot. So, yeah, a lot of people, are saying, you know, and and Nancy is definitely not alone in this. Right? That people a lot of people prefer Claude, who use AI just for content writing and who don't wanna necessarily go the extra mile, in prompt engineering, and they just wanna be able to get more human sounding output out of the gate, out of the box, and they wanna be able to, have it feel more like the, assistant understands you, like the AI chatbot understands you, right, which is something I think Claude's been great at. Again, for me, I'm a human or, like, I'm a human, but I don't know. I feel like thinking bits and bytes.

Jordan Wilson [00:18:46]:
I think in ones and zeros. So I don't necessarily need to feel understood, by an AI or anything like that. Right? But that's a great point there. So open and and this this is interesting y'all. So OpenAI says GPD 4.5 is not a frontier model, which is wild to think. Right? And that means that it it's the model that does not represent a groundbreaking or revolutionary advancement over its predecessors. Right? OpenAI strip said this. They're like, yeah.

Jordan Wilson [00:19:14]:
This isn't like benchmarking off the charts. This is not they literally said this is not a frontier model. Right? Frontier models are are those large language models that are supposed to be revolutionary. Right? This is not it. This is more of a foundational model. But I think here, what we're actually doing doing is this is building for the future. I think this is all about the training data. This is about how we're interacting with this model, and OpenAI is obviously collecting all of that.

Jordan Wilson [00:19:45]:
They're not collecting the data that you upload. Right? FYI. People always get that wrong. People are like, oh, anything I upload into ChatGPT, it's it's like, you know, it's like I'm printing it on the Internet. No. It's not what that is. Turn off, turn off your your data sharing, and then you're not sharing anything. Right? But you always have an opportunity.

Jordan Wilson [00:20:02]:
Chat g p t will ask you. Sometimes it'll give you two responses. Which one's better? Right? That's being sent to OpenAI. Right. If if you say this is wrong, that's being sent to OpenAI. So what I think is actually happening here is there's a big, large expensive they didn't say how many parameters this model is, but apparently, it's it's enormous because the API costs are insanely high. Right? I don't know who's going to be using the GPT 4.5 API. I'm gonna show you the prices.

Jordan Wilson [00:20:32]:
It doesn't compute. Right? Using that much compute doesn't compute. So this model has to be enormous, but I do think that this is going to be the last enormous model, from OpenAI, because I think what this is setting the stage for is to get that data, on how users like and interact with the model and also those that don't turn off data sharing. Right? And I think this is gonna lead for better and smaller, distilled models for, like, as an example. When we talk about the GPT five and, you know, the o four, o three models of the future, I think are just going to be distilled based off of this super big model. So it's expensive. And, also, this thing maxed out OpenAI's compute. CEO Sam Altman literally said, yo.

Jordan Wilson [00:21:23]:
We're out of GPUs. Right? At least that's what he said. You know, he said, hey. We can't bring this out to, all chat GBT plus users right now because it'll burn us. He literally said we're we're out of compute. We're out of GPUs. Right? Which means this thing is enormous. Like I said, the benchmarking improvements modest, not meaningful.

Jordan Wilson [00:21:43]:
Right? A lot of times when you get a, you know, new frontier model, it completely shifts the conversation on benchmarks, and you're like, well, this one went through the roof. This is not that. And it's designed for more natural human interaction. And like I said, I think this is laying the foundation for future smaller models and it more combines the, EQ with the IQ. Right? That emotional intelligence with traditional intelligence. That's what this is. This is a much more human touch. Alright.

Jordan Wilson [00:22:14]:
Here's what OpenAI said specifically about 4.5. Said we're releasing a research preview. Yeah. This is a research preview y'all. Keep that in mind. Our largest and best model for chat. They didn't say our best model. They said our best model for chat.

Jordan Wilson [00:22:33]:
Best model for humans to chat with. Right? GPT 4.5 is a step forward in scaling up pre training and post training. By scaling unsupervised learning, GPT 4.5 improves its ability to recognize patterns, draw connections, and generate creative insights without reasoning. Early testing shows that interacting with GPT 4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater EQ, emotional intelligence, makes it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less. Remember those two words I said? I think it's gonna be more reliable and more relatable. So who can benefit the most? I told you kinda what's new, gave you some of the bullet points.

Jordan Wilson [00:23:26]:
Who can actually benefit the most? Right? Like, when are you gonna use this? So I think for more human like conversations, it's going to be ideal in the long run for companies to use this for customer support, or for you to use this for customer support. Right? Maybe you're in customer service and you just copy and paste a bunch of information in here and you're trying to work through, tough customer support problems. I think it is great for understanding nuances in human language. Right? It's probably going to be pretty soon better than humans at understanding nuances in human communication, which is weird to think about. Right? So I think it's ideal for customer support, therapy, education. It has, enhanced creativity that I think will help writers, marketers, and designers generate creative ideas. And for some, advanced coding and technical abilities will benefit developers and data analysts. Not everything.

Jordan Wilson [00:24:25]:
Right? This is not gonna be something that you're gonna, you know, plug in and and and use to code. I don't think that's what we're gonna see here. Although, interestingly enough, even though the benchmarks did not shoot up, this is pretty interesting. So Cognition Labs, right. So they have Devon, which is an AI programmer, very popular, and they were kind of looking side by side looking at different models for agentic coding evaluations. Right? And even though GPT four four point five is not necessarily supposed to be a coding tool, it did very, very well on their evaluation. So as an example, its predecessor GPT four point o got a 49 on this agentic coding evaluation and GPT 4.5 got a 65%. You know, only trailing SONNET 3.7, which got a 67.

Jordan Wilson [00:25:19]:
So, again, it's not gonna be used. You know, programmers aren't gonna use this. Developers are gonna use this because right now the API costs are high. It's slow. Right? Even using it, you know, obviously, anytime a model first comes out, it's always gonna be slower. But I expect this to be slower in the long run, so it's not it's not the fastest. It's not the best. It's not the cheapest, but it's very capable even when it comes to agentic coding.

Jordan Wilson [00:25:43]:
So that's according to Cognition. Let's talk about the emotional intelligence. So more natural human like interactions than GPT four o. It's going to be better at reading and responding to emotional cues. And according to OpenAI, it was preferred by users in, about 56 to 63% of different use case tests against, GPT four o. So right? Showing kind of two, responses side by side. So the majority of the time, users prefer this to GPT four o. So, it will presumably and in my limited testing so far, this is true.

Jordan Wilson [00:26:20]:
Great at storytelling and generating ideas and just generating written content. Right? Which, that's something that you use AI for, which I know a lot of people do. I think there's so many use cases people should be using AI for, but they're not. And they're just like, yo. I need help writing this blog post, or I need help writing a paper. Right? And, ultimately, that's what they're using it for, which I why I think a lot of people flocked to Claude early on, and I'm like, nah. Like, you can do this in chat g b t. You just gotta know how to use it.

Jordan Wilson [00:26:48]:
Right? So I think it's going to, become a much better writer. It's going to write more clearly and concisely, and, also, I think it's going to be much stronger in design and creative tasks. This is a good news y'all. This is good news. Second straight day that the sun is shining in my face, and I had to close the curtain here. Oh, bless up. There's nothing worse than waking up for a livestream in the months of winter, and it's just dark outside. Right? Sunshine.

Jordan Wilson [00:27:20]:
Right? Maybe I need to go outside and touch grass, and then I'll appreciate this more human, side of GPT 4.5. Alright. Let's talk about some of the technical features. So, according to OpenAI, it is trained. This was trained with 10 times more computing power than the previous GPT models. Also, a 28,000 token context window for deeper conversations. I'm going to be testing that one ASAP because for years, OpenAI has said and maybe they just didn't differentiate and maybe this is just the API and they didn't say, hey. It's a 28,000 in the API versus 32,000, when you're using chat g p t.

Jordan Wilson [00:28:02]:
So I'm gonna be testing this. Don't worry. And I'm gonna talk about that in our part two of the show because for so long, when you're using the chat version, right, the front end chat gpt.com, it hasn't had a hundred and 28,000 token, context window. Right? So that means that chat g b t will start to forgive things much sooner. So it actually had a 32,000, token context window, which is about 26, 20 seven thousand words of back and forth interaction with chat g b t, and then it would start for to forget things. So I'm excited to test that out to see if that is just, on the API side or if that's going to be also in the chat window. If that is in the front end chat, that's gonna be big. Also improved coding, especially in complex tasks.

Jordan Wilson [00:28:45]:
So, again, according to OpenAI, this was pre, pre trained simultaneously in multiple data centers, which I believe will be the first, for a model like this. Right? That says something. When, you know, OpenAI has access to some of the biggest, data centers and the most compute in the world, and they're like, yo, we can't train this in one place. This is too big. I don't know how big this model is. Right? It's gotta be multitudes larger than the original GPT model. Right? If it's costing this much, in compute, if it's costing this much via the API, if it's causing OpenAI to run straight up, run out of GPUs, it's gotta be huge. I don't know.

Jordan Wilson [00:29:30]:
Right? Reportedly, earlier versions of GPT four were about 1,800,000,000,000 parameters. I don't know. This thing's gotta be double that. Maybe, maybe more. I don't know. But I that's not sustainable in the long run. Right? Which is why I think this is actually a foundation for OpenAI to better distill and to make better smaller hybrid models, as they switch to that, that kind of setup. Also, can handle task involving visual understanding.

Jordan Wilson [00:30:01]:
So I did some tests on this. Vision capability is pretty good so far. We're gonna do more more on that in our part two show. We're probably gonna show some comparisons between GPT four o and GPT 4.5. But out of the box does very well with visual understanding and, being able to see, and synthesize, information in, photos. So, yes, this is multimodal. FYI, right now, it has access to all of the tools. I should have maybe started with that.

Jordan Wilson [00:30:29]:
Right? Because, these reasoning models, a lot of them, the o one series doesn't have access to all the other tools. Right? Canvas and, you know, Dolly and and Vance data and and chat g b t search. Right? All these other tools that really make a large language model agentic. Right? Didn't have access. So right now, you do have the rest of the tools. Although I would like and hope that eventually we will see tasks, get GPT 4.5 as well as GPTs. My gosh. OpenAI.

Jordan Wilson [00:31:04]:
I know there's, you you know, a lot not a lot, but plenty of you all listening because you reach out and let me know. Can we update GPTs, please? These poor things. They're just like that that poor forgotten about child in the corner. Right? This is Macaulay Culkin in Home Alone. You know, we're leaving to the airport without GPTs. GPTs, y'all, enterprise companies, they hire us and they want to build us GPTs. Right? And I'm like, y'all, like, I don't know. We might have to build you projects instead because poor GPTs are in the corner and they haven't been updated in forever.

Jordan Wilson [00:31:37]:
Right? So, hopefully, we see the GPT 4.5 model eventually be rolled out to other things like tasks and like GPTs. Here's the thing, reliability. Let's talk about accuracy and knowledge because I started the show up by saying a lot of companies didn't get on board with AI because they're like, yo, it lies. It hallucinates. Is GPT 4.5 free of hallucinations? Absolutely not. If you know how to use it, you're probably going to see a great reduction in hallucinations. But, according to OpenAI, it knows more and hallucinates less. So, the hallucination, rate has gone down significantly, higher accuracy and factual questions.

Jordan Wilson [00:32:17]:
But what's important to know, the knowledge, cutoff has actually been rolled back. So, GPT four o has a knowledge cutoff of June 2024, which is reasonable to work with. This one is October 2023. Right. So I'm sure they'll be updating the knowledge cutoff in the future, but just know if you're using GPT 4.5 right now, in the chat or when you're using it, when it rolls out to chat GPD plus users, you should probably, in many use cases, use our refined cue method that we teach in our free prime prompt polish prompting course. Okay? You need to bring in more accurate and more up to date information for whatever it is you're working on to get started with, or make sure you go retrieve that by using chat g p t search. Right? Here's the thing. I'm gonna have to do a dedicated episode just on training data and what this means.

Jordan Wilson [00:33:09]:
Right? So people think, oh, that means it knows every single thing, and it's a % accurate and up to date by October 2023. No one doesn't. Right? A lot of these datasets that companies use to train their models, you know, by saying, oh, it cut off in October 2023. Well, what happens to that data set is updated once a year. What happens to that data set has some extremely outdated information. You hope that through reinforcement learning with human feedback, you know, a lot of that older information gets kicked off when they're going through and they're, you know, training the model, but not necessarily. So keep that in mind. The knowledge cutoff is rolled back.

Jordan Wilson [00:33:44]:
You need to do a better job. If you are using GPT 4.5, need to do a better job at making sure it has more accurate, more up to date, and relevant up like, fresh information if you are relying on it for accurate up to date outputs. The entire world changes around us every single day. So to work with knowledge, a knowledge cutoff from 2023, you gotta be careful. Right? It is computationally demanding. So like we said, it's very limited right now. OpenAI's ability to scale this out to users, because of GPTs, sorry, because of GPUs. Also weaker, it does have weaker comp performance and complex reasoning compared to specialized models.

Jordan Wilson [00:34:29]:
Alright. Let's talk a little bit about accuracy and knowledge. Alright. So this is simple QA, which is actually OpenAI's own benchmark. Right? I would really like other people to start using this or something like it. But this is essentially like, is this getting things correct? Right? So simple QA accuracy where higher is better. This is just is it factual? Is it getting questions correct? Can it, recall information in the right way? So on this, g p some of GPT's, previous models or some of OpenAI's previous models. So GPT four o scored a 38% on this where GPT 4.5, not double, but pretty pretty close, got a 62%, where even the reasoning models got a 47% and a 15%.

Jordan Wilson [00:35:32]:
So if you're wondering what's the point of this model, Boil it down to two words. It's relatable, and it's reliable. It has a much higher accuracy. And let's let's be honest. We just sometimes look past large language models, and we just assume that it's always accurate and we can always rely on them. That's bad. I don't know why people are trying to take human out of the loop, and we expect large language models to always be a % factual and accurate. Right? They're trained off the Internet.

Jordan Wilson [00:36:08]:
Is the Internet a % factual and a % accurate? Absolutely not. Right? I read, in in in in an article on I think it was chat g b t from a huge publication last week, and it was completely wrong. All their facts were wrong. Right? I'm not gonna name shame them. Maybe I should, but a a publication we've all heard of. Every everyone out here reads it. I was thinking about, like, roasting them on Twitter and fact checking it, and I'm like, this is all wrong. This is all not correct.

Jordan Wilson [00:36:37]:
Right? But guess what? All this information that people put out on the Internet, sometimes people intentionally put out misinformation, disinformation. Sometimes people don't know what they're talking about, but all that goes out on the Internet. Models gobble this up, and you hope that humans can pick out, you you know, information that's in the training data that's not right versus what's right right through reinforcement learning. But much more accurate. Almost twice as accurate. And what's what's pretty interesting for me at least is o three Mini there with a 15% on this simple QA accuracy and a GBT four five with a 62%. I love o three Mini. It's probably my most used model.

Jordan Wilson [00:37:19]:
Right? Again, I do a good job at making sure I feed it the accurate and relevant information that it needs, and I'm not necessarily always relying on it to, go and seek and find the absolute truth on its own. But out of the box, Jeep GPT 4.5 according to OpenAI's own internal benchmarks extremely reliable. And let's talk about hallucination rate. Same thing. Much lower. In this case, lower is better. So in their test, it's only, it's a thirty seven percent hallucination rate. Now I want you to keep in mind, that doesn't mean it hallucinates thirty seven percent of the time.

Jordan Wilson [00:37:53]:
In these tests and in these benchmarks, they're hard. They're tricky. They are made, to get the model to kinda screw up. Right? So a very, very, very low hallucination rate actually for 4.5, with a thirty seven percent where o three Mini as an example, eighty percent, and a the g b t four o at sixty one percent. So, again, these are intentionally very difficult questions that are meant, to make models hallucinate. So it's more reliable. It lies less. Alright.

Jordan Wilson [00:38:26]:
Other benchmarks, again, nothing here is jumping off the page. Many of the major benchmarks, this is not OpenAI's best model. Right? It's in some cases, it's actually about the same or on par with GPT four o or it's just behind, o three minutei, which, again, that is my workhorse model. I'll probably do let me know, livestream audience. Let me know yes or no. Should I just do a show where I tell you what models I'm using and for what? I might have to wait, a couple of weeks to see how and where I'm using GPT four o. I had some people ask about it recently. I didn't think it was that interesting, but, you know, if it's interesting, let me know.

Jordan Wilson [00:39:12]:
And maybe I don't know. Maybe it will be more interesting now that we have, like, nine models to choose from. But one thing that I thought was pretty impressive, about these benchmarks. So it did score better in the which is the multimodal, equivalent of the MMLU, and it scored fairly well on the MMLU, which is the multilingual equivalent of MMLU. So MMLU has historically been one of the, you know, it's one of the benchmarks that we talk about most. I say it's like the ACTs, for AI models. Right? So it did perform well or better than GBT four o, in those models pretty significantly, but, Sweet Lancer, I love this. So this is, an actual test that, OpenAI developed and, you know, other model they use other models.

Jordan Wilson [00:40:03]:
Right? And, essentially, like, when Claude came out, Claude was better. And OpenAI said that. They're like, yo. Claude does way better at Sweet Lancer. Right? This is essentially a a test where, it goes out and performs the type of task you would see on, like, Upwork. Right? But this one outperformed, the other models, by far. It it completed 32.6 of tasks, whereas OpenAI o three Mini completed 10.8%, and GPT four o completed 23%. So that's interesting.

Jordan Wilson [00:40:36]:
Also, o three Mini, I'm guessing at the time, did not have access to all the same tools. O three Mini does have access to the Internet, which is huge because the other o models do not have access to the Internet. Alright. Here we go. Here we go. The, costs. I don't know. You know what? I actually can't wait to talk to companies that are using this on the API because I'm not sure who is gonna use it.

Jordan Wilson [00:41:06]:
It costs $75 per million input tokens and a hundred and $50 per million output tokens. So expensive. So I guess it's it's gonna be those people that, those companies that really value, a reliable and relatable model. Right? So this is just if you're using it on the back end of the API. Right? So if you're logging into Chadgpt.com, you don't gotta worry about this. Right. But I do I would assume that when this does roll out to, plus users, it's it's it's gotta be limited. I don't see them rolling out, this extremely expensive model that they're probably gonna be losing money on.

Jordan Wilson [00:41:54]:
When you look at the API cost, I don't see, ChatGPT plus users getting unlimited access to this. I I would assume that there would have to be some rate limits. So, let's go ahead and look at some of the cost comparisons. Ready? So I said $75 per million input, GPT four o, $2.50. 2 dollars and 50 cents. So we went from $2.50 to $75. Yikes. Yikes.

Jordan Wilson [00:42:31]:
30 times more expensive. Is that right? Did I I just did that math in my head. Hopefully, that's right. And then the output, 15 times more expensive. The output for 1,000,000 tokens for GPT four o, $10, and then on GPT 4.5, a hundred and 50. Right? And, like, everyone was losing their marbles when Claude three point seven saw it, less than a week ago came out and their API pricing didn't change. Right? And everyone's like, oh, Claude SONNET is so expensive. Right? And aside from if you're using it from coding, there's no need to ever use Claude SONNET 3.7 via the API.

Jordan Wilson [00:43:06]:
Now you're looking at g b d four five, and you're like, alright. Well, Quad 3 7 SONNET doesn't sound like that bad. Right? $3 per million tokens, input and 15, 4,000,000 output. So yes. I mean, we were looking at the, Claude three point seven SONNET versus GPD four o, and Claude is like, oh, it's like, okay. Well, that's 50% more expensive for output. Oh my gosh. And then GPD four or five comes out and says, hold my GPU.

Jordan Wilson [00:43:37]:
Right? You won't believe this price. I I don't believe it, but we'll see. We'll see who uses it. Clearly, someone's gonna use it. So let's talk about the strategic impact. Like I said, I think this is a base model for future AI development. I think it moves OpenAI towards integrating soft skills with technical skills. Right? That's like when we talk about, oh, this is a vibes model.

Jordan Wilson [00:44:01]:
This is an EQ model. Right? Where I think previously, which is why I never had a problem with it. I don't need vibes. I don't need EQ, but I think a lot of people do. I think looking at it even as soft skills. This is a soft skills model, which I think why it's actually a big step forward. But it's a big step forward in areas that we're not used to looking at. Normally, we look at big step forwards in in AI models, in benchmarks, in features, but we're not looking at it in terms of it being more relatable and reliable like a human.

Jordan Wilson [00:44:36]:
So I think that's a big thing is this is going to be probably the most human model out there. Is it gonna be the best at certain tasks? No. Is it going to be the most reliable and relatable model? Probably. And I also think that this is indicating a possible limits to continued scaling with the current GPT versus o series architecture. So, yeah, next, we're gonna see this hybrid setup. Alright. So that's a wrap y'all. Look at that.

Jordan Wilson [00:45:11]:
We didn't go one full hour. Bless up. So we're gonna do a part two next week. Let me know what use cases do you wanna see. Let me know in the comments, today in the newsletter, so make sure you go sign up to it. You can just reply to the newsletter. Let me know what do you want to see. Do you wanna see writing use cases? Do you wanna see a a creative strategist? What do you wanna push the boundaries on? So we'll we'll do this show, likely, either next week or the week after.

Jordan Wilson [00:45:41]:
I gotta look at what we have scheduled. We have some great guests coming up y'all. I'm very excited. I know sometimes, you know, the show I do a lot of the shows. Sometimes we go through periods where it's a lot of gas. Sometimes it's a little bit in between. We have some fantastic guests coming up. But let me know what type of, hands on you want to do.

Jordan Wilson [00:46:00]:
We're gonna do it live. We're gonna do it hands on. Let me know what you want to see. Also, go to our website, youreverydayai.com. Sign up for our free daily newsletter. We're gonna be recapping this, but like I said, high level here, GPT 4.5 is out. It's only out for, pro users right now on that $200 a month plan or if you're paying through it via the API, which is crazy expensive. This is not a groundbreaking model by traditional metrics.

Jordan Wilson [00:46:31]:
Alright? But I do think it could be a groundbreaking model by just the vibes, by how we feel, by how we interact and think about the AI. Alright. A couple questions. Let me see. Douglas was saying, how would a mixture of models compare to the idea of a reasoning orchestrator and then transformer agents specialized? Oh, Douglas, you're really trying to push this episode to more than an hour. Alright. I I think I covered most of that in the AI predictions, show. So say I'll say go listen to that, Douglas.

Jordan Wilson [00:47:02]:
Maybe maybe I'll leave you a more thoughtful comments, on the live stream later and explain that. Yashel, sorry if I got that wrong, is asking, Jordan, would you use 4.5 instead of four o moving forward? Here's the thing. It's better. 4.5 is better than four row. Right? It's just not better in the same step that normally a new model would be. Right? There's very few metrics, or, instances where four five is going to be worse. Right? Which is interesting because a lot of the chatter so far around, SONNET three seven is a lot of people are saying it's worse for certain situations than three five. It's still too soon to answer that, but from everything that I've used it for, unless I need speed out of four o, which is usually not something I'm looking for.

Jordan Wilson [00:47:53]:
Right? I'm I'm patient enough, but I don't think that I'll be using four o much, except, you know, in GPTs. Right? Except with tasks. But for the most part, if I'm looking at a non reasoning model, I'm probably gonna be using 4.5. Samuel asking, does 4.5 support live voice, Canvas, etcetera? So, voice is still powered by four o, but you can be in, four five mode and use voice. I did test that last night. So it's not a new voice model, but four five still integrates and work with voice mode. It does also work with Canvas. I did test that last night.

Jordan Wilson [00:48:35]:
I also test the combination of the two. So you could be in, GPT 4.5. You can use voice mode and it will update in Canvas. So pretty cool. Samsara from YouTube is saying, why is Google so bad? They're so bad that no one wants to compare their models against Gemini. I think Google's great if I'm being honest. Right? I think their front end Gemini chat really was neglected until about five to six months ago. I think the new Gemini models are fantastic.

Jordan Wilson [00:49:03]:
I think their integration into Google Workspace leaves a lot to be desired. I think their AI studio is extremely powerful. But, no, I think the Google models, I mean, they're top of the charts for many benchmarks, including the LM arena. Yachell with another great question here. Is there a benchmark to measure EQ for AI products? As far as I know, no. Because I was researching the same thing. So, yeah, that should be interesting. How can you benchmark these soft skills? I don't know if there's going to be one that's developed.

Jordan Wilson [00:49:37]:
I would assume after this model, there will be one that developed. Right now, there isn't one. Doug was asking, are you gonna look for a PPP update that has transformer model and reasoning model for the different methodologies? Great. Great question, Douglas. So, the the PPP and it's still gonna be free. The updated PPP is still gonna, be based on the GPT infrastructure and the PPP Pro also free. We'll go over, prompting for reasoners as well as some other advanced features. Alright.

Jordan Wilson [00:50:11]:
We got through most of the questions y'all. Thank you for tuning in. I hope this was helpful. Let me know in the comments. Please share this with your friends. If this was helpful, you know, I like, our team spends a lot of time putting this together. We want you to be the smartest person in AI at your company. So if this was helpful, please let me know and let others know as well.

Jordan Wilson [00:50:30]:
Share this and go to youreverydayAI.com. Thanks for tuning in. Y'all will see you tomorrow and every day for more everyday AI. Thanks, y'all. And

Jordan Wilson [00:50:41]:
that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit your everyday AI Com and sign up to our daily newsletter so you don't get left behind. Go break some barriers, and we'll see you next time.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI