Episode Categories:
Resources:
Join the discussion: Ask Jordan questions on OpenAI
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course!
OpenAI's Leap Towards AGI: Understanding the New O1 Series
First Model of Its Kind OpenAI recently introduced a ground-breaking AI model, setting the path for novel developments in predicting and understanding human-like reasoning processes. This unique model - codenamed "Q Star Strawberry" - is the first of its kind to encompass a chain of thought reasoning not seen in previous AI iterations.
A New Perspective on AI
Distinct from the GPT series, this new model presents a clear divergence, leading towards the promising realm of Artificial General Intelligence (AGI). Its introduction is synchronized with the expectation of frequent updates from leading AI developers, hinting at a rapidly advancing AI landscape.
OpenAI o1 Breaking Performance Barriers
While this new model is currently limited by factors such as pricing and API availability, its performance benchmarks have proven to be noteworthy. The model outshines predecessor AI models, impressively mirroring PhD-level human performance, especially in MMIOU tests.
OpenAI o1 A High-Growth Future for AI
The unveiling of such innovations forecasts not just frequent updates from companies such as Anthropic and Google, but also rapid advancements from Meta. The anticipation of a competitive reaction from significant players lays the groundwork for a large language model arms race, leading to potential leaps in AI development.
OpenAI o1 Pricing Insights and Challenges
The new generation of AI models isn't yet as accessible as one might wish. With substantial pricing speculation suggesting rates as high as $2,000 per month, the model might be limited to enterprises or those willing to invest heavily in AI's potential.
OpenAI o1 Performance Evaluation
When tested with numerical codes for locked boxes, complex tasks, and word count tasks, the model demonstrated both its capabilities and its limitations. Despite detailed steps during its processing, the model was often confined to providing single correct answers, displaying a limit in its implementation of variability.
OpenAI o1 Limited But Promising
Though faced with certain constraints, such as the inability to browse with Bing or upload files and images, the model still suggests promising leaps for the future of AI. The performance of large language models in solving societal problems and contributing to shifts in the job market encapsulates the potential of this new age AI tech.
Looking Ahead
In essence, OpenAI's release of the new o1 series embodies a pivotal moment in AI history. While the price points and limited APIs may pose challenges, the improvements in reasoning and thought process map out the future of AI - a future where artificial intelligence helps to solve some of society's most pressing issues.
Remember to keep an eye on updates and developments in this rapidly evolving field. While the course may seem challenging, the potential benefits are transformative and far-reaching.
Topics Covered in This Episode
1. Introduction to OpenAI o1 Series
- First model featuring chain of thought reasoning
- Not part of the GPT family—distinct series with limitations like no browsing with Bing, file upload, computer vision
- Code-named "Q Star strawberry"
- Price and limited API availability
2. Performance Analysis
- GPT-4 and O1 Preview's performance on various tasks: locked box numerical code, word count, and creating a world hunger plan.
- Detailed step-by-step process of O1 Preview
- The limitations each model faces, pointing out areas of improvement
3. General Insights on Large Language Models (LLMs)
- Shifts in job markets due to AI advancements
- Potential of LLMs to solve societal problems
- Comparison of GPT-4 and O1 Preview's processing techniques
4. New Model Overview
- Introduction to OpenAI's distinct model series from the GPT family, known as the OpenAI o1 series
- Expected features of the o1 series like browsing, file, and image uploading
- Distinct functionalities and potential pricing tiers of the o1 series
5. Price and Subscription Insights
- Pricing speculation regarding the new series
- Comparison between the subscription price of ChatGPT and the potential enterprise versions
- Detailed pricing for the o1 Preview and the o1 Mini
6. Usage Limits
- Limits on message allotment for GPT-4 Plus accounts, o1 preview, and o1 Mini
- Current challenges users face in tracking message usage
7. Performance Benchmarks
- Unveiling of o1 series impressive benchmarks, surpassing PhD-level human performance
- MMLU benchmarks for the unreleased o1 model compared with previous models and domain experts
8. User Experience Concerns
- The drawback of limited message allowances and the lack of tracking tools
- OpenAI taking user feedback into account regarding these limitations
9. Historical Performance of Leading LLMs
- GPT-4.0's performance in MMLU
- Performance comparison with leading models like Claude 3.5 Opus, and Meta's 405B
10. Future of AI and LLMs
- Massive investment in AI by major companies
- Anticipated advancements and responses from major AI players
- Predicted impact of post-US election on LLM developments
11. Need for Regulation
- Persuasion for creators of LLMs to work closely with US government and federal regulators
- Acceleration in the development of LLMs following regulatory changes
12. Practical Testing of Models
- GPT-4.0 and GPT-4.0.1preview's performance on a river-crossing logic puzzle
- Importance of selecting the correct model for intended tasks
13. User Recommendations
- Issues with login and usability for tools like ChatGPT in browsers
- Overview of different model versions available
14. Experiments in Real Time
- Live demonstrations of LLMs solving logic problems
- The importance of agentic processing in problem-solving
Podcast Transcript
Jordan Wilson [00:00:00]:
This is the Everyday AI Show, the everyday podcast where we simplify AI and bring its power to your fingertips. Listen daily for practical advice to boost your career, business, and everyday life.
Jordan Wilson [00:00:16]:
OpenAI has just released a new model. Well, technically, 2 models. That's not the story. The story is these models are actually one of a kind, as in they don't have any competitors. It's something no one's talking about. This new o one and o one mini model from OpenAI are agentic in nature. They think like humans. They go through human level reasoning and take time to respond.
Jordan Wilson [00:00:55]:
Yeah. That's right. They aren't instantaneous and spit out hundreds of characters before you can even read. It goes through an agentic process and a chain of thought reasoning like a human would. I'm excited to talk to you about that today on everyday AI. And we're gonna be telling you also, not just about open AI's new o one and o one mini models, but also five things that I think you need to know about this new groundbreaking model. So what's going on y'all? My name is Jordan Wilson, and welcome to Everyday AI. This is a daily livestream podcast and free daily newsletter helping us all learn and leverage generative AI to grow our companies and to grow our careers.
Jordan Wilson [00:01:37]:
So technically, double dipping today. That's how big this one is. Yes. 2nd, livestream and podcast for today. So if you're normally tuning in to hear the AI news, by this time, we'll already have a second podcast out. So go check that one, And always on our website at your everyday a i.com. Let's jump straight into it. No fluff.
Jordan Wilson [00:02:00]:
No fillers. Give us about yeah. We'll see. 20, 25 minutes. I'm gonna tell you everything. We did all the research, took our time. We've been reaching out to people at OpenAI. I actually heard back from one of them.
Jordan Wilson [00:02:12]:
So, we're gonna tell you everything you need to know. So first, this is a new class of model, period. I don't know why anyone's talking about that. This is a reasoning based model. So when we talk about other large language models that we talk about all the time. Right? This is the future of work. We talk about anthropics Claude. We talk about Microsoft Copilot.
Jordan Wilson [00:02:37]:
We talk about, you know, technically, Microsoft Copilot uses GPT 4 o. We talk about GPT 4 o from OpenAI. We talk about Gemini, 1.5 from Google. This new model, o one zero one Mini, they are not the same. Right? So this is a reasoning based model, chain of thought, reasoning, extremely impressive, and apparently very good at coding and math. So, this is also considered a preview. Alright? And it is available now as in today. Yeah.
Jordan Wilson [00:03:15]:
No wait list. No blog post announcing. OpenAI shipped it, right after we've kind of been waiting for things like search GPT. We've been waiting for things like advanced voice mode. You you know, we kind of heard some rumblings, and they just shipped it today. So if you have a chat GPT, plus or a chat GPT teams, account, it should be available now as you're listening. We have both. We have free, paid, which is, sorry.
Jordan Wilson [00:03:46]:
We have the free, ChatGPT Plus, ChatGPT Teams, and ChatGPT Enterprise accounts, when companies hire us to teach them. So, yeah, if you want to, have your company learn, ChatGPT reach out. So, right now, the only people that have access, like I said, are the plus in the Teams. Enterprise and e d u accounts should be available within the week and free. All we heard is soon, but for just the o one mini. Alright. So as, Chad GPT did or OpenAI did, with GPT 4 o, there is a 4 o and a 4 o Mini. In this case, there is an o one and o one Mini.
Jordan Wilson [00:04:26]:
Yes. The naming, not a huge fan of, but it is a completely different class. Also, the API has been rolled out, but right now only to tier 5 users, so only those, heavier users right now. So let's get in. That's just an overview. So here are the 5 things that you need to know. Alright? And then we're gonna dive into each one of these 5 in a little bit more depth. So number 1, this is a first chain of thought, reasoning based model.
Jordan Wilson [00:05:03]:
Agentic thinking. Right? That's the reality. And this is a big step toward AGI. Number 2, this is not part of the GPT family. It is different. More on that later. Number 3, right now, it's very expensive, and the messaging rate is limited. My gosh.
Jordan Wilson [00:05:22]:
Yeah. Hopefully, you don't like it too much because you're not gonna be able to use it too much. Also, number 4, the benchmarks are surpassing, obviously, every model that there is, and it's not even close, but it is surpassing PhD level humans. Also, not even close. And number 5, I think this is going to set off a large language model arms race. All the big companies are going to have to respond. Alright? So more on that here. Let's now dive into each one of those five things that I think you need to know.
Jordan Wilson [00:05:58]:
So, like we talked about, number 1. And, also, podcast audience, check check the show notes. We are going to have a video for this as well. You know, if you're listening on the video, you're gonna see. But I'm gonna try to do my best to describe at the very end, we're gonna do run some simple tests. We're gonna do it live. Here's the thing y'all. I've been itching, itching to do this.
Jordan Wilson [00:06:24]:
Even though I have multiple accounts, I haven't even done this yet. So I'm I'm gonna be doing it live, unedited, unscripted. Hopefully, it works. Right? But there's the the the limits are very severe, which is why I'm saving it, and we're gonna find out live. So step number or thing thing number 1 to know, this is the first chain of thought reasoning based model. Huge step toward AGI. So it's like this. All every single model out there right now I mean, number 1, they're so fast, which is sometimes good or bad.
Jordan Wilson [00:06:56]:
Right? But, also, think. Models right now, how they work without proper prompt engineering. Right? If you're listening, you've probably taken our, prime prompt polish course, our PPP course, and you know what? We might have to completely rehash this because, this new o one and o one minutei changes the the rules. It is. Right? You can achieve similar outcomes, as you can in o one and o one minutei and other large language models. But it takes someone highly skilled. Right? I would say that's myself. And, you know, so if you're someone that, you you know, spends at least 5 to 8 hours a day inside of large language models like me, and if you've been doing that for multiple years, you know, even pre, you know, pre chat GPT.
Jordan Wilson [00:07:50]:
Right? Like, our team's been using the GPT technology since 2020. You have to be highly skilled and spend a lot of time to get the results that you can now get in a single prompt because of this agentic workflow. And you're gonna see this. I, kind of have this screenshot here, but there is a toggle after you run a prompt where you can essentially see the chain of thought. Right? In chain of thought, it is a prompting technique. That's why I say, technically, the outcomes that you can achieve right now have already been available, but it has required someone highly skilled, and it has required usually a lot of time. So you'll see it. It's almost like it's going through, in theory, an ideal outcome of what a very skilled prompt engineer would do over and over and over in multiple steps.
Jordan Wilson [00:08:41]:
Right? Extremely impressive. It thinks. It processes. And think of think of it like that. Right? Like, a smart human can accomplish a ridiculous amount of things with a powerful large language model in enough time. Right? Yes. This is slower. Right? So, yeah, it might take 10 to 20 seconds.
Jordan Wilson [00:09:04]:
Like, gasp. Right? But that's because you literally have an agentic workflow. This isn't where, you know, oh, you know, large you know, I hate when people call large language models next token predictors. Yes. They technically are, but they are so much more than that if you know how to use them. Right? So now think of that happening, seemingly dozens of times, in in one step. Super, super impressive. Alright.
Jordan Wilson [00:09:33]:
Second thing to know, this is not part of the GPT family, y'all. So right now, let me just read this from OpenAI. So it says, this is an early preview of these reasoning models in chat gpt in the API. In addition to these model updates, we expect to add browsing, file, and image uploading, and other features to make them more useful to everyone. So that means right now, o one and o one mini do not have these normal features that we expect. Right? The ability to browse the web, upload files, computer vision via image, uploading. Right? Also, it says we plan to continue developing and releasing models in our GPT series in addition to the new open a I o one series. 2 separate things.
Jordan Wilson [00:10:22]:
Alright? And, y'all, the cool thing about having a daily podcast that's been going on for, like, a year and a half, go check the receipts. Right? I cover I cover the rumors, but I've never jumped on the hype. I said OpenAI has zero reason to release GPT 5 anytime soon. This is different. Right? This is essentially, we've been talking about this strawberry. Right? OpenAI strawberry. Before that, it was QStar. That's what this is.
Jordan Wilson [00:10:51]:
This isn't the next iteration of the GPT family. This is essentially a new mode or a new way of thinking. Alright? So, you you know, it's almost like OpenAI just created a fork in the road. So at least right now, it seems for the short term future, it's it's going to be operating like that, at least according to this statement that they put out that, you know, they're essentially looking at it as 2 different series, which is interesting. Right? Because if they say that in this o one series, they're going to be adding browsing, file and image uploading, and these other features, well, then what would separate it from the, quote, unquote, GPT class series? I'll tell you one thing, price. Right? We covered we've been covering it here, on on, everyday AI is there's been these rumblings and rumors that, you know, internal discussions at OpenAI as they're, you know, right now, they're reportedly burning through a lot of money and trying to raise a lot of money, and they've been floating out higher subscription prices. Yeah. I think this is why.
Jordan Wilson [00:12:01]:
I think this is why. Right? Numbers have been floated out, like, 2,000 dollars a month. I don't think we're gonna get to that. Right? Right now, you have very limited use of it. But right now, it's still only, you know, 20 to $30 a month for your plan. So is that what we've been hearing? I personally think so. But if if if they are separating the 2, these 2 kind of, series, I believe that's what they said. Yes.
Jordan Wilson [00:12:35]:
But if they're adding all of the, quote, unquote, GPT features to the o one series, what separates them? Well, I would think price. Right? And maybe this is more for enterprise, companies, and maybe, yeah, you're going to have to pay a couple $100 or $2,000 in the future to use it. I don't know. But right now, go out and use it now. Right? Or or maybe it's just going to continue to be extremely limited. Alright. So, yes, that's our okay. So speaking of limits, number 3.
Jordan Wilson [00:13:10]:
Well, it's number 3 is it's expensive, but also very limited. Alright. Let's talk about the expense first. Alright? So I'm not talking about on the front end of chat g p t. That's the limits, but let's talk about the expense. So API. Alright? So so many companies out there, you know, they're building on top of OpenAI's API. So right now, the o one preview, again, it's called preview, the thing that, the public has access to.
Jordan Wilson [00:13:40]:
$15 per 1,000,000 tokens input and $60 per 1,000,000 output, which, like, 2 years ago was super cheap if I'm being honest. But after the 4 o, that's pretty expensive. Right? Comparatively, 4 o is 5 and 15, whereas this is 15 and 60. So big jump up. But, again, we are getting we are getting agentic workflows in an API. Right? That's wild. What this means for businesses, I am excited. Someone hire me right now.
Jordan Wilson [00:14:18]:
I can't wait to go wow. I mean okay. Anyways, let's stick to price here. But but but but o one Mini, 80% cheaper. $2.50 per 1,000,000 tokens input and $10 per 1,000,000 output. So the o one Mini is looking nice. It's looking nice. Alright.
Jordan Wilson [00:14:42]:
Got the chart there for the rest of you, but, let's talk now about messages. Yeah. This is pretty pretty limiting here. So there we go. Here's here's the downside. Right now, you get on gpt4 o so if you are on the plus accounts. Alright? The $20 a month. Right now, you get 40 messages every 3 hours, which isn't isn't bad.
Jordan Wilson [00:15:13]:
Right? This for o one, you know, technically, o one preview. Right? So o one, you get 30 messages a week. Not every 3 hours, not every day. A week. Thirty messages a week. Oof. And for o one Mini, 50 messages a week. Ouch.
Jordan Wilson [00:15:40]:
Yeah. Remember when I said, oh, why are there these 2 different series, and why is everyone talking about, this this, you know, plan and, you know, potentially much higher prices. Yo. I've been saying this. Go back and look. Go check the tape. I've been saying, as models continue to improve, it's worth 100 of dollars a month as you get agentic workflows. I said that back in 2023 before any of these rumors came out.
Jordan Wilson [00:16:13]:
I said it's going to be worth 100 or 1,000 of dollars. So it's not shocking when this, the information reported this a week ago. I said, okay. Well, if you're getting agentic workflows, technically, if it works, $2,000 is a bargain. You you can't get a smart human to even power your chat GPT account for $2,000. Right? So it's good. Right? Assuming the limits. But right now, on that $20, $30 a month plan, not looking good.
Jordan Wilson [00:16:49]:
Hopefully, enterprise will have more. Alright. And, hopefully, here's the thing I don't like. So I reached out to the head of research at OpenAI, saying the 30 to 50, 30 to 50 a week is going to be lingering. You know? So asking if there's a way to see your limit. Right? What if you're in the middle of a big project, and you don't know. Over a week? Okay. How many did I use Tuesday? How many did I use Wednesday morning? Oh, did I use any on my phone on Friday? You have no way to no way to know right now.
Jordan Wilson [00:17:23]:
So, Boris here said good feedback. Not yet. So, yeah, there's no way to track it, so that's a bummer, but, at least they're listening. Alright. Let's look at number 4. 4th thing you need to know. The benchmarks are nutty. The benchmarks are nutty.
Jordan Wilson [00:17:39]:
They're surpassing PhD human level. Alright. So this is MMLU. Alright? So I'm gonna tell you right now, on the screen if if if you're watching, you got this. Otherwise, I'm going to describe it. So up until today, the most powerful model was GPT 4 o. Now if you have a paid account, the most powerful model you can access is o one, technically called o one preview. But then there is the o one that OpenAI has that they haven't released.
Jordan Wilson [00:18:06]:
Alright? So we technically have access to o one preview and o one mini, but the actual o one model, which is still under wraps, 92.3 MMLU. Alright? And, yes, I know there's problems with the MMLU benchmark. Alright? And to say it very, plainly, right, I I I try to keep things simple here at everyday AI for the nontechnical people. The MMLU is the multitask language understanding benchmark. Yes. There are maybe some better by now, but, historically, the MMLU, it's it's I I I call it, like, the ACT for large language models. So, about 4 years ago, right, the scores in the in in in the 40th percentile were, considered cutting edge. Right? The average human, right, the average educated human, not a domain expert, would score in the thirties.
Jordan Wilson [00:19:06]:
Alright? It's out of a 100. So, you know, if you're a smart human, you'll get in the thirties. Models 4 years ago were in the mid forties. Domain experts, what that means. Literally, think of the smartest human in the world. Get all of the smartest humans in the world, and they all take the m m MMLU, experts estimate, that they'll get about an 89%. Alright. 92.3.
Jordan Wilson [00:19:37]:
92.3. All the other models right now were in the 80 eights. So, you you know, GPT 4 o, Claude 35 Opus, Metas, 405 b. All of the leading models had been stuck, essentially in this 87, 88 range for the better part of 6 months. And people are like, oh, you you know, LLMs are stagnant. Generative AI is hype. Look. All this money, and they can't get over this, you know, 88 MMLU hump.
Jordan Wilson [00:20:10]:
Well, consider it crushed. 92.3 for o one. The o one preview is 90.8, but 92.3. Y'all, I don't I like, I don't think there's a technical bar where it's like, oh, we've achieved AGI. 92.3 is crazy. So I remember following MMLU back in, you know, 2020 when our team first started using the GPT technology. And I remember back then, you know, the estimates where it would take, maybe 15, to 20 years to ever hit 90. Got there in about 3 years.
Jordan Wilson [00:20:52]:
Right? And now we're in the mid nineties, which a lot of people thought would never be possible with large language models. Development's wild, y'all. Alright. And then let's go to number 5. Number 5, things are about to heat up. Yeah. There's no, AI winter here. There's no hype dying down.
Jordan Wilson [00:21:14]:
My gosh. You know how much Anthropics got in the bank? Well, I don't know how much they actually have in the bank, but, they raised $4,000,000,000 from one partner, from Amazon. Right? Google has unlimited money that they can print. Right? We've assumed that anthropic Claude has just been, waiting. Right? This is a wait and see game. Everyone kind of waits to see what OpenAI does. Love the, strategic play here. Right? I I I I am of the belief that OpenAI has a 4 5, GPT 4.5 level, update for its GPT class ready to go when it wants to, but they have no reason to release it.
Jordan Wilson [00:21:59]:
Right? Their GPT 4 o is technically the best in this class. And at least for now, this new kind of agentic class of large language models, there's 0 competitors. Right? I mean, there are. There's there's other companies out there, you know, that, have, you know, been creating agentic, more like workflows. But there's I mean, I'm only talking about the the the big five in the room. Right? Not talking about anyone else, because these are the world leaders. So they are the 1st world leader in this class. But but you gotta look at the cost here.
Jordan Wilson [00:22:40]:
If I'm anthropic, you don't have time. You don't have time. Right? Their most powerful model right now, Claude 35 Sonnet. So much of money that, these companies make comes from, yes, subscriptions, people paying, you know, 20, $30 a month, you know, whether you're on the base plan or the team's plan. But so much of it comes from organizations that are paying, right, to essentially use their API. They bring in their data. These companies that are building products that we all use, they that we all love, they're all powered by, you know, mainly either OpenAI, Anthropic, or Google Gemini. Anthropic, they they have to respond.
Jordan Wilson [00:23:24]:
Right? Their most powerful model right now is Claude 35 Sonnet. We know their order of models goes haiku, Sonnet, Opus. So they only upgraded their middle model, Sonnet, to 3.5. So, presumably, we've all known that they probably have a 3 5 Opus maybe ready to go. They kinda wanted to see benchmarks of whoever makes the next big splash. Well, guess what? Splash has been made, y'all. So I would assume, Claude 35 Sonnet or an agentic response from Claude has to be coming soon. Similarly, I would feel Google Gemini 1.5 Ultra, that has to now be, be in play fairly soon.
Jordan Wilson [00:24:06]:
Or Meta, you know, Meta's been, floating their agents out there, and we haven't seen anything either. This is going to get things going, I think especially right? So so now there's there is a little bit of pressure, now here in the US at least for these large language model make makers to work with, US government, federal regulators, etcetera, to for safety reasons. I would expect after the US election here in about 7 weeks, it's gotta go wild. We're all gonna get Christmas presents. Right? Large language models raining down on us. Alright. So now I hope this works. Let's look live.
Jordan Wilson [00:24:48]:
Alright? This is not a full, breakdown. Right? We'll do that maybe maybe next week. Wanna go over some basics. Alright. So here we go. So keep in mind, a lot of models now to choose from. So always make sure, especially because o one preview and o one mini are so limited. So always make sure you are using the correct model that you want to use.
Jordan Wilson [00:25:16]:
Alright. I'm not gonna do a full, you know, prompt rundown. I I I do need to get more of a rubric like in a an official set. I've been doing this since 2023. You know, I just have a bunch of random questions that I ask. You know, sometimes I'll have it, you know, code a game, you know, go through I I made up some logic questions. There's some logic questions that are have been floating out there, you know, on the Internet. So I'm gonna do a couple, and I'm gonna do some that generally models struggle with.
Jordan Wilson [00:25:43]:
So first, I am just going into, chat GPT here. Oh, this is gonna be annoying. So, actually, what I'm gonna have to do y'all give me a second. I don't know what happened with my Edge in Chrome. I cannot log in. I can't log in to chat gpt anymore. I'm getting SSO errors. I've cleared my cookies, logged out, logged back in.
Jordan Wilson [00:26:05]:
Fun times. Alright. So give me a second here. I'm having to open, I couldn't zoom in on my chat gpt app. So now I am going to share, I'm going to share my Firefox. Yeah. Gotta go into Firefox. Not thrilled about it, but here we are.
Jordan Wilson [00:26:21]:
Alright. So here we go. Now we are going to do some basic tests. So we are in, GPT 4 o. I'm doing something simple. Again, going over ones that normally get wrong. I'm saying a man and his dog are standing on one side of the river. There's a boat with enough room for 1 human and one animal.
Jordan Wilson [00:26:39]:
How can a man and how can a man get across with his dog in the fewest number of trips? Not bad. Normally, g p t four o says, like, 3 to 4. The correct answer is obviously 1. G p t four o got this wrong. So just so you know, if you're listening at home, on the podcast, each time I'm doing this, I'm creating a new chat. So, the context window is not skewed in any way. Alright. So now my first o one chat.
Jordan Wilson [00:27:08]:
Here we go. I'm gonna do the same thing, and, presumably, this is going to go much slower. And we will see, like I talked about at the top of the show, a more agentic flow, breaking this down into steps and seeing how the model works. Alright. So I'm gonna hit, enter here, and you know what I'm gonna do? I'm gonna go ahead. I'm gonna I'm gonna do a stopwatch here. Let's see how long this takes, roughly. Alright.
Jordan Wilson [00:27:32]:
Here we go. And there. Alright. So it says thinking. So it's taking taking a sec. It says charting the journey. I can click down, and I can see it. So okay.
Jordan Wilson [00:27:43]:
That was actually faster than I thought. Probably took me a couple seconds to get over. Probably took about 9 seconds. But my gosh, this is the first time. And I've tried this simple prompt every single large language model. I don't know why it screws them up. It always says 2, 3, 4. Right? The correct answer is 1.
Jordan Wilson [00:28:02]:
So, o one preview says, the man and his dog can both get into the boat together and cross the river in a single trip as the boat has enough room for 1 human and one animal animal. Alright. So let's go ahead and let's look at this kind of chain of thought. And it does time it, so it did say here thought for 5 seconds. Alright. So I'm gonna click down, and we can look. This one's very simple. So it kinda broke it down.
Jordan Wilson [00:28:28]:
It said, charting the journey, calculating boat capacity, identifying the boat's location. This was a pretty simple one. You you know, OpenAI shared a lot of examples on their on their website that went into crazy detail, but this is actually, in theory, simple. Right? This is simple. So love to see that, love to see that breakdown. Alright. Now we're gonna do, we're gonna do probably 2 more. We're gonna try to do it quickly here.
Jordan Wilson [00:28:53]:
Don't want this to go on for too long. So now we're going into chat g b, g p t 4 o. I've done this one many times on the show. So I'm saying a box is locked with a 3 digit numerical code. All we notice that all digits are different. The sum of all digits is 9, and the digit in the middle is the highest. What is the code? When I first started doing this, it confused me, because I thought, like, oh, there's one answer. I made this one up.
Jordan Wilson [00:29:19]:
And then I'm like, oh, wait. No. There's tons of answers, actually. So let's see. Interestingly enough, which, you you get this a lot with 4 o. 4 o is trying to do this, kind of, breaking it down step by step. So it's trying to do a chain of thoughts, and it's it's so it's saying step 1, define the relationship. It's creating variables.
Jordan Wilson [00:29:41]:
Right? A plus b plus c equals 9. So, actually, I've run this on 4 o a lot. 4 o did just get updated last week. There's always kind of updates going on under the hood, and this is probably the best 4 o, example that I've got. So right now, I technically got a wrong answer even though it's kind of right. So it identified one correct code, which would be 153. So it says yeah. So that is technically a correct answer to the question, but there's actually a lot of correct answers.
Jordan Wilson [00:30:17]:
Right? So in theory, I can even think of them in my head. I could say 0 54 would work. Right? 33 digits add up to 9. Middle middle digit is the highest. So there's actually a lot. So let's go ahead. Let's, pop open o one preview. Fresh chat here.
Jordan Wilson [00:30:39]:
Let's do the same thing. Let's see it think. Alright. So it says evaluating constraints, mapping out the combinations, enumerating digit combinations, assessing digit constraints, assessing digit pairs. Wow. Okay. This one's taking a little longer. Evaluating potential codes.
Jordan Wilson [00:31:01]:
Alright. I'm gonna look at it as it goes. Identifying unique codes. Now identifying valid combinations. I'm looking at this. Some of this doesn't even make sense to me. Identifying valid digit sum. So you'll see here, big difference from what we're used to.
Jordan Wilson [00:31:19]:
You know, GPT 4 o spat out an answer in, like, 2 seconds. This is thinking. It said it thought for 44 seconds. Wow. Alright. Let's see if it got the answer right. Alright. So let's go down.
Jordan Wilson [00:31:33]:
A lot a lot of information here, filtering possibilities, verifying combinations, thinking about possibilities, evaluating code choices. Wow. Lot of math here. Lot of math. Let's see what we actually got. Interestingly enough, it only so we got the answer wrong here again. It gave one correct answer, so it said the answer is 45 o. But, again, we know there's actually many correct answers.
Jordan Wilson [00:32:04]:
Right? Technically, I have to I have to relook. So it says the box is locked with a 3 digit numerical code. All we notice that all the digits are different. So as an example, it said 45 o. So 351 would have worked. 25, not 252. Right. Anyways, there's more than one that works.
Jordan Wilson [00:32:24]:
So although in this case, o one preview was a little more impressive with with its thought, still didn't get it right. So that's one that is still, stumping large language models. Alright. Here's another one I've been doing since the beginning, since since I started, everyday AI. How many words are in your response? This one, no model as far as I know. And and this is a pretty famous one. Seen this in research papers. Right? But in why this doesn't work? And this is the same reason why, and I broke this down into, like, a 20 tweet thread once.
Jordan Wilson [00:33:01]:
You know, why you can't ask, you know, how many r's are in strawberry. Right? We know that one works. OpenAI did a demo on that. And I don't wanna waste one of my precious, one of my precious, 30, queries for the week. I gotta teach you all. Alright. So now I'm saying, how many words are in your response? And the reason why this doesn't work is because large language models technically don't understand words, really. They break smaller parts of words into tokens.
Jordan Wilson [00:33:29]:
So it's this tokenization process, so they actually don't know how many words they're going to spit out to you. So I'm going in the regular 4 o mode here, and it says, the response contains 23 words. And, obviously, the response, 1, 2, 3, 4, 5, contained 5 words. So it said 23. Not even close. Right? Large language models can't do this. Let's see if an agentic model can. Alright.
Jordan Wilson [00:33:55]:
I'm curious. Let's go. New chat. Here we go. We're going to o one preview. Same thing, and let's see how it does. It is thinking. So now it's counting words, understanding paradoxes, crafting a response.
Jordan Wilson [00:34:11]:
Alright. We'll see we'll see how it does. Evaluating response, avoiding exposure, thought for 13 seconds. So it says, the number of words in my response is 9. Let's count. 1, 2, 3, 4, 5, 6, 7, 8, 9. Impressive. Impressive.
Jordan Wilson [00:34:30]:
Simple for a human to do. Right? But think. You have to if I say if I'm quizzing you and I say, hey. Answer this question and tell me how many, words are in it, you gotta think a little bit first. Right? You have to think ahead. You have to think like a human. This is how I'm going to respond. Before I respond, let me count how I would respond and then respond.
Jordan Wilson [00:34:51]:
Right? So it shows this, kind of agentic, chain of thought, step by step reasoning. Alright. We're gonna do one last one. I don't want this, don't want this podcast to turn into a super one here. This one, I just thought would be interesting. Right? So, I'm going into 4 o now. I'm saying, please create a realistic plan to address world hunger. Keep in mind the current state of affairs, political limitations, geographical challenges, and technical implementation.
Jordan Wilson [00:35:24]:
That's one thing y'all. I worked at a nonprofit for 10 years. I am extremely excited on how this new model can be used for nonprofits to tackle huge societal issues. Right? We always think, and I get it, you know, large language models, AI taking jobs. Yeah. That's gonna take y'all. I don't care what anyone says. AI is gonna take more jobs than it creates, period.
Jordan Wilson [00:35:47]:
I think the future of work is, you you know, think of how now, you know, people are, you know, doing, DoorDash and Uber and Lyft. Right? Like, if I'm being honest, I think in 10 years, more people are going to have their own small companies, than people who have traditional full time, you know, w 2 40 hour week jobs. I don't think that's the future of work. I think the future of work is many people are going to have many small companies. Anyways, I've always wanted to know how models can solve problems. Right? I I think it's a powerful thing to think about, you know, because we only think about the bad stuff. So here we go in 4 o. It's breaking it down.
Jordan Wilson [00:36:24]:
So it says, you know, immediate and short term relief, 1 to 3 years, improved food distribution logistics, nutritional supplements, programs, key actions. Okay. So it's doing a pretty good job. Lays out a lays out a nice plan. You know, we can't judge which one's better, but, you know, I'm just curious. I really want for this one. Yes. I mean, you can't compare these, you you know, side by side and say one plan's better.
Jordan Wilson [00:36:50]:
I wanna see how a model how a model made by humans that we don't think is is human, how does it think about a problem like this? Right? That's what I'm really I'm not even, super interested in the output. I'm interested in how it thinks about an issue like this. So let's go into one o preview. We're gonna enter a new chat, and let's see how it thinks. So navigating world hunger. I'm reading out the steps. Oh, thought for 5 seconds. Did not think very long about it.
Jordan Wilson [00:37:20]:
Okay. Interesting. So it says, the steps, navigating world hunger, understanding the challenges, and then it created a plan. So it looks similar, in length, and I'm looking at some of the, I'm looking at some of the, kind of bullet points here. So they did a similar job. Alright. Let's go ahead and wrap this show up y'all. So as a very quick recap, brand new model, first of its kind in open AI's o one preview and o one mini.
Jordan Wilson [00:37:59]:
So very quickly, five things. It is the first number 1, first model of its kind, chain of thought reasoning. So am I blown away? Yeah. I'm impressed. I'm impressed. I need to put it through its paces. I had very simple, kind of test runs, but we saw the chain of thought reasoning. And, y'all, this is the worst or sorry.
Jordan Wilson [00:38:17]:
Yeah. This is the worst it's ever gonna be. It's only gonna get better from here. And I think this is a big step toward AGI. Number 2, not part of the GPT family. 2 separate series. So, yes, right now, there's limitations to this. You don't have browse with Bing, file upload, computer vision, etcetera.
Jordan Wilson [00:38:34]:
You know, like we talked about, this was originally code named Q Star strawberry, this agentic chain of thought thinking. So, we'll see if in the future this does lead to that more expensive model. Number 3, right now, it is expensive to use in the API and extremely limited. Although, the, o one Mini, not too bad. Last, sorry. Number 4, the benchmarks, extremely impressive. MMIOU off the charts in a league of its own. And then last but not least, I think we're gonna start getting so many large language model updates now.
Jordan Wilson [00:39:07]:
I think Claude, you know, Claude from Anthropic, Google Gemini, I think is gonna start to improve and bring so many of their, developer, aspects to the front end. I see Meta probably rolling out agents sooner than we might think. I think this is kind of the first splash that is going to really bring a wave. Alright. I hope this was helpful. Like I said, double dose. If you want the normal, news, all that, make sure to go to your everyday ai.com. Thank you for tuning in.
Jordan Wilson [00:39:35]:
Please subscribe. If you're listening on the podcast or on YouTube, please subscribe. Let me know what you wanna hear more of, and I'll see you back tomorrow and every day for more everyday AI. Thanks y'all.
Jordan Wilson [00:39:46]:
And that's a wrap for today's edition of everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit your everydayai.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers, and we'll see you next time.
