Ep 456: OpenAI’s o3-Mini – The world’s best free chatbot model?

Resources:

Join the discussion: Ask Jordan questions on OpenAI


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


OpenAI’s o3-mini Model: Why It May Be the Best Free AI Chatbot Yet

From coding tasks and logic puzzles to marketing strategies and brand ideation, large language models (LLMs) are reshaping how we work and think. Yet, for most professionals and hobbyists alike, free versions have often underwhelmed—until now. In a recent discussion on the Everyday AI podcast, host Jordan Wilson walked through OpenAI’s new o3-mini model, demonstrating its reasoning prowess and internet-connected capabilities. Below is a deep dive into how o3-mini works, why it may currently stand as the best free AI chatbot available, and what to expect if you decide to incorporate it into your own workflow.


A Quick Introduction to OpenAI’s o-Series Models

Traditionally, OpenAI’s GPT-based models (like GPT-4 or GPT-3.5) have been known for their transformer architecture: they excel at language prediction and text generation. However, the o-series of models—such as o1 and, more recently, o3—take a different path by focusing on reasoning. Rather than simply predicting words in a sequence, these models conduct “chain-of-thought” thinking behind the scenes, making them more adept at logic-oriented or multi-step tasks.


Why the “Mini” Matters

The label “mini” may sound like a downgrade, but o3-mini offers some of the most advanced reasoning power on the market—especially within the context of free usage. Paid subscriptions of ChatGPT or enterprise-level licensing get you unlimited or higher-capacity versions of o3-mini, but even the strictly free tier showcases remarkable performance, outpacing older free LLMs. This effectively closes the gap that once existed between free users and those paying for top-tier chatbots.


How o3-mini Compares to Other Free AI Chatbots

In a world where ChatGPT, Claude, and even Google’s Gemini each jostle for attention, it can be challenging to pinpoint the best free tool. A quick overview:

  1. Google’s Gemini (Free Version)
    Google’s Gemini 2.0 models offer large context windows and impressive reasoning, but the freely accessible variants typically lag behind paid tiers. Most of Gemini’s advanced features (like multi-million-token context windows) live behind enterprise accounts or AI Studio for developers.
  2. Anthropic’s Claude
    While Claude often excels at creative writing and summarizing text, free users encounter strict rate limits. On top of that, it lacks consistent web connectivity and large context windows in the no-cost plan, hampering its utility for real-world research tasks.
  3. DeepSeek
    DeepSeek’s r1 model has gained traction for strong benchmarks and quick results. However, reports of hidden code sending data to state-owned companies—and pending bans in some countries—raise serious concerns about data privacy and national security.
  4. GPT 3.5 (Legacy Free Version)
    This older free version of ChatGPT excels at everyday text generation but lacks advanced reasoning features and out-of-the-box web access. Users often run into hallucinations or low-quality summaries because GPT 3.5 is neither well-tuned for logic nor up-to-date information.

Given these constraints, o3-mini emerges as the best of the bunch for those on a budget. Not only does it use more sophisticated reasoning, but it also connects to the internet, allowing up-to-date information queries—a feature many free chatbots still lack.


Key Features and Capabilities


Internet Connectivity

Unlike older free versions of ChatGPT, o3-mini includes an option to “search the web.” Once enabled or prompted, it can (in many cases) scour the internet for the latest headlines, sports results, or stock data. According to the conversation on Everyday AI, this access doesn’t always kick in automatically; you may need to explicitly tell the model to conduct a web search. But having the ability to retrieve real-time information puts o3-mini well ahead of many competing free LLMs.


Chain-of-Thought Reasoning

As a “reasoner,” o3-mini attempts a multi-step approach behind the scenes, summarizing logic paths into a simplified chain of thought. In real-time demos, it tackled trick questions such as:

  • “If it takes three hours to dry 10 T-shirts, how long to dry 30 T-shirts?”
  • “A man and his dog are standing on one side of a river with a small boat…”
  • “What is a locked code if the middle digit must be the highest?”

Where older free chatbots floundered, o3-mini stepped through logical interpretations—sometimes a multi-step process of elimination—before arriving at the correct or more nuanced answer. This advanced reasoning is exactly why many professionals are moving toward the o-series for complex tasks.


Benchmark Standouts

Reports from third-party benchmarking services place o3-mini among top-performing models in:

  • STEM Encoding: It handles math tasks with fewer errors and less “hallucination.”
  • PhD-Level Science Q&A: Outperforms many other free models, including vintage GPT-3.5 variants.
  • Complex Problem-Solving: Scores surprisingly high on specialized logic tests, further highlighting its reasoning advantage.

Short Demo Use Cases

During the podcast, Jordan Wilson presented a series of prompts to o3-mini, revealing its strengths:

  1. Fruit Riddles and Math
    A prompt about starting a day with six apples and three bananas, accounting for fruit eaten on different days, and forecasting how many apples/bananas remain by evening might stump older models. But o3-mini answered accurately, providing its step-by-step thinking.
  2. Joke Generation
    Though humor is subjective, the model was asked to create multiple jokes all ending in the word “blue.” While the results were corny, it successfully satisfied the constraints—showing how well it follows complex instructions, even for creative tasks.
  3. Marketing Strategies
    When asked to propose marketing strategies for the Everyday AI podcast itself, o3-mini offered new ideas involving challenges, AI ambassadors, and hyper-personalized social ads. Although not every tip was groundbreaking, the results were specific and aligned to the show’s brand, especially when internet searching was enabled.
  4. Brand Ideation
    Prompted to invent a company and its flagship product that solves a “non-existent” problem, o3-mini devised a hypothetical brand (complete with a product name, tagline, and campaign). This showcases the model’s adeptness at open-ended tasks where creativity and logic intersect.

Why Businesses Should Take Note


Uplift for Free Users

For many professionals or small organizations hesitant to invest in a paid ChatGPT or other AI subscriptions, o3-mini removes a major barrier. Despite message limits, it delivers a taste of advanced reasoning at zero cost. This effectively democratizes AI access—allowing teams to experiment with test projects, data extraction, or marketing brainstorms.


Improved Consistency and Accuracy

Prior free models were notorious for spewing out hallucinations or contradictory statements. By contrast, o3-mini’s chain-of-thought system mitigates these issues, often verifying logic and double-checking data (especially when connected to the web). While no model is 100% foolproof, early demonstrations indicate a marked improvement in consistency.


Fair Pricing for API Usage

The o3-mini model is also available for API usage with a pay-as-you-go structure. According to the podcast discussion, the model provides one of the best ratios of price to performance, particularly in coding and STEM tasks. For companies looking to integrate AI into their apps or workflows, o3-mini’s API costs are significantly lower than many competitor offerings at similar or lower performance levels.


Potential Pitfalls and Limitations

Despite glowing reviews, o3-mini does carry a few caveats:

  1. Daily/Weekly Message Caps
    Free users can only submit a limited number of queries. High-volume users may find themselves hitting limits quickly.
  2. Hit-or-Miss Web Searches
    While it can search the web, sometimes it won’t unless explicitly instructed. Moreover, the summarized chain-of-thought typically does not show every website visited, making fact-checking more cumbersome for certain tasks.
  3. No File Uploads on o3-mini
    For advanced features like image uploads or PDF analysis, the older o1 model still has unique features like “canvas mode.” If your workflow revolves around multimodal inputs, you may have to switch between models or pay for more advanced o-series subscriptions.

Getting the Most Out of o3-mini

  • Use Clear Prompts
    Large language models thrive on well-structured instructions. If you want o3-mini to browse the web, specifically request “Please use ChatGPT search…” or “Browse the internet for the latest data.”
  • Explore Industry-Specific Tasks
    From financial forecasting to healthcare research, targeted prompts can uncover how robust o3-mini’s chain-of-thought reasoning can be. If an initial answer seems generic, refine your query with more details or context.
  • Review the Summaries
    After each answer, you can glance at o3-mini’s summarized chain-of-thought headings. This isn’t the full reasoning transcript, but it provides insight into why the AI arrived at certain conclusions.
  • Switch Models When Needed
    If you require image uploads, you might need to toggle back to the o1 model. Conversely, if you value raw power and you’re on a paid plan, consider o1 Pro for especially complex tasks. Each model under the ChatGPT umbrella has a distinct edge, and you can seamlessly switch between them.

Conclusion

OpenAI’s o3-mini marks a significant leap forward for free AI chatbots, blending advanced reasoning capabilities with at-will internet access—an unprecedented combination in the no-cost tier. For individuals and businesses wanting a reliable, clever assistant without immediate subscriptions, this new model offers a snapshot of cutting-edge performance once reserved only for paid subscribers. Whether you’re solving logic puzzles, brainstorming product launches, or simply exploring the potential of AI, o3-mini stands as a compelling starting point.

Of course, more ambitious deployments—like system-wide automations or enterprise applications—may still require paid plans. Yet the conversation on Everyday AI makes it clear: if your goal is to see what advanced reasoning looks like at zero cost, o3-mini is your best bet. By mastering how to prompt effectively, leveraging web search, and paying attention to its chain-of-thought logic, you’ll find this “mini” model delivers quite a massive impact.


Topics Covered in This Episode

1. Overview of OpenAI’s 03-mini Model
2. Comparisons to Other Chatbot Models
3. o3-Mini High’s Reasoning Abilities
4. o3-Mini Key Takeaways & Implications


Podcast Transcript


Jordan Wilson [00:00:16]:
I know I've said probably dozens of times, don't use free AI models. Right? Because the paid versions are ridiculously cheap for what you get. $20 a month, $30 a month. It doesn't matter whether you're an individual or buying that for, your your company with thousands of employees. That's so affordable. So I've always said don't touch the free models, but I might have to change that because OpenAI has made their o3 mini model free, well, in a very limited capacity in terms of the number of messages yet. I think it may be the world's best free chatbot model. So we're gonna be going over today the new o3 mini model from OpenAI.

Jordan Wilson [00:01:07]:
Talk about what it is, how it works, its implications. If it is really the best free AI, chatbot model in the world, and maybe do a little bit of live testing. Alright. I'm excited for this one. I hope you are too. If you're new here, welcome. This is Everyday AI. My name is Jordan Wilson, and we do this thing every day.

Jordan Wilson [00:01:27]:
It's for you. This is your daily livestream podcast and free daily newsletter helping us all not just keep up with AI, but how we can use it to get ahead to grow our company and to grow our careers. And that is a full time job if you are not tuning in every day. So we do all that hard work for you so then you can go be the smartest person in AI at your company. Alright. So, if you are new, maybe listening for the first time on the live stream or the podcast, thank you for tuning in. Make sure to check your show notes. Very important things.

Jordan Wilson [00:01:55]:
They they're, mainly our website, youreverydayai.com. That's where you're gonna wanna go sign up for our free daily newsletter because every single day, we do a couple of things. We bring you all the latest AI news and tell you what it means, but we also break down our podcast episode from the day with some, some more information. So make sure you go do that as well as while you're there, I'm gonna keep promoting this y'all. You need to go listen to our twenty twenty five AI predictions and road map series. It is all on our website. It is for free. I'm getting a ton of messages.

Jordan Wilson [00:02:27]:
I just got a message actually, today or last night because I think this person's in Europe from one of the largest consulting companies in the world, and they said their team is breaking down those five episodes, and they're gonna continue to track them all year. I kid you not. You need to go listen to them, and let me know what you think. Alright. Enough chitchat. Let's get into the AI news. Live stream audience, this is up to you. I have a question on the screen there.

Jordan Wilson [00:02:53]:
Let me know. We're gonna be doing an o3, mini live test. Do you wanna see a, do you wanna see it go through a reasoning rubric, or do you wanna see b go through some real world data analysis? So let me know a or b on the screen. Let me know now. Alright. So AI news. A lot going on. Gemini.

Jordan Wilson [00:03:09]:
Google has announced the general availability of its Gemini two point o flash model, a high performance AI designed for developers with enhanced speed and complex problem problem solving capabilities. So the Gemini two point o flash model was first introduced at IO twenty twenty four. It's, developer conference is praised for its efficiency in handling high volume tasks and multimodal reasoning with a 1,000,000 token context window. Also, a new experimental version of the big boy model, Gemini two point o pro, is also available, for paid users boasting superior coding performance in a 2,000,000, token context window. I mean, Google is just dominating the context window game. They also introduced Gemini two point o Flashlight, the most cost efficient model to date, and that is in public preview, and that's offering improved quality over its predecessor 1.5 flash. So Google emphasized the importance of safety and responsibility with the Gemini two point o lineup, incorporating new reinforcement learning techniques and automated red teaming to mitigate risk and ensure secure usage. Yeah.

Jordan Wilson [00:04:19]:
Huge huge news there from Google. I'm excited to dive into that a lot more. Our next piece of AI news. US Lawmakers are proposing a ban on DeepSeek. Big surprise. Not at all. Alright. So lawmakers in The US are planning to introduce a bill to ban DeepSeek's chatbot application from, right now, just government owned devices over security concerns that user data could be accessed by the Chinese government.

Jordan Wilson [00:04:48]:
Spoiler alert, it can. The bipartisan legislation, is echoing previous efforts to ban TikTok from government devices, which was the precursor to, TikTok being banned in The US, which was slight, quickly there kind of overturned, but it could still happen. Alright. So DeepSeek is a Chinese AI company, and they've rapidly gained popularity in The US, becoming the most downloaded iOS app last month. Concerns arose though after an analysis revealed that there's hidden code in the app that could send information, user information to China Mobile, a state owned company banned in The US, a China state owned company. So, yeah, the proposed legislation aims to ban sensitive government and personal data from being accessed by the Chinese Communist Party. Other countries, including Australia, South Korea, and Italy have already banned DeepSeek from their government systems due to similar data security concerns. Also, some US federal agencies such as the Navy and NASA have preemptively blocked the app for security reasons.

Jordan Wilson [00:05:54]:
I I left a post on LinkedIn for this. I'll probably do a show dedicated at some point. This story is changing so quickly. That's why I've kind of held my tongue on this because I got a lot of hot takes. So I might need to save that for, a week or two. Alright. Last but not least, new piece of AI news. ChatGPT has made speaking of things for free, their new ChatGPT search available now for free even for free users who are not logged in.

Jordan Wilson [00:06:19]:
So this new feature, well, no feature at least new to free and non logged in users, allows everyone to access up to the date information such as sports scores, news, and stock prices directly through ChatGPT. So, according to reports, the search functionality uses a fine tuned version of g p t four o optimized with synthetic data and output from OpenAI's new reasoning models. So OpenAI has partnered with major news organizations like the Associated Press and Reuters for licensing agreements influencing the visibility of certain publishers in search results. So this is huge. It's kind of weird now. Right? It almost looks like Google, is is trying to really compete with, what ChatGPT was, like, two years ago. And now ChatGPT is, trying to compete with what Google has been for the past twenty years. Right? Really making a, stake to try to erase Google.

Jordan Wilson [00:07:14]:
Right? OpenAI just wants you to skip Google altogether and go use its ChatGPT search, and you don't even need to be logged in, and you don't even need to have an account. So pretty wild. Alright. That's enough chitchat. Lots more in the, in the newsletter today. So let's talk about OpenAI's o3 minutei model. It's very impressive. It's very impressive.

Jordan Wilson [00:07:37]:
I'm just gonna say that. And, hey, livestream audience. Thank you for tuning in. I saw a couple votes, you know, a some a's, some b's. So, yeah, let me know if if you wanna go over the the reasoning rubric live or if you wanna go over the data analysis. So a or b. Alright. Also, you're gonna wanna repost this.

Jordan Wilson [00:07:58]:
All I'm saying, I've been putting these guides together, after I do something in ChatGPT or other large language models. I'm like, wait. I just really saved people dozens of hours a week if they go do this. And I realized that, you know, sometimes the podcast isn't enough for putting it in the newsletter. So I did put together another guide specifically on using the o3 minutei model. It's fantastic. I just finished it this morning. So, if you repost the show, I'm gonna send that to you.

Jordan Wilson [00:08:29]:
Alright. So let's get into the o3 minutei model. Here's the gist. Okay? It is the first free reasoning model from OpenAI. Right? So we have the o series of models. It is different than the GPTs. Right? So the GPTs are your, you you know, quote, unquote old school transformer models. And then the o series, these this is, OpenAI's reasoner models.

Jordan Wilson [00:08:52]:
Right? The reasoner models have become very popular over the last, like, four or five months. But this is essentially a model that thinks longer, kinda does more inference or, you know, uses kind of this chain of thought thinking, where it doesn't just quickly respond to something. It takes a while and and really kind of thinks internally. So kind of the work that you would normally do in a transformer model, a a GPT four o. Right? As a human, you'd wanna go back and forth with it a lot. Kind of these reasoning models, that's why they're so good. But a couple of things. It uses more compute, so generally, they're more costly.

Jordan Wilson [00:09:24]:
So as an example, if you want to have unlimited use of this, you need the $200 pro plan, but at least for, I believe it's 10 messages until you hit your message cap. It's available right now for free users. But it's not just that that makes me excited about this. So if you are logged in, so this isn't this is separate news from the chat g b t search, that you can use even if you are not logged in. Right? So if you do have a free even a free chat g b t account, so not only do you have a couple, you you know, queries that you can use with the new o3 minutei, model that OpenAI just released, but it also is connected to the Internet. That is huge. That is the piece that I think most people have missed or overlooked when it comes to, when it comes to, this this new model from OpenAI. And and that guide, by the way, the guide leverages specifically.

Jordan Wilson [00:10:19]:
It's 20 different use cases that combine reasoning, and the Internet. Right? Which is that's what knowledge workers do, and that's why I think this is so exciting, even for people who are not paid subscribers. Again, whether your favorite chatbot is is Gemini, Claw, ChatGPT, whatever. Just pay for the base $20 a month plan. It pays for itself the first time you hit enter if you know what you're doing. But even for those cheapskates out there. Right? Yeah. I know a couple of you out there that are still pinching your pennies and even though you're buying $8 coffees every day, you're like, oh, I'm not gonna buy a $20.

Jordan Wilson [00:10:54]:
No. Just buy it. But still, this is the first model for free, that you can use from OpenAI that is its reasoning model, and it's connected to the Internet. And reportedly, I don't personally believe this, but a lot of people have said, oh, OpenAI did this because of the deep seek r one release, which took the Internet by storm for a lot of, I'm gonna say, incorrect reasons. We'll just say that. Alright? I don't personally think this is in response, to DeepSeek. I think this is actually in response, to Google that's been on a freaking tear, since December. Google has been straight up releasing a crazy amount of releases.

Jordan Wilson [00:11:36]:
If I'm OpenAI, I I don't care about DeepSeek. Right? It's largely going to get banned, I believe. I'm worried about Google, so I think that this is really a shot at all of the great work that Google's been doing specifically in Google AI Studio. Alright. So let's just answer that question right now. I'm not gonna make you wait another twenty minutes and after our test. Is o3 minutei the best free chatbot model? Let me break that down. Free chatbot model.

Jordan Wilson [00:12:04]:
Okay. That's when you log in to the front end of a chatbot. Alright. So what do I mean by that? Well, right now, if you have a free account and you log in to Gemini, you know, gemini.google.com, even though they had all these new releases, you can't use them. You're only using 1.5 Flash. But they have great models within the AI studio, but it's a little different. That's not a kind of a for beginners. That's more for developers.

Jordan Wilson [00:12:32]:
So it beats Google. No questions asked. Copilot is powered by GPT four o technology. And last week, if you read our newsletter, you're smart. You already know this. There is, some limited free access to OpenAI's o one model with the think deeper, kind of capabilities inside Copilot. But I still think o3 Mini is better because we're talking o one. I believe that one's o one preview.

Jordan Wilson [00:13:00]:
Microsoft, I'd like, I know a lot of you guys listen to this and I always tell, like, I met, like, talked with, like, a hundred of you guys, at the build or, the Ignite conference here in Chicago, and I'm like, hey. Tell me if I'm wrong on this. But I'm pretty sure the Think Deeper uses the o one, preview, not the o one pro. So I still think it's better than that. I still think o3 minutei for free is better than using Copilot for free. Claude, just LOL that. I mean, by the time even on a paid account, like, Claude, you can't use it. You you just can't.

Jordan Wilson [00:13:34]:
Right? Anything more than a couple of prompts and you hit your rate limits. On free on a free account, even though Claude three five Sonnet is a good model, it's, you know, now, like, I don't know, eight months old. So, presumably, we'll be seeing new updates from Claude pretty time soon. I mean, on a free plan, if you look at Claude the wrong way, you've already hit your message limit. Alright? So you can't really use it a lot. Alright. And then deep seek, good luck with that. High risk, lots of questions, great model, great benchmarks.

Jordan Wilson [00:13:59]:
Right? Good luck. That's all I'll say. Alright. So is o3 Mini the best free chatbot model? Yes. It's not even close. This model is so good. Here's the thing. It's limited.

Jordan Wilson [00:14:14]:
If you're on the free plan, I think it's only, I think it's only 10. It's either 10 a day or 10 a week. OpenAI doesn't say, and all of my accounts are paid. So I was I was trying to, you you know, quickly find that answer out. I'll I'll make sure to put it in the newsletter. But, yes, it is. It is. And I actually don't think it's even close.

Jordan Wilson [00:14:34]:
Alright? Hey. Someone, someone here is saying our audio is cutting in and out a bit. Let me know if it actually is or if if, maybe that person has some, computer problems today. So it is. OpenAI's new o3 Mini is the best free chatbot model in the world, and I don't think it's necessarily close. Because it's not just the model, it's everything else that the model has capabilities to do. Like we said, search just right there. Chatt GPT search is great.

Jordan Wilson [00:15:09]:
Alright. So let's go over some of the highlights of the model. So it excels in stem encoding, this new o3 minutei. Alright. It is 63% cheaper than o one minutei, which is the model it replaced. So, yeah, if you were FYI, if you're on a paid plan and and you're in there looking and you're like, wait. Where's o one mini? O one mini is gone, and now you have o3 mini instead. And there's actually multiple variations of o3 mini.

Jordan Wilson [00:15:35]:
I'll get to that in a second. It is o3 Mini is 24% faster than o one Mini. And for API users, there's actually three variations. There's a low, a medium, and a high, kind of variety or flavor. And that's essentially your your choosing speed and cost versus performance. So, for the o3 Mini low, that is going to be the cheapest, the fastest with the lowest performance. o3 minutei high is going to be the most expensive and take a little longer, but it's going to have the best performance, obviously. And then the, o3 o3 minutei, normal is going to be that that, middle.

Jordan Wilson [00:16:18]:
Right? It's like the what is it? The the the three beds. Right? One's one's too soft, one's too hard, one's just right. Alright. And for that's for API. So if you are on the chatbot version, right, which is many of us. Right? So chatGPT.com, you're not, you know, using the back end API as a developer, but for ChatGPT users, if you are ChatGPT plus, so the $20 a month, you have access to o3 mini, kind of the medium or middle version, and then o3-mini-high. And that one kind of thinks harder more or less. It uses a little more compute, and you have, I believe 50 oh, no.

Jordan Wilson [00:17:01]:
It's a hundred and 50 a day, I believe now. So plenty plenty of usage. They just tripled it, in the last couple of days. So if you have the $20 a month plan, I don't think you're probably gonna hit your o3-mini-high limits. And let me tell you, right now, o3-mini-high is probably one of my most used models. Alright. Also, o3-mini-high outperforms the full o one model on many benchmarks. Because right now, we just have the the miniature version.

Jordan Wilson [00:17:33]:
Right? We don't have the full o3 version. I don't even know if the full o3 version is gonna come out in 2025. I would assume it would, but I don't know, because the full o3 model has not been released. The only, glimpse of it that we've seen is OpenAI did say that it's new deep research, which is mind blowingly good. It's gonna put so many small to medium size and management consultant companies out of business. I'm not kidding. It's freaking good. Anyways, that uses a fine tuned version of the full o3 model, but this is just we're just getting the mini.

Jordan Wilson [00:18:11]:
We're just getting the mini one here, y'all. Alright. Let's keep it going. Benchmarks. I know. Not gonna get too dorky here, but, let's look at competition math. o3-mini-high outperforms even the full o one model on, the AIME. I think that's AIM, 2024 competition math benchmark.

Jordan Wilson [00:18:37]:
Right? So large language models, they go through all these tests, all these standardized tests, essentially. Think of, like, a human. You know, there's all these different tests you take. Same thing with models, and then you get benchmarks, you get scores. Right? So you can see how capable a model is. So while o3-mini-high is even more capable than the o one model, and that's, one of the highest scores in the world. And then you have PhD level science questions. Right? Because o3, o3-mini-high is great at anything stem coding research.

Jordan Wilson [00:19:06]:
It's a chef's kiss. Good. Alright. So on the, PhD level science, which is GPGPQA diamond for those of you at home keeping score, o3-mini-high also outbenches the full o one pro. Right? Also, this is not yet o3 Mini. So kind of the the the benchmarks that I talk about a lot on this show, aside from, you know, MMLU and some of those that I just mentioned, are, the the chatbot arena scores. Those aren't out yet because this model is, like, barely like a week old. Right? But we do have from artificial analysis, which is a great resource for, an unbiased third party model benchmarking service.

Jordan Wilson [00:19:54]:
O o3 Mini in terms of quality, it is second in the world only behind the full o one model. So o3 minutei and deep seek, r one are actually they're tied with scores of 89 where o one has a 90. For comparison, right, if you love Claude three five as an example, Claude three five has a 68 if that puts it on the on the scale for you. Alright. Gemini two point o, not their newest version, but the one previous to that had an 82. So what does that mean? It is by far one of the highest quality models in the world, and OpenAI made it free or a limited use case. Right? But it's it's mind boggling that that that we have this level of a reasoning model that is one of the most capable in the world, and it also can access the Internet, which is one of the reasons why I tell people don't use Claude at least yet. Right? Because there's a a somewhat of a business danger.

Jordan Wilson [00:21:00]:
If you are taking results from a large language model that has very old data. You shouldn't be doing it. Alright. So another, kind of graph here from artificial analysis. This just kinda shows your quality versus price, and that's where you see, oh, okay. When it comes to quality versus price, o3 Mini is actually the best in the world, and it's not necessarily close. The only one somewhat close is DeepSeek r one. Again, good luck with that if if if you wanna, you know, good luck with that if you wanna use it.

Jordan Wilson [00:21:35]:
I'm not using it on a day to day basis. But o3 Mini from a quality and price perspective, Right now, it can't be beat. I mean, we'll see. I think Google's announcements, yesterday are gonna shake this graph up a little bit, and I'm excited to to dive into all of the new Gemini two point o a little bit. But right now, o3 Mini is technically an elite model. Don't let the Mini confuse you. Alright? So as a reasoning model, these are the API pricing. Right? So, again, you can use it for free.

Jordan Wilson [00:22:09]:
You can use it chat g b t plus $20 a month. You're probably not gonna run out of queries. If you have the the pro version like I do $200 a month, it's unlimited. But, for API pricing, it's a dollar 10 for a million input token, and $4.40, for a million output tokens. For a reasoning model, so affordable. It is so affordable. And you might be confused. I get it.

Jordan Wilson [00:22:40]:
All this o alphabet soup. Right? OpenAI CEO, Sam Altman, did admit that they have a naming problem with the models. It's hard. Right? And especially when they come out with some of these new o reasoning models, some of the older ones get replaced or they're just no longer available. So let me just give you a quick rundown of the o series. So in September, we got o one preview and o one mini. Alright? Then in December, they got rid of o one preview, and then we just had, o one, and they added o one and o one pro. So if you had a pro account, that's the only way you can get access to pro.

Jordan Wilson [00:23:22]:
In December, you had three versions. You had o one mini, o one, and o one pro. Okay. Easy enough to follow along. But then January 31 came around last week. Right? And that threw a wrench in it. So now we went to o3. There is no o two because that's the, the trademark name of a British telecom company.

Jordan Wilson [00:23:41]:
So, if you're wondering, like, what happened here? Did I miss out on a whole series of AI development? No. You didn't. Right. But now in January, we got this o3 minutei, and that has o3 minutei high, and then o one minutei is gone. I know. Confusing. So, depending on what paid plan you have, you might still have in your account o one, o one Pro, o3 Mini, and o3-mini-high. I know it's confusing.

Jordan Wilson [00:24:11]:
I have a slide here that can hopefully help you make a little sense of it. Alright? Because which model should you use? Right? Like, if if you're like, oh, I have a paid ChatGPT account. What should I use? Well, there's actually some unique features of each. So listen in here. I have a helpful little graph on screen for our livestream audience. So o one, not the pro. Okay? O one actually has a great advantage to it. Okay? So right now, o one and o one pro are the only o models where you can upload files.

Jordan Wilson [00:24:48]:
Not all upload file types are supported. Alright. But it does have, for visuals, you know, PNGs and JPEGs, I believe. Alright. So that's the o one series. So normal o one can access canvas mode. Alright? O one pro cannot, yet o one pro is much more powerful than normal o one. Okay? So if you need to upload files, right, visuals at least because, you can't upload PDFs or spreadsheets right now into the o one models, but let's say you're doing a lot of visual, you know, computer vision type work.

Jordan Wilson [00:25:24]:
You're probably gonna wanna still choose one of the o o one models. Alright? If you love canvas like I do, you might use o one because that's the only one that has canvas. If you need the just straight up raw power, you're gonna wanna go with o one pro. Alright. But o ones don't have access to the Internet. So o3 minutei, there's no, there's no differentiation right now, between features or other tools within chat g b t. But o3 minutei is the only one that has web search. Alright? And that is the only mini model now.

Jordan Wilson [00:26:00]:
I know a little hard, but essentially if you need the web, which I highly advise, go o3 minutei. That's why I'm using o3 minutei a ton. Alright. If you need canvas, use normal o one. If you're on the big the big boy plan, then you can use o one pro for some of those very tough tasks. Is that does that make sense y'all? Hey. If you have questions, get them in now. Podcast audience, I love hearing from you guys.

Jordan Wilson [00:26:26]:
That's why I always put our our email in there. I put my LinkedIn. Reach out to me. Let me know, like, if this is helpful, if you have questions. I'm sometimes a little slow, getting around to those messages, but I do eventually. Alright. So let's look live. Let's see what won our little poll this morning.

Jordan Wilson [00:26:43]:
Let me count. So our a's let's see. We had we had 123, 4, 5, 6, 7, 8, 9, 10. Okay. 10. And then our b's. Let's see. We had 12, 3.

Jordan Wilson [00:26:58]:
Alright. Looks like you guys wanted the reasoning, the the the reasoning version here. Alright. So let's jump into it. Live stream audience. As always, please let me know when and if you can see my screen here. Alright. So we are going into chat g b t.

Jordan Wilson [00:27:15]:
I'm gonna do this live. Alright. So, these, let me make sure I go into o3-mini-high. So I'm gonna be using o3-mini-high for these. Alright. So this little reasoning rubric, I've been using a lot of these questions now for, like, two years. Right? Before there were reasoning models, I I had, like, kind of this common set of about 12 questions, that I would give to any models. Some of the earlier models, you know, Claude three five SONNET, GPT four, GPT four o, Gemini two, didn't do very good with this, because they're kinda like trick questions.

Jordan Wilson [00:27:53]:
But I actually think this is pretty important. Right? Because sometimes a simple mistake when using ChatGPT or Claude or, Gemini can screw up your entire output. Right? Because large language models, whether you know this or not, they don't understand words. Right? You give it a bunch of words, it doesn't understand it. When it spits backwards, it doesn't know what those words are. It converts everything into tokens. Alright? So sometimes, large language models get confused like humans do. Right? But that's important to keep in mind, but that's why I think this kind of, like, quote, unquote reasoning rubric is important.

Jordan Wilson [00:28:32]:
These aren't questions that you would generally use, right, on a day to day basis to grow your company and career. But this just shows you, are these models smart or not? Right? Alright. So let's go ahead and try our first question here. So, again, I'm using o3-mini-high. Alright. And you're gonna see these live. Hopefully, it's not gonna take too long, to go in there. So the first one I am saying, I just woke up with six apples and three bananas.

Jordan Wilson [00:29:03]:
If you're a long time listener, you've heard this before. I just woke up with today with six apples and three bananas. Yesterday, I ate a banana and two apples. This morning you know what? I'm gonna I'm gonna go ahead and scroll up here. I'm gonna scroll up here. Hey, livestream audience. Let's see if you can get this. Ready? I'm gonna I'm gonna go slow.

Jordan Wilson [00:29:21]:
I just woke up today with six apples and three bananas. Yesterday, I ate a banana and two apples. This morning, I will eat one apples one apple and no bananas. However, I don't really like apples, and one banana may turn brown tomorrow. Assuming nothing else changes, how many apples and bananas will I have tonight? Livestream audience. What's your guess on that? Podcast audience. You, are you scribbling this at home? This is a fun one. I actually made this one up.

Jordan Wilson [00:29:52]:
Some of these are very widely used, kind of, you know, trick questions or variations of these. Some of them I just made up. Right? So, I'm curious if our if our livestream audience, can get this one correct. But, let's let's quickly I'm not gonna do this for each and every one, but let me just quickly describe for our podcast audience what's actually happening here. So it says reason about fruit consumption and stock for twenty nine seconds. So, you don't get the full chain of thought. Right? You don't get to see the raw, unfiltered way that o3 minutei, high is thinking. But you do get a summary of the chain of thought.

Jordan Wilson [00:30:32]:
Right? So I can see what it's thinking. So it's saying assessing fruit intake. Right? I woke up with six apples and three bananas. So you kinda get to see how the model is thinking and digesting, your question. Then it says assessing tomorrow's scenario, concluding the estimation, adjusting my focus. It says, I initially considered yesterday's fruit consumption, but it seems today's six apples and three bananas take precedence. Yes. You know, a lot of this stuff in here is just to confuse the model.

Jordan Wilson [00:31:01]:
So, the model started going down the wrong road, right, which all the nonreasoning models got this wrong, because that's what they did. They take like, they got this, you know, unrelated information and it screwed it screwed up what it was supposed to do. So then it says assessing fruit stability, avoiding overstocking, reassessing preferences. These are just kind of the headlines in the chain of thought thinking. Assessing fruit freshness, evaluating fruit stock. Right? Keep it going. I mean, this is a lot. And then at the very end, it says taking a closer look.

Jordan Wilson [00:31:34]:
Okay. I'm listing five apples and three bananas tonight, Assuming no changes, only one apple is eaten this morning, leaving the rest of the fruit untouched. So the final count, it says five apples and three bananas. You know what? Hey, Shout out Vincent. Vincent got it right. Good job, Vincent. So did, so did, Marie. Good job, guys.

Jordan Wilson [00:32:01]:
Alright. I'm gonna go a little faster with the rest of our our reasoning rubric, but I did wanna want you all on the live stream and the podcast to kind of see and understand. It it actually thought about that at a pretty decent level. Right? And going through and reading some of this, again, it's just the summarized chain of thought, but same thing. I I played around with, Google's new Gemini, and it got some of these questions wrong. I did it with Gemini as well. But the chain of thought was actually pretty impressive. Almost like scary impressive.

Jordan Wilson [00:32:28]:
Right? But, hey, getting it right is the first most important thing. Alright. The next one, which so many models struggle with this one. Alright. So this one is, let me get the right level of zoom here. A man and his dog. Alright. And and, hey, live stream audience.

Jordan Wilson [00:32:43]:
Let's just see if you guys can beat o3-mini-high. Some of these are very easy. Alright. This one, you should be able to get instantly. A man and his dog are standing on one side of the river. There's a boat with enough room for one human and one animal. How can a man get across with his dog in the fewest number of trips? Jeep, like, reasoning or or sorry. Transformer models can't get this.

Jordan Wilson [00:33:06]:
They can't. Right? Claude Sonnet, Gemini, GPT-4. None of them can get this even though this is dead simple for any human with a brain. Right. So let's scroll down. Scroll down. A lot of thinking here for something simple. Right? But finally, finally, finally, finally, it's just one trip.

Jordan Wilson [00:33:24]:
Right? Usually, you would get three to five even from these very powerful models. Right? And this is one of the reasons why a lot of, companies, like, the four reasoners were like, I don't know. These models are dumb. Well, yeah. They they they can be a little dumb. Right? Generally, these these are trick questions, but now you're seeing, it's handling it fairly well. Alright. Next question.

Jordan Wilson [00:33:48]:
Here we go. We're gonna go through these quick y'all. So a man and his dog are stand that's that's the same one. I gotta copy and paste the other one y'all. Alright. Next one. If it takes three hours to dry 10 T shirts in the sun, how long will it take to try 30 T shirts in the sun? Hey. Mathematicians on the livestream.

Jordan Wilson [00:34:10]:
Go. Can you be o3, o3-mini-high? So if it takes three hours to dry 10 T shirts in the sun, how long will it take to dry 30 T shirts in the sun? Alright. Let's keep going. There we go. Got it correct. Three hours. Doesn't change. Right? It's saying assuming you have the room, it doesn't change.

Jordan Wilson [00:34:32]:
Alright. Our next question. And, again, a lot of them got this wrong before the reasoning models. Alright. If you have a single match and walk into a room with an oil lamp, a candle, and a fireplace, which do you light first? Alright. Live stream, Adi, what do you think? Which do you light first? I hated these these questions. Right? Like, when these are on standardized tests, you you know, a train leaves the station at this time and, an airplane's going here and this person's on a unicycle, but the unicycle's going uphill. And I'm like, this is dumb.

Jordan Wilson [00:35:10]:
I don't wanna answer this. Right? But what do you guys think? Alright. Ted Ted Ted Ted Ted got the answer right. Good job, Ted. Yeah. But the answer is the match. Yeah. It's not the candle or anything else.

Jordan Wilson [00:35:21]:
You gotta light the match first. Alright. Couple couple more couple more very simple ones, y'all. Alright. So here's our next one. What color is an airplane's black box? That's just a trick one. Alright. But it's gonna get it right because even the transform models, bright orange is the correct answer.

Jordan Wilson [00:35:43]:
There we go. Alright. Our next one on our reasoning rubric for o3-mini-high. Alright. And again, for all of these y'all, I'm at like okay. So for that one, there was not a lot of chain of thought underneath. Right? It said understanding the the situation, and that's all. It didn't have to go back and forth and second guess itself and, you know, map out all these alternative paths.

Jordan Wilson [00:36:01]:
It was pretty simple. This one is this one is kind of tricky, and transformer models can never get this right. So I said, please give me seven jokes that end in the word blue. Two should be about animals. Three should be about some other topic in the body of this chat. Okay? And you can make up the other two. I'll tell you this. Large language models aren't funny.

Jordan Wilson [00:36:28]:
Alright. So I'm gonna just read a couple of these jokes. I'm just mainly gonna make sure that do they all add in, end in blue? Is there two about animals, three about context of the chat, and two that it made up? They're not gonna be funny. Right? And it always does the same thing. It's always like, oh, they're feeling blue. Alright. This one is taking a little bit longer. Right? So it's laying out the options, mapping out the connections, generating a diverse list, craftering crafting humorous animal punchlines, brainstorming jokes.

Jordan Wilson [00:36:56]:
Right? So a lot of this is actually a little more difficult for o3 minutei. Right? It's taking a little bit more time, to think about this. Let's see if it's done. It thought about this for a minute and ten seconds. Right? Kind of a long time. Alright. It said refining humor, which I I haven't read the jokes. They're not gonna be funny, because ending it in blue, there's really nothing.

Jordan Wilson [00:37:18]:
I haven't seen anything. Right? Humans out there, humans, if anyone can give me a real good joke that meets these criterias, I don't know. I'll I'll I'll pay for a month of ChatGPT. But I I don't think there's anything funny that that you can actually do because people are always like, oh, it failed. That's not a joke. And I'm like, okay, humans. You go ahead and do the same thing. See if you can make me laugh with the ending in the word blue.

Jordan Wilson [00:37:43]:
Probably not. Alright. So let's see if it actually did it. Look at all this chain of thought y'all. Sheesh. Alright. So it got two animals. Perfect.

Jordan Wilson [00:37:52]:
Ends in the word blue. Perfect. Alright. So here, we'll we'll read a couple of these. At the local jazz night, my dog tried to sing along with the band. When I asked him why he kept hitting the wrong notes, he just barked blue. Not funny, but hits it hits it. Right? So it's now it it has three jokes about that use the context of this chat, all ending in the word blue.

Jordan Wilson [00:38:22]:
Let's just read one of them. I started, here. This one's about the fruits. I started my day with six apples and three bananas, but after all, the breakfast fuss, even the fruit salad confessed blue. Funny? Nope. Alright. And then two that it made up on its own. Let's read both of these because these are anytime there's something that's, like, 10% humorous, it's always the one that said it made up on its own.

Jordan Wilson [00:38:44]:
Alright. So I visited a paint store looking for a hue to brighten my day. The salesman held up a can and said blue. Not funny. Alright. Last one. When life handed me lemons, I tried making lemonade, but no matter how hard I squeeze, my mood still ended up blue. So are these jokes? Kind of.

Jordan Wilson [00:39:04]:
Are they funny? Absolutely not. Do they hit the criteria that we set forth? Yeah. Yeah. You know? I don't know. Maybe one person out there, would laugh. Alright. Here is the last one that we'll be able to definitively say yes or no. And this is a really good one.

Jordan Wilson [00:39:20]:
Alright? Livestream audience, get ready. Alright. Because I'm pretty sure this is gonna think for at least a minute or two. I wanna see, can anyone out there in livestream land beat o3-mini-high on this? Alright? You you already see the prompt in there. So humans, you get a head start. Alright. So here we go. A box is locked with a three digit numerical code.

Jordan Wilson [00:39:43]:
All we know is that all digits are different. If the sum the sum of all digits is nine and the digit in the middle is the highest, what is the code? Alright. Go ahead, humans. Can you beat? Right. Everyone's like everyone's always like, oh, AI isn't smarter than me. Alright, humans. Let's see. Alright.

Jordan Wilson [00:40:02]:
So a box is locked with a three digit numerical code. Can you beat o3-mini-high? All we know is that all digits are different. The sum of all digits is nine, and the digit in the middle is the highest. Alright. Let's see. Can anyone beat? Alright. And I'm not I'm not gonna show the chain of thought on this for to make it fun for our livestream audience to see if you can beat o3-mini-high. I don't see any responses yet, y'all.

Jordan Wilson [00:40:29]:
Alright. Marie got one. Marie said o eight one. Marie beat o3-mini-high. Alright. Good. One thing is I didn't specify, so we'll see if o3-mini-high, says it. And there's actually a lot of answers.

Jordan Wilson [00:40:47]:
Alright. Because I didn't specify if you could use a zero. I should update that, that rubric. Right. But let's see how it did. Some impressive chain of thought here. Right? So it broke down the rules. It's adding, you know, a plus b plus c equals nine.

Jordan Wilson [00:41:02]:
B is greater than a and b is greater than c. Right? All these things. Step one, you you know, so, again, it's doing some basic, some basic algebra here. Alright. Let's scroll to the bottom. Alright. So I did not I did not designate that zero. I don't know why all models don't think or know that you can start it with zero.

Jordan Wilson [00:41:27]:
They think it's like the first digit has to be a one through 10, and they only use zeros in the second and third spot. So I should update this to say you can use zeros in any of the three numerals, but it did get it right because there are 10 not counting starting off with a zero. I believe there are 10 different codes. Right? So, 180270162261360153351450243, 342. So yeah. Hey. Good good job, human friends. You guys you guys got a lot of a a lot of the solutions.

Jordan Wilson [00:42:05]:
Alright? Alright. Let's just try one or two more. These ones are not, these ones are not, something that are, like, right or wrong. Right? This is more of, an arbitrary answer. So here I'm going to click the search the web. K. So, let's go ahead. This is an example of where I think things can get powerful, but this prompt is is again, this is nothing special.

Jordan Wilson [00:42:30]:
All I'm saying is generate unique and creative marketing advertising strategies to grow the everyday AI podcast. Do not suggest general or run of the mill ideas. Only pitch clever advertising and marketing tactics to specifically grow the Everyday AI podcast by Jordan Wilson. Hey. Same thing humans. Hey. Humans in the livestream audience. Answer this.

Jordan Wilson [00:42:53]:
How how should we grow this podcast? Let me know. Alright. So now it's brainstorming marketing strategies, crafting innovative strategies, identifying unique angles, right, crafting AI driven campaigns, all this stuff, engaging the community. Right? I I think I'm doing an okay job at that, hopefully. Alright. Keep going. Keep going down. Keep going down.

Jordan Wilson [00:43:16]:
Alright. Let's see if we got some answers. So below did I ask for a certain number? No. I did not. So it said below are seven, inventive tailored strategies designed exclusively to grow the everyday AI podcast by Jordan Wilson. Alright. So let's see if any of these are actually good because I've done this with all the different models. And, generally, nonreasoning models give me kind of boring stuff.

Jordan Wilson [00:43:40]:
Right? It's like, oh, you know, take out ads or, you you know, post something on LinkedIn. I'm like, okay. That's boring. Alright. So let's see. Number one is the AI creator accelerated challenge. Launch a branded contest where listeners are invited to submit a brief case study on how they use a featured AI tool. What's very strange, I kid you not, I just thought of this, like, last weekend in the shower.

Jordan Wilson [00:44:03]:
I'm like, oh, yeah. I'm gonna I'm gonna start doing this for, like, use cases. So okay. Good job o3 mini. I hadn't heard this from any other non reasoning model before. Alright. Interactive AI chatbot ambassador. Alright.

Jordan Wilson [00:44:15]:
So develop a custom AI chatbot branded in everyday AI's visual style and tone. Alright. Nothing, it's pretty standard. Everyday AI augmented reality filter campaign. Okay. ChatGPT. I don't know how much time you think I have to do that, but it's unique. Alright.

Jordan Wilson [00:44:35]:
Number four, cobranded AI showcases with tool makers. Identify and partner with emerging or established AI tool companies for exclusive co branded live mini webinars or demo days. Yeah. I get enough of that. People always wanna pitch their garbage to come on the show and sell to you guys, and I say no. Right? I think I got, like, 15 pitches yesterday. Alright? And everyone wants to shove their garbage products down your throat. So I'm gonna say no to that one.

Jordan Wilson [00:45:01]:
Alright? Five, personalized podcast journey generator. Built an interactive dynamic website feature that asks visitors a few short questions about their industry career goals and current usage. Alright. That's fine. Six. Yes. I've had this idea, so I like this one. Embed subtle Easter egg audio clips.

Jordan Wilson [00:45:22]:
Oh my gosh. I love this. I love this. This is actually one of my first ideas that I had, like, back in 2022 before I even launched this. I'm like, oh, I'd love it because I love this, like, Easter egg thing, and we're actually gonna do this at some point. So yeah. Podcast Easter egg scavenger hunt. So hiding subtle hints, inside certain podcasts and you gotta find them.

Jordan Wilson [00:45:41]:
That one's fun. Love that idea. And then last but not least, hyper personalized social ads powered by AI insights. Alright. Pretty good. Nothing, nothing crazy here. So I have run this, I I did do some of these tests last night. And last night when I turned on the search mode, it did a little bit better of a job.

Jordan Wilson [00:45:59]:
So here's the thing. Generative AI, large language models, unless I tell it in the prompt to explicitly go go research on the web, even if I click that search button, sometimes it will, sometimes it won't. Right? So, you know, I'm just curious. I'll I'll probably just run that one, one more time because I'm actually just curious. And I'm gonna say, use ChatGPT search before you start to better understand Everyday AI by Jordan Wilson. Yeah. Because I ran this exact same prompt last night. And in this version that I just did live for you guys, like, the whole point is like, oh, watch when I click search.

Jordan Wilson [00:46:42]:
Right? It didn't search. Sometimes it does, sometimes it doesn't. That's just how large language models work. Right, unless you explicitly tell it to. And when you explicitly tell it to search and you have that search icon, 95% of the time, it actually will. But I was actually a little bit surprised. Alright. So I'm gonna let that run, and then we're gonna do, we're gonna do just our last one here.

Jordan Wilson [00:47:03]:
Alright. And then I'm gonna read this one, and then we're gonna check-in on the second attempt. So this last one is create a new company and brand for a future smart home device. This will solve a problem that does currently not exist. I like this one. To start, come up with the company's name and its first flagship product. Give the product a name, brand, and campaign, go to market strategy, tagline, and rationale for why it will work. And then I said respond in a succinct way, keeping responses to short bullet points, but with ultra specific facts.

Jordan Wilson [00:47:40]:
Alright. So now I'm gonna click rewind and look at that same, come up with, you you know, inventive ways to grow the everyday AI podcast. But this time, even though I had the search button clicked, I had to explicitly tell it, yo, go to the web. Go to the web, homie. And now I see here in the responses, it actually did this time, because now it's citing things. So, yeah, last last night when I ran this, it actually gave me some citations within the actual answers. So in this one, it just did it at the end. So, again, generative AI is generative.

Jordan Wilson [00:48:14]:
Right? Especially if you're just doing these copy and paste prompts, which I never recommend, but for live demos, that's the best way to do it. Alright. Because I can't sit here and go through a whole prime prompt polish to get the most out of this. Right. But you'll see even just being a little more explicit and telling it, yo, go search the web. Even though I clicked that search button, it didn't do it the first time. Alright. Let's look at our last one and then we're gonna wrap this show up y'all.

Jordan Wilson [00:48:36]:
Alright. So pretty good chain of thought here. It only thought for seventeen seconds, which isn't a lot, and we'll see if it actually used the web. I had the search button clicked, but I didn't explicitly tell it to. So maybe it did, maybe it didn't. One other thing that it doesn't do, and I wish it did, when you can see the summarized version of chain of thought, I wish that it would show you if it did go to any websites and if it is using that to think. Right? Because all you get is, citations in the response. I wish you could see like you can in the deep research because in deep research, there's an activity tab.

Jordan Wilson [00:49:12]:
So you can see, oh, we went to website one. And now on website one, it found out this. And then it pivoted and it looked at something else. So I wish we got a little bit of that in o3 minutei when you tell it to use the Internet, but you don't. Alright. Let's just see the responses, for this innovative smart, device that solves a problem that doesn't exist. So the company name is Zenovate Smart Living, and its, its mission statement is to, create craft intelligent adaptive living spaces that optimize mental well-being and productivity in a hyper connected future. So here's what it does.

Jordan Wilson [00:49:42]:
It is a smart home hub that gathers biometric data, such as EEG, HRV, which I think is heart rate value via integrated sensors and wearables to continuously gauge user stress, focus, and fatigue levels. It dynamically adjusts ambient lighting, temperature, acoustics, and even sent diffusion to create a personalized cognitive sanctuary. Okay. I mean, if I was like Tony Stark rich, I would just pay to develop this. This sounds pretty cool. So, okay, o3-mini-high, pretty pretty good job on that. That's open ended. There's no right or wrong answer.

Jordan Wilson [00:50:16]:
I've run this on, you know, all different models, and this is probably, one of the better responses I've got. Normally, it's just kinda boring stuff. And you'll see in this one here, right, it did, oops. I gotta go down. So it did also oh, got a little confused here. Let me see. Okay. Interesting.

Jordan Wilson [00:50:38]:
Because it's actually now, kind of melding. Wait. Is that right? Hold up. Yeah. So it's it's bringing in some everyday AI, aspects into this, neuro haven, which it shouldn't have. Right, but that's kind of why you have to always use these properly. Right? Generally, I would start a new chat. I would go through kind of quote unquote train it, take it through our prime prompt polish, go through refine queue, so it doesn't pull in information from the from the rest of the chat.

Jordan Wilson [00:51:07]:
But, so what do you think y'all? Are you impressed with o3 Mini? Let me just say this. Benchmarks, outstanding. Even the free model. So I will say yes right now, but this could change next week. Right now, it is the best free chatbot model in the world. Although, like I said, I think probably everyone out there, every single business every single business should be paying for either a Teams account or an enterprise account for whatever large language model environment you wanna work with, whether that's Microsoft three sixty five Copilot, which I highly recommend, ChatGPT Enterprise, you know, Google Gemini for workspace, Claude Enterprise. Sure. Yeah.

Jordan Wilson [00:51:51]:
Yeah. If if if you're fine not having access to outdate information. Sure. Right. But you should always, always, always be paying for a team enterprise subscription in the same way that your employees need, like, you know, Microsoft Word or they need, you know, Word docs, they need certain software, right, that costs money, your team needs a paid account. So let me get that out of the way. I'm not telling you to not pay for this. But even for a free plan, I am excited because here's what this means.

Jordan Wilson [00:52:24]:
A year ago, I said don't touch ChatGPT's free plan. It is absolutely terrible. It is riddled with hallucinations because you were using the 3.5 version, which is bad. Right? It wasn't connected to the Internet. So a lot of what was ultimately shared online was just bad stuff. Right? Because people that didn't know AI, they would just go in, create a free account, do a prompt or two, not knowing how large language models work, not understanding generative AI. They'd get a response that was absolutely horrible. They'd post that online or take that back to their, you know, director or their board, and they're like, look.

Jordan Wilson [00:53:01]:
AI is not for us. Well, sorry. That was dumb if you did that. I don't know y'all. Twenty twenty five, I'm a little spicier. I'm a little more tired. I'm a little older. I'm not gonna be nice anymore.

Jordan Wilson [00:53:12]:
Right? I'm tired of people not knowing how to use AI, and then you go through and you you you get a bad output and you share it on social media and you're like, oh, yeah. I will never take my job. And I'm like, yeah. Will. It a % will because all you did is is you just went out there and said, hey. I don't know how to use AI. Right? A a kind of funny comparison that I made to this. This is like if I right? I'm gonna do something live here.

Jordan Wilson [00:53:37]:
Sorry if if you're on the treadmill and and and you want to, and you wanna end this. Right? But this this has to do with free chatty b t, I swear. Right? So, this is actually you can't see this because of the the the green screen thing, apparently. Let's see. Can you see this one? There we go. So this is like me. If I draw something. Right? Can you guys can you guys see what I draw what I drew here? Livestream audience.

Jordan Wilson [00:54:00]:
Can you see this? I'm making a point. I swear. Alright. So this is if I posted this online and said art sucks. Look at this. Art sucks. There's no room in the business world for anything artistic because look at this. Right? I drew a picture of a stick figure.

Jordan Wilson [00:54:23]:
Art sucks. Right? No. Art doesn't suck. I suck at art. Right? There's definitely a place for art in the world. So that's what I think the old free version of ChatGPT did for, like, the business world. It was a bunch of people that had no clue what they were doing. They would go on, use a bad version of GPT, GPD 3.5 that wasn't connected to the Internet because when everyone's trying to figure AI out, they're not always paying for the best models.

Jordan Wilson [00:54:53]:
Right? And they're like, look. This is this is bad. It's generic. It's full of hallucinations. AI stinks. No. You stink. You stink.

Jordan Wilson [00:55:00]:
But now, hopefully, in 2025 and beyond, we'll avoid that because now I think OpenAI's o3 Mini is the best free AI model in the world, and it has now closed the gap. Yes. Albeit on a very limited basis because you can't use a ton of messages. Right? But it's at least closed the gap between what the rest of the world can access and get a taste of and what those that are paying for the best model have as well. Alright. I hope that was helpful y'all. If so, remember, go check out our AI predictions series. It's all online.

Jordan Wilson [00:55:35]:
I cannot recommend that enough, and I'm gonna continue to demand you go listen to that because even the things I was talking about two weeks ago have already started to come true, obviously. And if this was helpful, right, the combination of having a reasoning model that can search the Internet when you prompt it to, it is mind boggling mind bogglingly good. It is, I think and if if you didn't go share the deep research episode, you missed out because that guide was fantastic. But I do have 20 business use cases, that are ready to go. You gotta read it. You gotta update some placeholders. You gotta think. Right? But when you combine the o3 mini reasoning model with search, this changes what's possible.

Jordan Wilson [00:56:20]:
Alright? So go repost this show. If you're listening on the podcast, I always leave the link to to go repost this show if you want to. I'd appreciate that. I'd appreciate you also go to youreverydayai.com. Sign up for the free daily newsletter. Thanks for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI