Ep 469: Claude 3.7 Sonnet: – World’s first hybrid AI model. How it works and when to use it

Resources:

Join the discussion: Ask Jordan questions on Claude


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Exploring the Next Frontier in AI: Hybrid Large Language Models

The world of Artificial Intelligence is witnessing innovations at a rapid pace, and the recent release of Claude 3.7 SONNET by Anthropic signifies a pivotal milestone. As the first publicly available hybrid large language model, Claude 3.7 SONNET opens up exciting new avenues for business applications in AI. Here's what business leaders need to know about this groundbreaking development.

Introducing Hybrid AI Models

Anthropic has set a new benchmark by launching Claude 3.7 SONNET, a hybrid AI model combining traditional transformer models and advanced reasoning capabilities. This innovation marks a shift from having separate models for tasks like content generation and logical reasoning to a unified model that can intelligently choose between generating quick responses or engaging in complex, extended thinking. This flexibility is poised to change how businesses leverage AI for diverse tasks, from coding assistance to strategic decision-making.

Hybrid Models: Bridging Old and New AI Paradigms

Historically, AI models have been divided between "old school" transformer models and newer reasoning models. Transformer models have been adept at content generation tasks, while reasoning models excel in logic and complex problem-solving. Claude 3.7 SONNET embodies the best of both worlds, offering businesses a versatile tool that adapts to their specific needs, ultimately reducing the need for multiple AI solutions.

Implications for Enterprise Application

Claude 3.7 SONNET is a game-changer in fields requiring intricate problem-solving, like coding, where it already has shown significant improvements over its predecessors in industry-standard benchmarks. Additionally, its ability to output 28,000 tokens, a 15-fold increase from its previous output capacity, is particularly beneficial for processing larger datasets and more extensive dialogue without interruption.

Refinement in AI Thinking and Visibility

A hallmark of Claude 3.7 SONNET is its visible thought process. Businesses can now observe the model's reasoning, providing insights into how it reaches conclusions—an invaluable feature for tasks that require transparency, accountability, and fine-tuning of AI interactions. Developers can customize thinking durations, opting for more intensive reasoning or speedier, less resource-intensive responses, thereby optimizing AI resource allocation.

The Financial Perspective: Cost and Value Considerations

Despite its advanced features, the API pricing for Claude 3.7 SONNET remains premium, reflecting the industry trend where companies may lean towards hybrid models for increased profitability. This pricing could be a consideration for businesses, especially when other models, such as OpenAI’s GPT series, offer competitive options with different cost structures.

Strategic Moves: The Road Ahead for Businesses and AI

With Claude 3.7 SONNET’s capabilities, businesses can integrate AI more effectively into their operations, discovering new efficiencies and insights. As AI models continue to evolve, companies should strategically assess their needs for reasoning versus generative tasks, potentially leveraging newer models like Claude Code for enhanced software development and operational agility.

Conclusion: Embracing the Hybrid AI Era

For business owners and decision-makers on LinkedIn, the introduction of hybrid AI models like Claude 3.7 SONNET represents a transformative step forward in AI technology. By embracing these new tools, companies can remain at the cutting edge, redefining how they approach challenges and leverage AI in a fast-evolving digital landscape. The strategic implementation of hybrid AI models promises not only to enhance efficiency but also to drive innovation across industries.


Topics Covered in This Episode

1. Overview of Claude 3.7 Sonnet
2. Performance Benchmarks
3. Potential Use Cases
4. Discussion on Hybrid AI Models


Podcast Transcript


Jordan Wilson [00:00:16]:
Another week, another state of the art large language model release. But this one from Anthropic is a little different. It's actually the first of its kind because when Anthropic just released its Claude three point seven SONNET, they became the first company to release a hybrid large language model. Alright. So we're gonna be talking today about what that is, what it means, how it works, and when you should actually use this new model from, Anthropic. I hope you're excited for this show. I am. Welcome if you're new here to Everyday AI.

Jordan Wilson [00:01:00]:
What's going on y'all? My name is Jordan Wilson, and this is Everyday AI. This thing is for you. This is your daily livestream podcast and free daily newsletter helping us all not just keep up with Jenna AI advancements and LLM updates, but how we can use it to get ahead. I want you to be the smartest person in AI at your company, and this is your cheat code. So if you haven't already, go to youreverydayai.com. That is where you can sign up for our free daily newsletter. Yeah. Maybe you're listening to this podcast for the first time.

Jordan Wilson [00:01:33]:
If so, thank you. Make sure to check out the show notes. There's gonna be a lot of other information, but probably the most important is our website because each and every day, on in our newsletter, we recap exclusive insights from this exact podcast as well as giving you every other piece of news and update that you need to stay ahead, in the generative AI space as well as you can go listen to, like, 500 episodes, on our website, all sort of by category. So make sure you go check that out. Alright. So I am extremely excited, to talk, about the new Claude three point seven SONNET. I think it's gonna change, how a lot of people are using large language models both for the good and for the bad. But before we get into that, let's first start out as we do most days by recapping the biggest AI news.

Jordan Wilson [00:02:22]:
So first, Google has launched a free version of its AI powered coding assistance, Gemini code assist, aimed at solo developers, students, freelancers, startups, and hobbyists. So the new free public preview offers up to 180,000 monthly code completions significantly exceeding the 2,000 completions, for free offered by competitors like GitHub Copilot's free tier. So it is powered by the Google Gemini two point o model, and it can generate entire code blocks, autocomplete code, debug, and assist developers via a chatbot interface. So users can instruct the assistant in natural language, such as asking it to create specific code snippets or modify existing applications. So Gemini code assist supports 38 programming languages and integrates with popular developer environments like Visual Code Studio, GitHub, and JetBrades. Alright. Our next piece of AI news, Apple, a couple years, maybe too late. I don't know.

Jordan Wilson [00:03:25]:
But they're making a splash with a reported $500,000,000,000 investment over the next four years into AI infrastructure, signaling a major push into not just AI, but American manufacturing and technology. So according to Apple CEO Tim Cook, this commitment reflects confidence in the future of American innovation and aims to strengthen the company's role in AI and advanced manufacturing. So a key part of this investment, 500,000,000,000 with a b, includes the development of a key new manufacturing facility in Houston, providing thousands of jobs to produce servers designed for AI cloud computing. These servers will feature the Apple silicon and offer, cutting edge security and performance capabilities. So the integration of Apple Intelligence, the company's AI platform, could further transform health care specifically by leveraging its global network of over 2,000,000,000 active devices to provide innovative health tracking and data insights. So Apple investments come amid a broader AI spending race with competitors like Meta spending 65,000,000,000, Amazon spending a hundred billion, and project Stargate, which is 500,000,000,000 over five years, also ramping up their AI infrastructure and innovation budgets. Speaking of, that's our last piece of AI news. Microsoft, just on the same day that we get reports that Apple is going all all in with a $500,000,000,000, investment, reportedly, Microsoft is canceling leases for a couple of hundred megawatts of US data center capacity equivalent to about full, two full data centers, and that's according to a report from TD Cohen.

Jordan Wilson [00:05:14]:
So this move raises concerns about whether Microsoft, obviously a global leader in AI investment, may be securing more AI computing capacity than it needs in the long term. So the cancellations involve agreements with private operators and a slowdown in converting statements of qualification, which are typically precursors to former formal leases. So TD Cohen speculates that OpenAI, which is backed heavily by Microsoft, may be shifting some of its workloads to Oracle as part of a new partnership, which may be causing Microsoft to cancel or change some of its longer term investments. So Microsoft, which owns and operates many of its own data centers, is also reallocating billions of dollars in infrastructure investments, potentially shifting focus back to The US from international projects. So despite these adjustments, Microsoft reiterated its $80,000,000,000 spending target for an AI data center infrastructure for the fiscal year ending in June. So analysts suggest that Microsoft could be in an oversupply position may meaning it may have overestimated the immediate demand for AI computing power. So it'll be interesting to see how those stories play out especially happening at the same time. I mean, Apple, you know, making a huge splash with a $500,000,000,000, investment where we get reports that Microsoft may be slightly scaling back.

Jordan Wilson [00:06:41]:
Alright. Enough. If you want more AI news, make sure you can go get it at our website. Sign up for the free daily newsletter, youreverydayAI.com. Alright. Let's get into it. Let's talk the world's first large language model hybrid. Alright.

Jordan Wilson [00:06:57]:
And that is with Claude three point seven SONNET. Alright. So it is the world's first publicly available hybrid AI model. So what that means and we're gonna get more into this. Right? I've been talking about this on the show now for, I don't know, at least six months since OpenAI kicked off this reasoner race. Right? So you essentially think of it when it comes to generative AI in large language models. I know I hate to use terms like old school, when technically, you know, this space is only, like, I don't know, six years old. You know, at least commercially available, you know, the GPT three technology, right, I would say would be the first large language model that was popularized, commercially a couple of years before the ChatGPT release.

Jordan Wilson [00:07:43]:
So you have your kind of quote unquote old school transformer models, and then you have your quote unquote new school, reasoning models that kind of use this advanced thinking. Alright? And right now, those are two very separate things. So as an example, OpenAI, the leader, in large language models, you know, they have their GPT four o, still an industry leading model even though it's technically older, but at that is its kind of quote, unquote old school transformer model. And then they have their newer kind of reasoning models that use logic kind of under the hood, and that is, you know, o one, o one pro, o3 mini, o3 mini. Hi. Yeah. These names suck. Right? But, you know, you essentially have these two very different types of models that excel at two very different kinds of tasks.

Jordan Wilson [00:08:30]:
So now with Anthropic, they are essentially merging this together in cloud 3.5 sorry, cloud 3.7, SONNET, and it kind of does both. And, this new hybrid system will kind of decide on its own when it should use more of this advanced thinking versus when it should just straight out spit an answer to you without really thinking it through. Alright. So, let's go over a little bit, and it's not just the 3.7 SONNET. They also announced Claude Code, which I think is, an extremely big move from Anthropic and, kind of tips its cap into where it's actually competing in. So more on that in a second. But, hey, livestream audience, thanks for joining. Appreciate y'all tuning in.

Jordan Wilson [00:09:18]:
We got an international audience today. So, thanks to our, you know, YouTube audience. We got big Bogeyface and Sandra and Sam. Michelle, thanks for joining. On the LinkedIn crew, we have doctor Harvey Castro, Christopher, Woosie, LinkedIn user, Marie, Danny, Douglas, Cecilia, Jean. Thank you all. Jamie, Karina. You know? We got The UK and Italy in the house.

Jordan Wilson [00:09:40]:
Love to see it. Mac, Max holding it down from Chicago just like me. But I'm curious, livestream audience. Do you care about these, like, this new hybrid approach? Because it's something that OpenAI is also going to be adapting to as well. I think it's there's actually some downsides, and we're gonna talk about this. But, you know, livestream audience, I'm curious. Number one, do you care about this hybrid approach? Do you think it's gonna be good or bad? And have you used, Collard three seven SONNET yet? I know it's only been out for a couple of hours. If you do have any questions, get them in now.

Jordan Wilson [00:10:13]:
I'll try to tackle them, at the end of the show. Alright. So, let's get into an overview, and this is from Anthropic. So they're saying today we're announcing Claude three seven SONNET, our most intelligent model to date and the first hybrid reasoning model on the market. Claude three Sonic can produce near instant responses or extended step by step thinking that is made visible to the user. Yeah. So that part's important. You can kind of see, like you can.

Jordan Wilson [00:10:40]:
It is a summarized chain of thought. So chain of thought is actually a prompting technique that was popular, popularized, you know, over the last couple of years using transformer models. So this kind of chain of thought or, you know, how a a person would think about a problem. So now a hybrid model does that and it shows a summarized version of the chain of thoughts. You can kind of see how this new, three seven SONNET is kind of thinking about your prompt if it is using, the advanced thinking. So now back to the release API users also have fine grained control over how long the model can think for. So Claude three seven sonnet shows particularly strong improvements in coding and front end web development. Along with the model, we're also introducing a command line tool for agentic coding called Claude Code.

Jordan Wilson [00:11:29]:
Claude Code is available as a limited research preview and enables developers to delegate substantial engineering tasks to Claude directly from their terminal. Alright. So, a lot to unwrap there. So you don't have to read. I think there's, like, three separate releases that Anthropic put out. I'll just give you the high level. So like we said, this is the first hybrid reasoning model with with visible thinking process. And the extended thinking is on paid plans only.

Jordan Wilson [00:11:59]:
So if you are a free user, to, anthropic Claude, you will see the 3.7 SONNET model available, but you do not get, kind of this advanced thinking, available on the free plan. Alright. Couple other high level, kind of, points here. It scored a 70.3 on the SWE bench or SWE bench verified. Best in class by a lot for coding. Like we talked about, the Claude code program, for agentic development, and then it has a 15 times longer output token capacity. So a 28,000, tokens that it can output versus previously, Claude could only output 8,500 tokens. So that's just the amount.

Jordan Wilson [00:12:48]:
Right? So if you ask Claude to do something, before it would spit things out, sometimes if you ask for a lot, it would spit things out in little chunks. So now at least according to Anthropic, that is a one, 28,000 token output, capacity. Personally, I'm not seeing that yet. We're gonna do a live test here, y'all. We'll see if we actually see that. I was still getting it, breaking it out in small chunks. They did say that is in beta, so not sure if that's fully rolled out yet or if that'll be coming out in the coming days or weeks. But I don't know.

Jordan Wilson [00:13:24]:
I'm not seeing it. Also, which is important, this is available across all platforms. So a lot of what we're gonna be talking about is using Claude on the front end as a front end user. Right? So going to Claude a I, Claude.ai and, you know, using your free account, your paid account, maybe you have a Teams account. Right. But, obviously, Claude is available on the back end, and it is it is a very popular model on the API side mainly due to its, proficiencies in coding and software development. It is, historically been the most used model, at least when you're looking at open router statistics. It is generally the most used model, on the API side, at least those that are using open router.

Jordan Wilson [00:14:07]:
Right? Open router is one of the more popular, services where you can essentially sign up from one service, connect all your different API keys. So they have good data, but that's not every single model. That's just those using open router. Alright. So let's talk a little bit about Claude's thinking because this is the big, the big, you know, chain of thought, a reasoning model. So let's go over some of the highlights on how this actually works. So it uses deeper reasoning for complex tasks. So what that means is it has an extended thinking mode that lets Claude spend more time and come and and, compute effort solving challenging problems or answering tougher questions.

Jordan Wilson [00:14:53]:
Okay. It has user controlled thinking budget on the back end. So, developers can set a thinking budget to determine how much effort Claude should apply for a task. I think that's where things get a little tricky. We'll talk about that here in a second. It is the same model more effort. So the extended thinking doesn't rely on a different model. It is still the Claude three seven SONNET.

Jordan Wilson [00:15:17]:
Right? So hybrid model whereas OpenAI, as an example, has their o one, o3, and then they still have their workhorse do everything model GPT four o. Not like that with Claude. It is just Claude three seven SONNET. Right? It's not three seven SONNET thinking. It's not o3 seven SONNET alphabet soup. It's just three seven SONNET. It's the same thing. One model does it all.

Jordan Wilson [00:15:38]:
I think there's pros and there's cons. The extended thinking, like I said, doesn't rely on a different model. Visible thought process, that's a big new feature at least for Claude. Right? So users can see the it says the raw reasoning steps. I don't know. We'll have to see if if that's the raw reasoning. When I'm looking at it, it still looks like a summarized chain of thought. I could be wrong.

Jordan Wilson [00:16:03]:
We're gonna look at it live. The other thing, Claude is historically terrible for limits. So, you know, I was able to test this, a lot last night, and I wanted to do a ton more testing this morning before this live show. But, you know, even though I'm on a paid plan, Claude's limits have historically been the worst in the industry, and it's not even close. Right? So, I wish Claude would give paid users a little more leeway, in order to test these things. So, a lot of these things, I've already done them a couple of times, but normally, I would like to play, with an LLM for at least, you know, six to eight hours before doing an even simple show. Not always, an an option, at least using Claude on the front end because those limits are terrible. Alright.

Jordan Wilson [00:16:51]:
Also improved accuracy over time. So Infropic says that extended thinking boosts performance on tasks like math problems or complex evaluations by allowing Claude to refine answers iteratively. Alright. So let's talk about the Claude update timelines because if you're wondering, wait. Has it been a minute since we've heard from Anthropic? Yeah. Kind of. Right? When now the the leaders, Google and OpenAI seemingly are announcing new models every month. It has been like light years and then some since we've had an actual step improvement from entropic.

Jordan Wilson [00:17:31]:
So the original 3.5 SONNET was back in June 2024. Alright. Then they had this upgraded 3.5 SONNET, which was confusing because they just called it 3.5 SONNET new. They didn't use 3.6 even though, a lot of people online, myself included, said this is dumb. Why are you calling it 3.5 SONNET new? And then they obviously skipped 3.6, which lends me to believe that, yeah, that 3.5 SONNET new, which really didn't bring anything terribly new. It was more of an under the hood update, the type of updates that, you know, Google and OpenAI do almost on a biweekly basis. It didn't seem like anything major, but we saw the, the Claude three five SONNET new in October. Then in November, we saw Claude three five Haiku.

Jordan Wilson [00:18:26]:
Right? So, essentially, Anthropic has historically had three model sizes, small, medium, and large for small tasks, medium tasks, and large tasks. So Claude, Haiku is the small, SONNET is the medium, and OPUS is the large. So you'll see here now, finally, February 24, we got the Quad three seven SONNET. So I will say the three five, you know, new update in October, I don't know. That wasn't much. I used it plenty. I use quad three fives on it every day. I didn't see anything new, anything noticeable, at least for my daily use case, which I know is different than a lot of people's.

Jordan Wilson [00:19:06]:
Right? But so I'll say this. For the most part, it's been since June. It's been a good eight months since we saw a top class model real update from Anthropic. So it's been a hot minute. So let's also talk about what's next because Anthropic did release this little, I guess you could call it a timeline, but looks very much in step with OpenAI's kind of five phases to AGI. Right? So you have your, you know, your reasoners, your agents, etcetera, from OpenAI. Claude takes a little different approach here. So they said 2024 was Claude assist.

Jordan Wilson [00:19:44]:
Then they said 2025 now is Claude collaborates. And then they said in 2027, Claude will pioneer. So is this kind of their AGI artificial general intelligence timeline? I'm not sure. It kind of looks like it. Right? They're saying it looks like Claude is just gonna be a collaborator. It goes from assist to collaborates, from 2024 to 2025, and it is going to be a pioneer in 2027. So I don't know what that means. But it's Tuesday, y'all.

Jordan Wilson [00:20:18]:
Should I come in with some hot takes? Let me know how spicy. I gotta get a sip of sip of coffee here for livestream audience. But how how hot should I make these hot takes y'all? And, yeah, if you do listen on the podcast, this is a live stream. We do it every single day. It's unedited, unscripted, realist thing in artificial intelligence. 07:30AM. I know it's a little early. That's why sometimes I take a little second to sip on the coffee.

Jordan Wilson [00:20:43]:
But, yeah, livestream audience, should should I be nice? Should I bring some heat here with my hot take takeaways? It is Tuesday after all. So alright. Let's get get to some of my, some of my takeaways here, and then we're gonna get back to the facts, the figures, the stats. We're gonna do a live walk through as well. So let's talk about this concept of hybrid models. Big Bogey Face said sweat emoji. Alright. Allison says just spicy.

Jordan Wilson [00:21:22]:
I'll keep it just spicy. Maybe I won't go, you know, five alarm, hot chili, hurts in the toilet, spicy. Alright. Not a fan of hybrid models right now. I'm not. But I'm also a power user, so I have to understand most people are not. I ultimately think these hybrid models are just gonna be a way for companies to make more money, right, which I get and I understand. I've said all along whether you're talking about $20 a month for Claude, you know, paid plan, $20 a month for ChatGPT plus, $200 a month for ChatGPT pro.

Jordan Wilson [00:22:08]:
Same thing with Gemini. Whatever. Companies, for the most part, are losing money. So I get it. I get you. You gotta make money. You gotta be profitable. But on the API side, if I'm a developer and and I've been using, you you know or maybe looking at switching from OpenAI to Claude, I am not incentivized to do so.

Jordan Wilson [00:22:31]:
Because when you have this new Claude three seven, SONNET, yes, you have this kind of slider control over how much thinking you can apply to certain situations. But when there's companies out there that literally their business model is essentially creating a helpful wrapper around an AI model for their customers for a certain niche. You need a little more control over a simple slider, you know, over saying like, you know, let's apply this much thinking unilaterally across the board. I don't think there was anything wrong from a back end API perspective. Right? So I'm hoping that Anthropic and others will not get rid of, you know, as an example, three five SONNET, and we'll still allow companies to have, you know, three five haiku, three five SONNET. And the reason why is because the API prices for three seven SONNET are ridiculously high. Ridiculously high. Alright? And if you don't have an option to have, like, you know, 3.7 SONNET regular and 3.7 SONNET think, I mean, there's a reason why right now OpenAI is winning the AI race.

Jordan Wilson [00:23:48]:
I mean, number one, they were the first with ChatGPT. Number two, even though it's confusing for front end users to stare at eight different model selections, it's extremely important for back end developers, companies that are essentially running their business off this technology to use the right model for the right time, for the right purpose, and the, costs that are associated with it. So Claude three point seven SONNET is extremely expensive. So for certain use cases, no brainer. Coding, software development, etcetera. Right? You're gonna pay it, because right now, quad 3.5 or quad three seven SONNET is the best in those areas. It is. It's a great model.

Jordan Wilson [00:24:31]:
I'm not a huge fan of it, and I probably won't be a huge fan of it, when OpenAI does that as well. So CEO OpenAI CEO Sam Sam Altman said that OpenAI is shifting once GPT five comes out. GPT five is going to be more of a system, and it will also use this hybrid approach. And it will say, you know, hey. Here's, you know, here's when you should use a, a a reasoner versus when you should use a a transformer model. So I just think this is just a way for these companies to make more money, if they eventually take away the option to use older models that are not hybrid. It's all I'm saying. And as a front end power user, I hope I always have the option as well.

Jordan Wilson [00:25:18]:
Right? I got some sun shining in my face. Alright. So I hope as a front end user, I'll still have the option in the future to say, oh, I don't wanna use a reasoning model for this, or I need to use a reasoning model and only a reasoning model. Right? You might have to over prompt engineer if you're giving, you know, Claude three seven SONNET, you know, something on the front end and you want it to use reasoning and it's not, then you just have to go and take that extra step, you know, do a little extra prompt engineering to get it to use, this logic. Right? So there's huge downsides that I don't think people are talking about. Right? Everyone wants to wrap it in a bow and say, oh, it's the world's most powerful. It's hybrid. It's all in one.

Jordan Wilson [00:26:06]:
Okay. There's times and use cases that all in one is great, but I don't think this is one of them. Again, I'm a power user, so maybe my viewpoint is skewed. I personally like going into, you know, ChatGPT and seeing eight different models. Right? Because I'm using probably five of them for very specific use cases. I don't want one. Right? I don't. I could be wrong on that.

Jordan Wilson [00:26:35]:
Alright? Next, poor Opus. Poor Opus. Opus hasn't been updated in, like, a trillion years. So it looks like, I don't know, Anthropic may have just abandoned their big boy model, Claude Opus. Maybe they're waiting until they're kind of clawed at four point o models to bring back Opus. I'm not sure. But at least for now, poor Opus is bye bye. Also, I I don't know.

Jordan Wilson [00:27:04]:
So I I I saw, I saw a comment here. Let's see who said this. There we go. Douglas from LinkedIn said cur curious how this will improve Cursor into Windsurf. Right? I think now, Anthropic is competing with them. Right? Even though these IDEs right? So that's an integrated development environment. So, you know, like we talked about at the top of the show with the news, you know, Gemini, Codesys, GitHub Copilot, Cursor, Wind Surf, Lovable, Bolt. Right? There's all these kind of, IDEs or essentially now AI coders, right, where you can literally you can talk to it.

Jordan Wilson [00:27:46]:
You can type to it. Think how we have these large language models. Right? We have ChatGPT. Right? The g p t models. Gemini, Claude. Right? And then we have now this newer breed. They use a model, so you choose which large language model, but it is an AI powered IDE or integrated development environment like Cursor. Right? Cursor by default uses Claude.

Jordan Wilson [00:28:12]:
But it looks like with Claude code, which, you know, audience, let me know if if if you want us to go into that. Not today. At a later time, we'd have to have a show or two dedicated. It is more technical, but I think Claude Code is really cool. But it looks like Claude wants to compete more with those IDEs, then it looks like they want to compete in the strictly large language model space. And I think that makes sense. I think it makes sense because it looks like over the years, Claude has kinda carved Claude has kind of carved out its niche. And I'm not saying they're abandoning, you you know, general business use cases.

Jordan Wilson [00:28:54]:
They're not. But it looks like especially with Claude Code, especially with the MCP protocol that they put out, you know, computer use, even though it's clunky, it did get updated with now this, three seven SONNET, so we'll have to see if it's any better. But it looks like Anthropic is maybe just wanting to compete more in that space, especially by making Claude Code a free beta preview. Also, another hot take since you wanted it spicy. I don't think most companies are going to end up using Claude. A lot of people were waiting for this release because they assumed that Anthropic would be cutting their API prices because that has been the trend across the industry. Right? OpenAI, you know, has cut their API prices by more than 90% over the last eighteen months when it looks at their top state of the art model. Google, same thing.

Jordan Wilson [00:29:53]:
Just ridiculous a API pricing cuts. Right? Anthropic, not so much. They didn't change their pricing at all. Right? Yeah. It's a more powerful model, but you're paying the same price. But I think for the most part, businesses are not going to use Claude general use cases. They won't. Maybe I'm sure Anthropic knows this, but I will say 90% of businesses that are looking for a large a general use case, large language model to use on the API back end, whether that's for, customer success, whether it's for sales, whether it's for an internal knowledge base, I I'd say non coding, non software development, 90% of companies will not look at Claude, and I don't blame them.

Jordan Wilson [00:30:40]:
The prices are more ludicrous than the early two thousands wrapper. It's they're they're they're insanely not practical for everyday use cases. They're not. Alright. So let's look at those API prices. So clawed 3.7, 3 and this is per million tokens. It is a $3 input, $3.4000000 tokens input, and 15 for output. Alright? So, yeah, it's a hybrid model.

Jordan Wilson [00:31:18]:
Sure. But I'm still going to go if I'm a business leader, GPT four o Mini is great because you can chunk. Right? You can chunk different tasks to different models. And that's why I'm I'm going to this whole, like, this API, you you know, and developers using it. No one's gonna use 3.7 unless you specifically need software development, coding. Right? Unless you're in one of those categories, maybe some some stem areas. Right? But, otherwise, who's gonna touch it? When you look at GPT four o Mini is 15¢ versus the $3 input and then 60¢ versus $15 on the output side. Right? And when you can chunk it and when you can say, hey.

Jordan Wilson [00:32:04]:
For these type of questions, for customer success, for for sales, etcetera, we're gonna use GPT four zero mini because we don't need a hybrid 3.7, model to do 90% of what we would use it for. I mean, the cost savings there, it's like, I don't know, ten, six, like, 30 times as expensive? Like, absolutely not. Or 25 times as expensive? I'm doing math live on the fly. I don't know. From an API perspective, this does not make sense. I I was really expecting Anthropic to slash their prices, but it looks like they're not necessarily concerned with competing for everyday business use cases. They're like, yo, If you wanna use agentic tools, if you wanna use software development, coding, etcetera, maybe some engineering, like I said, some stem use cases. But for everyone else, nah, we're good.

Jordan Wilson [00:33:03]:
Because the combination as an example of GPT four o Mini at 15¢ and 60¢ and o3 Mini at a dollar $10.04 40. Duh. Right? And then the same thing with Gemini. Gemini two point o Pro, dollar 25 and $5, and then they have their flash, and then they also have flash thinking. I probably should have put that up on the chart. But it just doesn't make sense. It doesn't make sense. Their pricing on the back end does not make sense.

Jordan Wilson [00:33:31]:
And I think as, the other models get essentially better at coding and software development, because right now, yes, Anthropic Claude and with their 3.7, they have a huge lead there. So let's look at that. So some of these benchmarks here, we're looking at the SWE bench, s w e bench verified, and looking at some of the different benchmarks and, you know, you have, the version here without, you you know, they're calling it custom scaffolding or without that extra thinking. Even without the extra thinking, Claude three point seven, SONNET on Sweebench, sixty two percent, where their last version, 3.5 Sonic, was 49. OpenAI's o one is 48.9. o3 Mini, forty nine three. DeepSeek, forty nine two. But with the extra thinking, Claude is a 70%.

Jordan Wilson [00:34:24]:
Right? So that's what I'm saying. If you're doing any type of software engineering, nothing else right now comes close. Same thing with agentic tool use. So the TAU bench, I think it's pronounced TAU bench, but t a u bench. Same thing. This is when you essentially have a model, you give it access to tools, and you have it go, complete, some technical tasks. Same thing, quad 3.7. Sonnet here with an 81%, on the towel bench, retail and then OpenAI seventy three percent.

Jordan Wilson [00:35:01]:
So not close. Generally, with a lot of these benchmarks, you know, especially some of the nontechnical, non software engineering ones, one point difference can be huge. Right? So in this use case is, Claude three point seven SONNET is light years ahead. Interestingly enough, when we look at the regular benchmarking, kind of marks here, between Claude three seven SONNET, Claude three five, OpenAI o one, OpenAI o3 Mini, DeepSeek r one, and Groc three beta beta. This is from, Anthropic's website. Interesting. They didn't include on this main one anything from Gemini, and they're these are definitely cherry picked. But something that I found interesting when, Infropic was putting out its own benchmarks, on its website is they didn't use the same benchmarks as they had previously, when they announced Claude three point five's on it, when they announced, the Claude three family of models.

Jordan Wilson [00:36:00]:
Specifically, they're keeping out, these benchmarks comparisons like MMLU, and then, the m, the m l, the the the multimedia version one. Right? There's essentially kind of, I think it's MMLU and MMLU pro. Now I'm blanking on it, but it's kind of the standard. It's been this golden benchmark, but, I mean, you can see it here. Anthropic is just kinda like, nah. We're good. We're just gonna stick to these more technical, benchmarks. Right? Visual oh, there we go.

Jordan Wilson [00:36:31]:
MMMU. You know, it's with non extended thinking at a 71%. It's not better than OpenAI. Right? OpenAI, on the MMMLU, which is the multimedia version of the MMLU, which I would say is the standard or has been the standard, benchmark. OpenAI is better than it. Right? So it's interesting to see here. It doesn't look like Anthropic is trying to overfit for certain benchmarks. Right? And I would like to see, once there is the MMLU and not the multimedia version of it, where Anthropic's new quad 3.7 stands because I'm guessing it is gonna be not in first.

Jordan Wilson [00:37:14]:
I'm guessing it might not even be in the top five, but I don't think that anthropic necessarily cares. Because like I said, it looks like they're just trying to compete, and they're trying to be more of a just a coding assistant. Right? So maybe their biggest competitors might also be some of their customers like Cursor, like Windsurf, like, Lovable, like Bolt. Right? Or maybe some of their competitors might be GitHub Copilot. Alright. Let's talk a little bit about Claude Code. So, yeah, live stream audience. Let me know.

Jordan Wilson [00:37:47]:
Should we tackle this, at a later point? I think it's pretty cool, but you you have to have a little bit of techno now. So here's how it works. So Claude Code, essentially, you you go to GitHub, you kind of install this GitHub repo, and then it can work with a code base on your computer. So, you know, if you're on a Mac and you open Mac terminal, essentially, you can have, Claude code. So this is a new, essentially research preview that's free, that's even for free users can use, which I think is great. So you can work with an entire code base. Okay? So let's say you have a folder. Alright.

Jordan Wilson [00:38:33]:
So nontechnical people bear with me, and I'm probably gonna get some of the technical details wrong here. So, you know, if if if you are a coder, bear with me as I explain it to a nontechnical audience. But let's say you have a code a code base. So you build an app, or something, and you have a folder and there's, you know, seven different files in there. You know, maybe there's a a JavaScript file, maybe there's an HTML file, a CSS, etcetera. Right? So the cool thing about quad code, well, number one is it works locally, on your machine, right? So you don't have to go into a a third party environment. You're just working in the terminal, which I know might be intimidating for some. But then you essentially just talk to Claude like you would as if you were inside Claude's Claude dot a I.

Jordan Wilson [00:39:19]:
Right? And then it can code and it can update your entire code base. So it will search, edit, test, and push code from within the terminal. So it's not gonna say, oh, here. Here's the new code for the HTML. Here's the new CSS code. Here's the new JavaScript code. Go copy and paste this. Right? It just does it all for you.

Jordan Wilson [00:39:40]:
It works with and updates your entire code base. It has GitHub integration. It's pretty good at debugging. So Claude Code was part of this, you know, 03/07 Sonic release, and I think it may end up being more impactful than the model itself because I think this signals infropics shift to really wanna compete more in that space. And you might be wondering why, and I actually don't hate it. Because if you listen to our 2025, AI, predictions and road map series, one of those things is nontechnical people are going to be spinning up apps for themselves to use. And now quad code is might be the easiest way to do that. Yes.

Jordan Wilson [00:40:25]:
You can, you know, use cursor. You can use, Windsurf. Some of these other tools, I think the, I think the learning curve might actually be a little higher. But quad code can allow everyday people to just go create apps, talk with it. You can even be like, yo. I have no clue what this means. Explain it to me. Or, hey.

Jordan Wilson [00:40:45]:
Make it prettier. Make it shinier. Make it more useful. Right? You know, make a a a data visualization. Right? You can just dump all your data, give it to Claude via this Claude code, create a program that runs locally on your computer that helps you solve something. Right? I do think enterprise software, if I'm being honest, it doesn't have the same future that it has today. I do think everyday nontechnical people are gonna be using AI and large language models to spin up their own software for very niche use cases, and I think Cloud Code might be that first big step, toward bringing that to everyday people. Right? Yeah.

Jordan Wilson [00:41:22]:
You might have to get used to, you know, here's what a GitHub repo is, you know, but it does it all working with your entire code base where, yes, I love using, you know, o3 minutei or o one pro or something like that, but then you still have to copy and paste, you know, all of those, all of those, different files. You might have to use something like replit to run it. So quad code, pretty cool. It does it all kind of for you. Alright. Let's look live, shall we? What could go wrong? What could go wrong doing a live test of a brand new model that has terrible, terrible limits? Let's try anyways. You guys say you like these live tests, so let's go ahead and do them. Live stream audience, let me know if you can see my screen here.

Jordan Wilson [00:42:23]:
Alright. So here's a couple of things to keep in mind. When you are choosing Claude, make sure you are using Claude three point seven SONNET. Also, you'll see this new thinking mode. So it's it it it's kind of ironic. You still have to have this extended, and you're only gonna see this, on the paid plan. You wanna make sure you have that extended box checked. Okay? So you can choose a normal thinking mode, and this is as a front end user, or you can use the extended thinking mode, and this is best for math and coding challenges.

Jordan Wilson [00:43:04]:
Alright. I'm gonna go ahead. I'm gonna put a giant prompt in here. Alright. You okay. Thank you. Marie is always the first to say, yes. I can see your screen.

Jordan Wilson [00:43:17]:
Thank you, Marie. I always appreciate that because I never know. Alright. So I have a giant prompt I'm gonna put in, and we're gonna use this extended thinking in quad three seven. Alright. So here's what it is. I've done this on the show before. This is what I did, when I first tested o one pro.

Jordan Wilson [00:43:39]:
Alright. So essentially, I'm saying these are my podcast stats. So I I'm using the same exact prompt. So I say today's date is January 16. This is when I did the o one pro show. I wanna have a consistent comparison across, quote, unquote, reasoning or hybrid models because that's what we're trying to do here. We're trying to say, like, okay, how is this how is this model? Right? And you'll see here at livestream audience, it's working. I'm gonna try to keep my eye, on the model here.

Jordan Wilson [00:44:09]:
I'll actually just let you watch it, and I'll read the prompt that I put in. Alright. So I say these are my podcast stats. Keep in mind, today's date is 01/16/2025. For all questions, always exclude the top 2% and bottom percent of episodes, unless otherwise noted. So then essentially, I have I give it a series of 11 questions, and these questions are extremely specific. Then I copy and paste, I believe I give it let me count here about data from a 50 podcast episodes. Okay? So this has the name of the episode, the episode number, then it has the, the number of downloads in the first or sorry, the last seven days, the last thirty days, the last ninety days in all time downloads.

Jordan Wilson [00:44:58]:
And then over the course of these 12 different questions, I am asking, in this case, you know, three, Claude, 3.7 SONNET with the, extra thinking, right, with the extended thinking. I'm asking it some very advanced questions. Alright? 13 of them. So as an example, you know, question number two, I say give me the complete list of all episodes with a new performance percentage of over or under the adjusted average because I'm saying take away the top 2% and the bottom 2% of episodes because sometimes there's anomalies. Right? And I don't really care about those. And so I'm saying, hey. Find trends. And then I'm saying question three.

Jordan Wilson [00:45:43]:
Give me top 10 and bottom 10 episodes in their respective percentage that they're either over or under the adjusted average. So what I'm trying to do is, you know, sometimes there's episodes that kinda go viral. Sometimes there's episodes that for whatever reason don't get, like, any downloads. And I'm like, okay. There must have been a a problem retrieving data. So I wanna find kind of that that median or mean, and then I wanna find types of episodes that get more downloads than that kind of, adjusted average. And then I want to over the course of all of these questions, I'm asking it to find different trends and patterns so I can create better episodes for you all. Right? It might be something as simple as how do I name these episodes better, right, and having it spot different things.

Jordan Wilson [00:46:27]:
It could be I'm asking some questions about days of the week. Right? So as, as an example for, question four, I'm saying for the top 10 episodes, above the adjusted average, please suggest three slightly adjusted title names for each if I were to rerun them. Right? Yeah. Probably a couple times couple times a month, I'll rerun shows. You know, I might get sick or, you know, a guest might, you know, have to bail at the last minute, and I might need a episode to rerun. So I'm saying, hey. Give me a new title. Alright.

Jordan Wilson [00:47:00]:
So it looks like let's see. Okay. So it looks like it says it thought for twenty three seconds. It looks like we're so it looks like it's done thinking. Let me see. Okay. So when I go through and I look at this thinking, this does not look like the raw chain of thought. Okay? So it says, I need to analyze podcast stats from the provided data.

Jordan Wilson [00:47:30]:
The first step is to extract the data and organize it in a way that makes it easier to analyze. Let me go through the instructions carefully and understand the task. Right? And then it breaks it down into six, different subsections, and then it says, let me first extract the data. So it's going through. It's kind of showing its step by step, but I'm looking at this. If this only thought for twenty three seconds let's see. Okay. I'm trying to see if there's more there's no more chain of thought.

Jordan Wilson [00:48:01]:
Okay. So it did get through this fairly quickly. It is still answering the questions. Alright. So I'll have to go through and give this a good scrub, but I'm kind of surprised that it only thought for twenty three seconds. And you know what? Hey. If you share this episode, I'll share the complete stats and prompt that I sent. I will share the exact output, from, that we got here from Claude because it's still going, and I will share the exact output that we got from o one pro as well.

Jordan Wilson [00:48:39]:
So if if you really wanna dive into the details, I'm not gonna have time. It would take another half an hour, to read all of this. I'm gonna go ahead offline once this is done and look at the comparison. But I will say this. Overall, it looks like it did a decent job. Although, I did look at the responses I got from o one pro this morning. The responses from o one pro were exponentially more impressive. They were.

Jordan Wilson [00:49:11]:
Right? The findings here let me just go ahead, see if I can read let's see if I can read, one or two of these answers, that maybe we can, have a little more nuance. Right? So let's say number seven. Okay. So the question for number seven was how does release day impact episode performance? Please exclude Mondays as that is usually our AI news that matters days, and we don't usually run any other type of shows on those days. Four and then I say, you know, here's here's what today is so you don't get confused. So it says impact of release day on episode performance. So it says Saturday, which I don't know why they're Saturday because we, as far as I know, have never released an episode on Saturday, so that's a little weird. So then it says Wednesday is 6% above adjusted average.

Jordan Wilson [00:50:08]:
Friday is minus 3%. Tuesday is minus 4%, and Thursday is minus 5% below adjusted average. I'm not sure if that is true. Right? Because we didn't release any episodes on Saturday. So unless there was some, weird thing in the formatting, Saturday shows should not be there. And if so, maybe it was a Friday like, one Friday show that got posted super late. I don't know. But this I mean, in short, this data does not look correct.

Jordan Wilson [00:50:42]:
It does not look correct that, three of our weekdays are below the average and only one of them is above the average. Doesn't make sense. And then it says key findings. Wednesday episodes perform particularly well for technical tool guides and platform specific, content. And here it says Saturday episodes, which again we haven't done, perform well likely due to less competition and more listener leisure time. Thursday is consistently the worst performing episode day, especially for industry specific content, which I know is not true, because I look through, daily downloads every single time. Thursday is not a bad day. Thursday is usually, our second best day.

Jordan Wilson [00:51:29]:
It says Tuesday episodes underperform, especially for news or recaps. Also false. So, you know, I'm gonna have to go through and and look a little bit, but not great responses. And I'm wondering, this this should have taken, I think, many minutes, many minutes. And at least it says here that it fought for twenty three seconds, which number one doesn't seem right. But it doesn't look like I got great results if I'm being honest. Right? It looks like it did go through and answer all the questions, which is good because in my first testing of this last night, it actually stopped. It only answered the first three questions, and then it essentially said that it it went through the the the context window.

Jordan Wilson [00:52:14]:
Right? Which didn't make sense because I'm like, yo. This is supposed to be, you know, a hundred, you know, a hundred some thousand, context window. So here's this one. Here's the one that I did previously. And at the bottom, it says, Claude hit I know it's kind of small there. It says Claude hit the max length for a message. And has paused its response. You can write continue to keep the chat going.

Jordan Wilson [00:52:39]:
And you'll see in this use case here let me go to the top. In this one, it thought for three minutes and eight seconds. Okay? Why did in one use case, it thought for three minutes and eight seconds and could not give me the entire output? Yet in the second one, it says it only thought for twenty three seconds, and it gave me the entire output. And I will have to do a little bit more offline comparison, but, you know, mixed bag mixed bag so far. Alright. So let's do this. I have a very short rubric that I normally do for reasoning, models. I'm gonna go through this one quickly.

Jordan Wilson [00:53:22]:
I'm gonna make sure that I have the extended thinking on this. Alright? So let's go ahead. Go through some of the questions that we would normally run. Alright. So this one, I'm saying I just woke up with six apples and three bananas. Some of these I made up. Some of them are, you know, kind of widely used across the Internet. Some are just modified from, pretty popular ones.

Jordan Wilson [00:53:46]:
So this I like, I know I need to make an actual, like, reasoning rubric, but these are just some that I generally use. So I said, I just woke up today with six apples and three bananas. Yesterday, I ate a banana and two apples. This morning, I will eat one apple and no bananas. However, I don't really like apples, and one banana may turn brown tomorrow. Assuming nothing else changes, how many apples and bananas will I have tonight? So, let's see here. It says it thought for five seconds, and I can go through and look at the chain of thought. Again, I don't know if this is raw.

Jordan Wilson [00:54:19]:
Maybe this is the raw chain of thought, not the summarized, chain of thought. So let's see. It says, let's work this through step by step. Starting point, six apples, three banana. Yesterday, the person ate, then it says this morning, they will eat. The question asks how many apples and bananas they'll have tonight after eating what they describe. So So then it's going, let's calculate, and then it says, wait. I need to double check the wording of the problem.

Jordan Wilson [00:54:43]:
The person says, I just woke up today with six apples and three bananas. So these are their current quantities after whatever happened yesterday. So, yeah, Claude gets a lot of the information that I put in here is just meant throw a model off. Most of these models, including, Claude three five SONNET do not get these questions correct. I'm assuming that, Claude three five SONNET with thinking. Got it correct. Yes. It did.

Jordan Wilson [00:55:05]:
The correct answer, right there. Got it correct. It is, five apples and three bananas. Alright. Let's do a couple more. We're gonna go through these quick y'all. Alright. Same thing.

Jordan Wilson [00:55:22]:
We have extended thinking on. Alright. So this one, a man and his dog are standing on one side of the river. There's a boat with enough room for one human and one animal. How can a man get across with his dog in the fewest number of trips? There we go. The man and the dog can cross the river together in just one trip. So even some of the original, you know, very powerful state of the art models would always get this wrong. It's very simple.

Jordan Wilson [00:55:53]:
Right? These these questions are simple. Any human knows right away, oh, that's one trip. For whatever reason, a lot of large language models, including, you know, GPT-4O when it first came out, Claude Sonnet three five would get that wrong. Alright. Here's another super easy one. Let's go ahead and, ask this one. I'm saying if it takes three hours to dry 10 t shirts in the sun, how long will it take to dry 30 t shirts in the sun? Alright. So here we go.

Jordan Wilson [00:56:28]:
We got a little, little bit of chain of thought here. Not a ton. Three seconds. Correct. It's saying it will still take three hours to dry 30 t shirts in the sun because the process is parallel. Right? So as long as you have room, you can dry as many as you want. Alright. Our next question.

Jordan Wilson [00:56:47]:
Let's see if, we can trip it up here. So I'm saying if you have a single match and you walk into the room with an oil lit with an oil lamp, a candle, and a fireplace, which do you light first? The answer, which, Claude got correct here, you would like the match first. Alright. I hate how jumpy this sidebar is. Alright. Next question. I don't think any large language model has ever got this one wrong. What color is an airplane's black box? Alright.

Jordan Wilson [00:57:19]:
It is actually bright orange. Correct. Alright. Here's one that a Claude model has never gotten correct. And let's okay. I'm asking for jokes. Also, a model has never really been good at humor. Alright.

Jordan Wilson [00:57:35]:
So I'm saying, please give me seven jokes that end in the word blue. Two should be about animals. Three should be about some other topic in the body of this chat, and you can make up the other two. Alright. So, again, seven jokes, they should end in the word blue. Two are about animals. Three should be about something else in this chat, and it can make up two. So, most models can at least now follow the basics of the directions.

Jordan Wilson [00:58:01]:
Some models still struggle, but let's see if they're actually funny and if it follows directions. So only thought for three seconds, and it says here are the seven jokes. So let's see. The two about animals. Why was the fish such a great musician? Because it knew how to play the scales and always stayed in the deep blue. Is that a joke? Not really. Is it following the directions? Yes. What did the sad elephant say at the waterhole? Even my memory feels blue.

Jordan Wilson [00:58:31]:
Not funny, but it's there. Alright. And then we are gonna have three about the context, from the context of this chat. So why did the banana go to therapy? Because it couldn't stop feeling peeled and blue. Alright. Why did the boat stay on the river after failing to cross? You're making me feel blue. Why did the t shirt refuse to dry in the sun? It was feeling too blue. So it's using the same, you know, oh, it feels blue as the ending of the, I guess, the punch line.

Jordan Wilson [00:59:01]:
It's not even really a punch line. Alright. And then two of the ones where it made it up themselves. Why did the painter refuse to use any other colors? He was going through a phase where he felt only blue. And then what did the match say after lighting the candle? So interestingly enough, it used another one from the context of this chat even though I told it to only use three. It ended up using four. So from an instruction following standpoint, I guess it I I said you can choose, so it just chose to use another one from the context of this chat. It said what did the match say after lighting the candle? That was my moment to shine.

Jordan Wilson [00:59:37]:
Now I'm just blue. So are they funny? No. Did it follow the directions? Yes. It did. Alright. I think this is going to be, the last one we do here just due to time. So, a box is locked with a three digit numerical code. All we know is that the digits are different.

Jordan Wilson [00:59:56]:
The sum of all digits is nine and the digit in the middle is the highest. What is the code? Alright. So now we can look at the chain of thought. It's going through. It's saying let's work through this step by step. It's assigning numerals. Let's call the digits a, b, and c. A does not equal b does not equal c.

Jordan Wilson [01:00:16]:
Alright. And then it says b is greater than a, b is greater than c. Alright. So it's breaking it down the way that a human probably would. It's going through some potential use cases. It's throwing some numbers out to see if they work. Let's scroll to the bottom here. So it's going pretty fast.

Jordan Wilson [01:00:33]:
So that's good. Even though the model's new, probably a lot of people are hitting, are hitting it right now. It's going pretty quickly. Alright. Let's see if it actually gets it correct. Here we go. So it's saying, actually, wait. Let's reconsider the constraint.

Jordan Wilson [01:00:53]:
All digits are different. Does this include zero? Yes. So many models skip over zero for whatever reason. When they look at numbers that would be on a padlock, they only think one through nine, but zero would be there. Alright. So now it's saying, I'm actually sure zero is allowed since the valid digits for a code are typically zero through nine, but zero does result in non uniqueness. There we go. Alright.

Jordan Wilson [01:01:23]:
Let's scroll to the bottom. So this is actually, aside from the first version, of the podcast one, the podcast stats that I did, I can't even remember. That was either late last night or early this morning, that caused it to think for three minutes. This is the one where it's taking the longest to think. And the chain of thought is pretty impressive. It's doing, seemingly a pretty good job here. And luckily, I haven't hit my rate limits yet. What a miracle.

Jordan Wilson [01:01:55]:
But I intentionally did not use it a lot, just so I wouldn't hit my rate limits. Alright. Live stream audience, we're gonna give this one a second to finish up, but what are your thoughts here? What are your thoughts here as we get ready to wrap? Big bogey is saying, I have Gemini two tell me a joke every day. It's not getting better. Yeah. These large language models are definitely, not good. Although, Denny is saying that AI makes dad jokes look good. Alright.

Jordan Wilson [01:02:27]:
Let me see. I wanna make sure if you did have any questions, let me just double check. I wanna make sure. I don't see any specific questions, but I do have a couple dozen. Okay. Here we go. Woozy saying and for our podcast audience, SONNET three seven is still thinking through this combination problem. So Woozy is asking, is there anything that you see in the thinking that makes you adjust specific things in your original prompt question? Woozy.

Jordan Wilson [01:02:58]:
Thank you for that question. That's an amazing question. Yes. A %. And I've mentioned this multiple times, and I called this out, I think, twice when I did deep research, when I did my deep research comparison show. You should be doing this all the time because when you're using deep research as an example, it shows you what it does. I think deep research is one of the best use cases for generative AI, FYI. But I always go back and I always say if you're gonna use a model like a deep research or a reasoning model that takes its time to think, you might as well squeeze the juice out of it, and get a good return on your time invested and go ahead and do it again.

Jordan Wilson [01:03:37]:
In these cases, when there's a finite answer, right, the number of apples and bananas, I'm not gonna do it. Right? I'm not gonna run that a second time. There's a yes or no answer. When it re when it's, more of putting it on a task that there's not it's not as finite, %. Thank you for that question, Woozi. You should always, always, always look at the chain of thought. This is the biggest one of the biggest advantages to having both chain of thought using using these reasoning models as well as being able to see exactly how these deep research tools research is you go through, you read it, you see, oh, this was a good decision that it made. Oh, this was not a good decision.

Jordan Wilson [01:04:19]:
I take notes on the side, and then I adjust my original prompt. You need to be doing that all the time. Woozy. You just added a ton of value, to our audience here. Let's see. Douglas asking says, I think hybrid is interesting. Pure transformer is missing the benefit of the thinking. The thinking is slow compared to the transformer.

Jordan Wilson [01:04:41]:
Hybrid could be a good bridge, but what is being sacrificed for the capability? Yeah. Yeah. So quad three seven SONNET is the first hybrid model. Well, what's being sacrificed here is fine tune control. Right? Especially for front end users. Yeah. You you you have a little slider on the back end, if you're using the API. But like I said, I personally don't like this.

Jordan Wilson [01:05:05]:
Right? But I'm a power user. Right? When I I love, which I know I'm in the in the minority, I love logging into ChatGPT and seeing, like, eight different models. And there are sometimes I know I'm going to o one pro. There's times that I know I'm going to o3 mini high plus web. There's times I know I'm going to deep research. There's times I know I'm going to g p t four o. Right? And I'm going there with intent because I know the pros and the cons. Does the average user need seven models? Probably not.

Jordan Wilson [01:05:35]:
Will the average user benefit from a hybrid model? Probably. But it does. I don't care what anyone says. For front end users, a hybrid approach lowers the ceiling. Right? Because there's gonna be times that it uses, thinking. It uses extra compute when you don't want it to. There's gonna be times, when conversely, right, when it does the opposite. So I think what this ultimately means for power users, you're gonna lose some flexibility.

Jordan Wilson [01:06:08]:
And for everyone, I don't care what you say. I think it lowers the ceiling just a little bit. The floor goes up. The floor goes up. The ceiling comes down. So like I said, for the average everyday user, I think hybrid models are great. For power users that are using it on the front end, I don't like it. I don't like it.

Jordan Wilson [01:06:28]:
And on the back end, the companies are gonna make much more money. Right? I think you're gonna be paying more if you're using a hybrid model. And like I said, I hope, hope, hope that all these, companies, even five years down the road, are still gonna have these non hybrid options. Because if there's only hybrid options, regardless for that very reason, especially when you're using the API, If if if you're a software company, let's say. Right? Or you're just using, you you know, OpenAI or or or or Claude for customer, you know, customer service. Right? To to get to support tickets faster, you use some rag. You put in your company documentation, and someone's chatting, with an AI chatbot with your company's information, but using one of these models. Right? If there's only a hybrid model and those hybrid model costs are much higher and you don't in the future, you don't have an option to use like a a GPT four o Mini or a or a Quad three five Haiku, and you only have the hybrid model, your costs go up.

Jordan Wilson [01:07:32]:
Your costs go up exponentially. Like I said, I think we've been getting a steal for the last couple of years. Alright. Let's go ahead and look at this, result and wrap this show up. So let's see. How long did this one think? It did finish, here a minute ago when I was on a side tangent at, answering, Wuzy's question. So this one, interestingly enough, thought for three minutes and fifteen seconds. So let's see if it got it right.

Jordan Wilson [01:07:58]:
So I'm gonna go past, the chain of thought. Alright. It says to solve this problem, I need to find the three digit code where all digits are different. They the sum is nine, and the middle digit is the highest of the three. So it says, 180, 2 70, 3 60, 4 50, 1 60 2, 1 50 3, and 243. Okay. And it says, since the problem asked for the code implying a single answer. So, yeah, I did say what is the code, but it should know, as most of the reasoning models do, that there's actually many answers.

Jordan Wilson [01:08:44]:
So I would not give this a pass on this necessarily because even though it gave me other options that worked, right, one eight zero works, two seven zero works, three six zero works, it ultimately chose a single answer, which I don't know. It says one fifty three is the smallest valid code that satisfies all the conditions. No. All these two. But there's tons of others. There's tons of other, codes that work. So as an example, 180 works. 243 works.

Jordan Wilson [01:09:20]:
What about 342? Right? So it didn't do a good it didn't do a good job. It fought for a long time. I would say that it did not pass this. So it is kind of a trick question. Even though I asked for the code, there's many codes. And it thought that, and it found that out with its reasoning, but still decided to almost overthink this issue, and say kind of just the wrong answer. Alright. I know this was a long one.

Jordan Wilson [01:09:48]:
I hope it was helpful. Like I said, if this was helpful, go ahead share this. I'll share that complete prompt, the one going over the podcast stats, and I'll share o one's complete answer as well as Claude, three seven SONNET's complete answer. So, yeah, if you are interested in just the, kind of the raw thinking capabilities, go ahead, share this. I hope this was helpful, but I'll tell you I'll tell you my quick takeaways. Claude three seven SONNET, amazing. I think it is going to dominate and continue to dominate, for any companies that need software engineering, coding, etcetera. I think even the three five new model in many of those use cases was already better.

Jordan Wilson [01:10:32]:
So it was already making kind of one of the best models in the world exponentially better. So for three seven SONNET, coding, software development, etcetera, through the roof. Even what this means for artifacts, huge. Right? I'll probably do, a dedicated episode just on three seven SONNET, artifacts and what that means even for nontechnical people. But for everyone else, for API prices, I think it's a big loss. I think overall, hybrid model, like I said, brings the brings the floor up, brings the ceiling down. I don't think hybrid models, at least right now, are great for power users who are using this on the front end. I actually think I will probably be using, Claude three point seven SONNET, maybe a little less, and I'll probably just be using Claude three five SONNET a little more.

Jordan Wilson [01:11:23]:
Right. And I know that seems backwards, but that's probably the reality. So I think that there's some, some super promising aspects of this new Claude three point seven SONNET, some highs, some lows, but hopefully, this episode was helpful. Alright. So like I said, if this was, please share this. Also, if you haven't already, please go to youreverydayai.com. Sign up for the free daily newsletter. We're gonna be recapping this episode.

Jordan Wilson [01:11:47]:
Maybe you missed something. You know, I know this was a longer one. Trying to explain things live is always a little time consuming, but that's what a lot of people that I hear from like. So speaking of that, go subscribe to the newsletter. Reply to today's newsletter if you're still listening, and tell me what you wanna see more of. This show actually was because of you. I put a pullout, on my LinkedIn. It was actually was only decided by one vote.

Jordan Wilson [01:12:12]:
So we're gonna have the other winner, which is how to prompt o models, o one and o3 models. We'll probably do that show tomorrow or maybe next week. We'll see, if time allows for it. So thank you for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI