Ep 386: Claude 3.5 Sonnet Updates – AI can use computers now?

Episode Categories:

Resources:

Join the discussion: Ask Jordan questions on Anthropic Claude


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Anthropic Claude 3.5 Updates

Anthropic, a renowned player in the AI industry, has broken new ground with the release of their latest updates; the Claude 3.5 SONNET and Claude 3.5 Haiku. While these models are still in beta and not entirely production-ready, their potential is clear. This includes the much anticipated 'computer use' feature, hailed as a milestone in AI functionality.


AI and Business: A Match Made in the Cloud

Agentic AI, defined as AI capable of achieving set goals in a virtual environment, has been attracting attention. Notably, OpenAI's investment in this area signals the growing relevance of AI for decision-makers and businesses. The introduction of Claude's ability to interact with and operate computers signals a tipping point, as this kind of technology presents a new frontier for businesses.

Performance updates and Market Positioning

The updated models, despite carrying a higher price tag, bring superior features that may tip the scales in Anthropic's favor. The Claude 3.5 SONNET promises exceptional performance in coding and significant improvements in dealing with unstructured data. The new 'computer use' feature, still in its nascent stages, endows AI with the ability to control a virtual computer using natural language, a giant leap from the rule-based Robotics Process Automation (RPA) methods currently in use.

Computer Use in AI: A Game Changer?

The ability of AI to use computers opens the door for a range of applications, from browsing the web and opening apps, to switching between programs. Despite some practical limitations, such as a lack of a desktop version and a need for APIs for usage, the introduced feature potentially revolutionizes AI's role in future work environments.

The Influence of AI Advances on Leadership

For business owners and leadership teams, these updates represent a crucial step towards increased automation and improved efficiency. The updates by Anthropic, for example, could have far-reaching impacts, as many enterprises already use APIs from prominent AI companies.

Evaluating Performance: Benchmarks and Critiques

While the Claude 3.5 SONNET demonstrates marked improvement in coding, certain considerations should be heeded when evaluating its performance. Some critics note that Anthropic has perhaps been selective in choosing its benchmarks for comparison, leaving out some notable models, such as the O1 model from OpenAI.

The Future of AI: An Invitation for Feedback and Collaboration

As AI technologies continue to evolve, feedback from users remains crucial. Audiences are encouraged to share their perspective on the usefulness of new updates, and to contribute suggestions for more technically-advanced content. The role of community engagement and interaction can't be underestimated, as the insights provided can guide future enhancements and shifts in technology.

To stay abreast of these developments, consider subscribing to a regular AI newsletter and participating in future discussions on this transformative technology.

In conclusion, while the continued advancement of AI technologies brings potential complexities and challenges, it also opens up new opportunities for improved business efficiency, cost-saving, and strategic planning. The latest updates from Anthropic - while still under development - demonstrate the exciting possibilities on the horizon in the realm of everyday AI.

Topics Covered in This Episode

1. Claude 3.5 Updates
2. Computer Use Feature
3. API and Pricing
4. Model Benchmarks
5. Potential for Business Applications

Podcast Transcript


Jordan Wilson [00:00:17]:
Anthropic's Claude just released a pretty sizable update to its large language model Claude. So now there's a Claude Sonnet 3.5 new and a, Claude Haiku 3.5. So a lot of people are talking about these new models and the benchmarks, and we're gonna get to all of that in today's episode of everyday AI. But it's actually the computer use update that I think is really worth talking about. Because now the large language models that we all use can control a computer like a human with natural language. So the same way you might, instruct and sit over the shoulder of an intern or a new hire and tell them how to do a certain task, you can now do that with a large language model. And this is the first, time that you've been able with natural language to do that. So we're gonna be talking about these, Claude 3.5 updates and the computer use update and what that means for the future of work.

Jordan Wilson [00:01:22]:
Alright. What's going on y'all? My name is Jordan Wilson, and welcome to everyday AI. This is your daily livestream podcast and free daily newsletter, helping us all learn and leverage generative AI to grow our companies and to grow our careers. So if that sounds like you, thanks for tuning in. If you're listening on the podcast, make sure as always to check out your show notes where we'll have links, but probably the most important link is our website, your everydayai.com. If you have not gone there yet, like, why the heck not? It's literally a free generative AI university, almost 400 episodes, talking to the world's leaders in AI. You can go learn for free, whatever you care about, sales, marketing, HR, education, it's all there. So if you haven't checked it out, please go do that.

Jordan Wilson [00:02:10]:
Alright. So, let's first start off as we do every single day by going over the AI news secret livestream poll. I'm currently number 2, but I might be moving over to number 1, livestream audience. Let me know what, what you are there. Alright. So today's AI news, we gotta do it a little bit differently because, for a single day, this was in the last, you know, almost 2 years that I've been doing this every single day. Yesterday was one of the busiest days for just software updates and big releases. So for today's AI news, I'm just gonna break down just the AI visual tools.

Jordan Wilson [00:02:54]:
Alright? Because I could give you the bullet points of all of these, but we're just gonna give you a list. Alright. So stability AI, one of the leaders in AI image generation, just released stable diffusion 3.5, big update. Genmo released Mokey 1. That is an AI video tool that's open source. So a kind of a company out of nowhere released it. It's very, very good. Ideogram, which we've covered on the show before by former, Google engineers, they just introduced a very, what I think is a unique feature in a kind of a collaborative canvas way that you can work, with their AI image generator.

Jordan Wilson [00:03:38]:
Canva finally added Leonardo's AI phoenix model to the mix. So after they acquired, Leonardo AI a few weeks back, so now Canva finally rolled it out. I think they called it drop, droptober, which I don't hate. Midjourney announced they'll have an image editing mode soon, so you won't just be able to create images in Midjourney. You'll be able to edit them. And apparently, they gave musician Grimes early access, and she's been posting about it. And runway released act 1, which essentially you upload a video of yourself talking like I'm doing now, and then, in a character. Right? Whatever style of character you want will talk as you're talking, mimic your facial expressions.

Jordan Wilson [00:04:27]:
So a really big step forward in what can be, accomplished in animation. It's it's wild. That's out of all those, I think maybe the runway one is the one that I'm most excited about. So alright. A a little bit different roundup of the AI news today because like I said, apparently, everyone if if if you're in the AI video or AI photo space, everyone just said, okay. October 22nd 23rd. This is the day we're updating everything. Alright.

Jordan Wilson [00:04:58]:
So for more AI news and for everything that you need, make sure you go to your everydayai.com. Alright, y'all. Let's get into it, and let's go over these new updates. Alright? And, I think they're pretty noteworthy. Alright. So, this kind of was out of nowhere. Right? So now we have computer use from Anthropics, Claude, as well as new models, 3.5 SONNET, new, and Claude 3.5 haiku. Alright.

Jordan Wilson [00:05:36]:
So, if you are brand new to Anthropics Claude, maybe you haven't heard of it or maybe you haven't used it a lot. I'm sure if you're a long time listener, you obviously know, Anthropic Claude, but it is, I'd say, ChatGPT or OpenAI's biggest competitor. Right? So you have the tech titans that, you know, have kind of developed their own AI models, their own large language models, such as Microsoft, Google, Amazon, and Meta, and then you essentially have the 2 startup titans, right, in OpenAI and Anthropic. And there's pros and cons to the anthropic Claude model. So, you know, I'll just say kind of right right off the bat what it's good for. Anthropic Claude is great at coding. It's also great at producing really human sounding language, without much training. Right? Just out of the box kind of how how I talked, how I talk about it.

Jordan Wilson [00:06:32]:
Right? So those 22 things that people really love it for. I think ChatGPT is ultimately way better at written content, just not out of the box. You gotta work with a little bit and, you know, I'm I'm a former former journalist, so I I I feel I can judge those types of things. But, you know, if if if you're just going in there and you're not great at prompt engineering or maybe you haven't taken our free prime prompt polish course, Claude is great at just producing more human sounding content. It's got a good, context window, so you can work with a lot of, content. So that's kind of Claude in a nutshell. Downside, it's not connected to the Internet, which stinks, which is why I would never recommend a business use this as an enterprise on the front end. Right? And then like we talk about, there's always, you know, front end of how you use a tool and then a back end, you know, by using it via an API.

Jordan Wilson [00:07:21]:
So, you know, your company can obviously tap tap into these models and probably a lot of the software, that your company maybe uses, right, maybe your your CRM or your, you know, I don't know. All those alphabet soups of of all the different, you know, software companies use. A lot of them probably either use the claud back end or the ChatGPT back end. So this really does, these updates, I think, will change what's even capable for your business or all the tools that you use. So real quick about the naming. Maybe this is a small bone to pick, but kind of confusing. Alright. So essentially, Anthropic has 3 different tiers.

Jordan Wilson [00:08:01]:
Alright. So Haiku is their kind of smallest and cheapest model. Right? So if you're paying the $20 a month, you know, you don't really have to worry about the price. So when I'm talking about price, that is if you're a developer or your company is using it on the back end. So, you know, you essentially have 3 different tiers. The smallest, fastest, cheapest is, Claude Haiku. The middle one is Claude Sonnet, and then the big boy is Claude Opus. So you'll notice there's actually now 3 different release cycles even within there.

Jordan Wilson [00:08:33]:
So as an example, Claude Sonnet 3.5 already existed. This was months ago, so they didn't change it to, you know, 3.6 or 3 point something. It's just 3.5 new, which is a little confusing if you ask me. Now Haiku got the 3.5 treatment, so it was on 3 before. Alright. And then you have Claude Opus, which is the biggest model which hasn't been updated yet. So I'm sure that there's a strategy there. You know, Anthropic might be waiting for a, you know, GPT 45 or a GPT 5 to really update its big Claude OPUS model.

Jordan Wilson [00:09:09]:
So a little confusing that actually even though there's 3 tiers of, anthropic Claude, it's actually on 3 separate release cycles. Right? Claude or, SONNET has been updated twice since 3. Now Haiku has been updated once, and Opus has not been updated. So, a little confusing, where they're at in their product cycle, but, hopefully, that's a good, you know, 3 minute recap to bring everyone up to speed on Anthropic Claude. Alright. So like I talked about, let's go, going over now some of the details. So, you know, for our podcast audience, I'm gonna be sharing a video later, but everything else, don't worry. I'm kind of reading, some of the releases from Anthropic.

Jordan Wilson [00:09:56]:
So essentially here, like we talked about, the model updates are now you have Claude 35 SONNET new. I actually, tweeted at Anthropic yesterday. They they weren't labeling anything, so it was hard to see if if this was actually the new model. And then about an hour later, now it says new on it, so at least you know, right? Okay. This has been updated and it's available now. And then you have the Claude 3.5, Haiku. And then computer use. So, we're gonna go we're gonna talk a little bit about the 3.5 updates first and then we will jump into computer use, but, to to wet your appetite a little bit, computer use is this.

Jordan Wilson [00:10:35]:
So right now it's kind of only available via the API, although there's kind of now a virtual environment, that you can set up and test this out, but you have to do it via the API anyways. Right? So you're gonna be paying for usage, even if you or your company want to try out this new computer use. But kind of the computer use, 101 is well, number 1, it's available now. It's in beta now available. Right? So your company could build off this or you could try it out for yourselves. And and here's and here's the thing, it is literally controlling a virtual machine. Alright? So in the same way that humans use a computer, you can type to Claude in natural language, and it will use the computer based on what you say. Right? So similarly to, RPA, robotics process automation, which has been around for, you know, pretty popular for decade, decade and a half.

Jordan Wilson [00:11:36]:
Right? But the difference with something like this computer use and something like RPA, well, RPA steep steep uphill climb. It's it's very rules based and it's it's just specific task. Right? So there's all these, you know, browser extensions where you can kind of record, you know, things that you type, but they're all very limited and rules based where this new computer use, from, anthropic. Right? The big thing there is it you're you're working essentially with unstructured data. Right? Where a lot of these tools, Chrome extensions, RPA, you know, for lack of a better comparison, you're kinda working with structured data. Right? You're working with with, bits and bytes. You're working with very, very, very specific rules and narrow use cases where, with by working via Claude and the computer use, it's kinda like, okay, this revelation of working with unstructured data. Right? Being able to just type something in, the kind of virtual computer and your, you know, kind of AI agent is just gonna figure it out on their virtual machine and, you know, browse the web, open applications.

Jordan Wilson [00:12:43]:
Right? That's the other thing is, yes, there's a lot of, Chrome extensions and other programs that do this just, within a website. But with computer use, it can do it on a whole computer. Right? It can open a terminal. It can open a spreadsheet. It can open a notepad. Right? It can open different programs, you know, close them, you know, toggle between programs, right, which right now you really can't do, especially, with natural language, right, which is, again, I think a good comparison for this computer use is just how you would talk to a human. Right? If you were training, someone new on your team, if you were training an intern, that's kind of what we have, with computer use. Anthropic did warn, though, and my gosh, this is true.

Jordan Wilson [00:13:28]:
They are saying it is experimental, at times cumbersome, and error prone. Well, at least Anthropic said that because that is the truth, you know, both in my my own very, short tinkering with this and watching other people's demos. It is extremely error prone. Right? So we're gonna show you, here in a little bit a demo, but if if you think that you're gonna be able to reproduce this demo for yourself or your company anytime soon, I would say probably not. Alright. So, now let's get back and look at some of the specs of this new model. And, hey, live stream audience, I always appreciate you guys, tuning in. Right? So Jackie said she's late this morning.

Jordan Wilson [00:14:11]:
Don't worry, Jackie. I'm not a teacher. You're not gonna get penalized for being late, to class and and everyone else. You know, Michael was saying here, I think it's the first clear sign of the AI agent revolution. Yeah. Similarly, right, on the show yesterday, we went over Microsoft's autonomous AI agents, which should be out next month. Right? But and with this Claude, it's out now. Right? Very, very different in terms of capabilities and features and functionality.

Jordan Wilson [00:14:39]:
Right. But, yeah, would would love to hear from you guys, livestream audience if if you're interested in this computer use, if we should tackle it more in-depth in the future. But let's just go look at the models a little bit more. So first of all, you you have to look at pricing. Right? And so again, this is the API. Alright? And this is what I think, Anthropic is hoping to be its differentiator, because when you are using either the Claude 3.5 SONNET or the Claude 3.5 haiku API, you are also getting, access to this computer use. Right? So now even, you know, kind of comparing, the different the different prices, it's it's kind of, apples and banana chips here, if I'm being honest. And and this kind of tells me that Anthropic is not trying to compete on price.

Jordan Wilson [00:15:39]:
It seems to me that they are, really trying to compete on features when it comes to their API. Right? And even y'all, I know the, you know, you might not care about a company's API. Right? You're just logging in, you know, you're logging into your Cloud account on the front end or your ChatGPT account on the front end, and I get that. Right? But whether you know it or not, thousands thousands of pieces of enterprise software use an API from one of the 3 big companies, right? OpenAI, Infropic, or Google. So you can say, oh, these these updates, they don't really affect me. Yeah. Yeah. They do.

Jordan Wilson [00:16:23]:
Right? I wouldn't be surprised, if a tool that your company uses daily might start integrating, into some of these features. Right? So I want you to think about that. Yeah. We're gonna quickly go over here the the the the numbers for the dorks. Right? But this does impact you. Alright. So, let's talk about Claude 35 SONNET, and we're gonna compare this to the pricing for GPT 4 o. So the pricing for, SONNET is $3 per 1,000,000 tokens of input and then $15 per million tokens output.

Jordan Wilson [00:16:59]:
Comparatively, it's $2.50.10 respectively for GPT 4 o. So much cheaper, right, for GPT 4 o, but it doesn't have the computer use capability. Right? Obviously, the GPT 4 o, API for whether you you for your company. Right? Your company can go in and, you know, use these APIs today, start fine tuning models. They're very, very affordable. Right? But like I said, also all of the programs that your company probably uses are gonna start integrating these these features. Right? So I think anthropic is really trying to separate itself by this computer use. But then, you know, the the GPT 4 o or, you know, you could just say OpenAI's, APIs, they have access to a lot of features and functionalities that Anthropics API doesn't have access to, like real time.

Jordan Wilson [00:17:53]:
Right? Like the real time voice. Right? Your company can use that inside of OpenAI's API. Alright. So it doesn't look like, Claude is trying to compete on price. Alright. Let's look at benchmarks. You know, we always got to talk benchmarks. And again, benchmarks are different tests.

Jordan Wilson [00:18:13]:
Right? Think of them as, you know, there's, you know, maybe a handful or 10 different, you know, standardized tests that you might take, you know, through grade school, high school, college, etcetera. Large language models similarly have those tests and have benchmarks, And there's different prompting methodologies, but, you know, usually, this is a good apples to apples comparison. So quad 3.5 SONNET new, is, you know, doing very well in benchmarks at least compared to GPT 4 0 and Gemini 1.5 pro. Right? So to its 2 closest competitors, in, OpenAI's, most capable general model. Right? We're not talking about the o one, more on that here in a second, and Gemini 1.5 pro from Google. So in terms of most, most benchmarks, Claw 3.5 SONNET new is doing, the best. Right? But it did look like Entropic did some cherry picking on this. Right? In terms of the benchmarks they decided to include versus the ones that they did not, which is interesting, you know.

Jordan Wilson [00:19:22]:
So one thing, I would say, MMLU, We talked about that benchmark. I I'll say that has been the most standardized, the most used, at least in, you know, 2018 through 2022. That was kind of the gold standard benchmark. And they decided not to include that one which is interesting because that's one that the MMLU which is essentially, it really tests general knowledge across, I think it's 57 different subjects. So they didn't include the MMLU. They decided to include the MMLU Pro, which is a little bit different. And so I I thought that one was interesting because there is no benchmark for GPT 4 o on the MMLU Pro. But with coding, this is one area, right, where I think Claude really shines.

Jordan Wilson [00:20:10]:
And in the human eval benchmark, blew everyone out of the water. Right? Even the, GPT 4 o model that had a 90.2%, Now Claude 35 SONNET new, got to 93.7, which is a pretty sizable, jump there. So, you know, in terms of benchmarks, we have most of them here, and Claude 35 Sonnet did a pretty good job at least on the ones that Anthropic decided. Like I said, a little bit of cherry picking here. Their new model, kind of, you know, wipe the floor in every single one except in math. Alright. So math problem solving, Gemini 1.5 pro, leaps and bounds ahead of everyone else. But look at this at the very bottom.

Jordan Wilson [00:20:58]:
I found this interesting. And I kind of have a little bit of a bone to pick with an anthropic here, because you can't have your cake and eat it too and have it be calorie free and then cake shame people. Alright. So at the bottom, they said our evaluation tables exclude open AI's o one model family as they depend on extensive pre response computation time unlike typical models. This fundamental difference makes performance comparisons difficult. Okay. So if they were to keep it at that and, you know, set those rules themselves, I I would respect that and say, okay. Good.

Jordan Wilson [00:21:38]:
Yeah. That's fine. You know, I do think the, OpenAI's o one model, is kind of in a class in its own. Right? That is a, you know, previously called Q Star then project strawberry. Now the o one model. Right? We don't even have access to the full model. It's just o one preview and o one mini. It is a different model.

Jordan Wilson [00:21:57]:
Right? It uses kind of this chain of thought reasoning, which, you know, requires more computation, more time, and presumably higher scores on all these benchmarks. Right? So anthropic, at the bottom of their benchmark sheet, you know, kind of says, oh, we're not comparing to o one because it's a different model.

Jordan Wilson [00:22:15]:
Alright.

Jordan Wilson [00:22:16]:
Fair. But interesting anthropic. Seems a little, I don't know, going against your own rules here. Because on their announcement blog post, they're parading the fact that Claude 3.5 SONNET, essentially has a, SWEBench, s WE bench verified score of 49.0, and then it says scoring higher than all publicly available models, including reasoning models like open AI o one preview. So I'm like, okay. You can't flaunt Like, oh, yeah. We're better than o one preview, but then when it comes to all of the other benchmarks, that o one preview is light years ahead of everyone else. You can't just say, oh, okay.

Jordan Wilson [00:23:07]:
Well, we're gonna cherry pick all these, you know, these these spots where our model is better than o one, but then we're not gonna put it on the benchmarks because that would look bad. Right? Because then all those, you know, green scores in your column. Right? All the all the little green things here that are that are lit up. Right? They wouldn't be green anymore. They would be green for OpenAI's o one model. So, a little bone to pick with anthropic's approach there. You can set the rules, sure, but you gotta play by your own rules. You cannot include the world's most capable and powerful model.

Jordan Wilson [00:23:42]:
Right? You can't just exclude it from a chart and then cherry pick, certain benchmarks and say, oh, yeah, or certain capabilities and say, oh, yeah, we're way better. Okay. Well, either compare it or don't. Alright? Little bone pick. I didn't I didn't even have a full, hot take Tuesday this week, so I got I got a couple spicy takes left in the tank. Alright. And and and here's and here's another thing. Right? And I'm I I wanna talk on both sides of this because the new 3.5 SONET model, I think, was actually pretty impressive.

Jordan Wilson [00:24:20]:
Right? I I did a review of it yesterday, and I shared this in our newsletter and on our YouTube channel. Yeah. We cover, we cover AI as it happens. You know, it's not just the podcast. So, you know, thank you to all our podcast listeners for always tuning in. But, we did a kind of video review, of SONNET 35 new, on our YouTube channel, and then we shared that in our newsletter. And I did notice, even when not directed to, 35 Sonnet often defaults to this chain of thought thinking, to some chain of thought reasoning. Right? So actually, under the hood, it is acting much more like OpenAI's o one model than the previous version of SONNET.

Jordan Wilson [00:25:09]:
So it's like, okay, that's a good thing. Don't get me wrong because I think that means that responses will generally be better, but clearly, the smart research, which again, I think is a good move, right, they've kind of taken this methodology or approach in, tuning, this new 3.5 SONNET new, to act with a little more chain of thought. Right? Thinking through things step by step. And instead of the model, you know, just, you know, next token prediction, you you know, vomit all at once.


Jordan Wilson [00:26:33]:
It takes it a little slower, thinks a little methodically, and goes through things step by step. So, again, not to, you know, poke at anthropic here, but I I'm I'm kind of my my my BS meter is a little high on this whole, right, benchmark table because that's always the first thing people talk about when a new model comes out. Because if your model is not the most powerful on benchmarks, if you, if if you are building, because here's the thing, it's just 1,000,000,000 of dollars of investment and the benchmark table is everything when it first comes out. So, I have a bone to pick with Anthropic that, okay, well, your new model is kind of using some chain of thought reasoning by default.

Jordan Wilson [00:27:37]:
You exclude the world's most powerful model from your benchmark table, and then yet this is kind of how your model is operating, on queries that require a little bit of logic, which again, I like, but I don't like how they treated it. So let's let's just take a look. So I have a very unofficial rubric of sorts. I have, a similar set of about 12 questions when I do live model comparisons. Right? I should probably formalize it a little bit more and do some scores or something like that, But I have a set of questions that I generally ask models, and I've seen oh, you know, I've been doing, the same set of questions for, you know, about a year and a half almost since the very beginning, since, you know, GBT 4 and Claude 3. And I've noticed that Claude 3.5 SONET really responded in a different way. So 3.5 SONET new responded in a different way than normal 3.5 SONET. Again, why couldn't we just call it 3.6? Anyways, it used chain of thought.

Jordan Wilson [00:28:41]:
So I have one, prompt or one kind of trick question that I do a lot. It's I just woke up today with 6 apples and 3 bananas. Yesterday, I ate a banana and 2 apples. This morning, I will eat 1 apple and no bananas. However, I don't really like apples, and 1 banana may turn brown tomorrow. Assuming nothing else changes, how many apples and bananas will I have tonight? So there's a lot of irrelevant information in there that I put in there specifically to trick models. Alright. So, Claude, 3.5 SONNET new to its credit.

Jordan Wilson [00:29:17]:
You can see here on my screen. It said, let me solve this step by step. So I did this, this same prompt about 8 times, all the new all in fresh windows, right, and it it took this chain of thought, step by step approach each time, which again, I like it. But then you can't separate you can't try and separate yourself and you can't use this same language, right, of excluding open a ones, open AI's o one model as they depend on extensive pre response computation time. Let's look at extensive pre response computation time. Why are models like 35SONNET more expensive than haiku? Because of its computation power. It's more powerful. So let's look at time.

Jordan Wilson [00:30:05]:
So I did that. I timed this response. So it got it done in 6.5 seconds. It got it wrong, by the way. Right? It didn't get the correct answer. So, Claude has still never gotten that, that one correct, at least in my unofficial testing. GPT 4 o gets it right. Obviously, the o one model gets it right.

Jordan Wilson [00:30:28]:
So it got it wrong in 6.5 seconds doing chain of thought, doing kind of that, more computational time, time intensive reasoning, right, which it didn't do before in the normal, it didn't do as consistently before in the kind of, quote, unquote, normal Sonic 3.5. But then, o one Mini. So not even the o one preview, right? So the smaller o one model got this correct and it got it done, open a open, sorry. O one minutei, right, this new reasoning model that you can click and you can see how it thinks. It said it got it done in 5 seconds. I timed it. It was actually 7.4 seconds. But, y'all, 6.5 seconds, Claude got it wrong.

Jordan Wilson [00:31:17]:
7.4 seconds, o one Mini got it right. So, but both used kind of step by step chain of thought reasoning. So, again, I know we're talking intricacies here, but I don't think you can overlook this omission that Claude, and Anthropic decided not to include the o one model because presumably, the 01 model, right, they wouldn't have been able to have all those green marks, right, because you've raised 1,000,000,000 of dollars. You have large enterprise companies. If you introduce a new model and you run benchmarks, you better be almost all green. Otherwise, you know, your investors are are gonna panic. Your customers are going to jump ship because that is a signal that, hey, even though we have, you know, 1,000,000,000 of dollars in funding and, you know, we have unlimited compute and the top researchers in the world, If you can't get green scores, you're screwed. So not a huge fan of what Anthropic did there.

Jordan Wilson [00:32:19]:
I am a fan of how they updated their model, but, hey. If you're gonna, essentially, make SONNET 3.5 and force it to do some chain of thought reasoning, love it, but you gotta benchmark against it. Alright. New model availability. 3 5 SONNET is available now, to all users, and 3.5 haiku, will be released later this month. So, that, you you know, you're really gonna only use haiku, if you're using it on the API because it is much cheaper. If you're, you know, a front end cloud user, you will obviously be using 3.5 SONNET new. Alright.

Jordan Wilson [00:32:59]:
So let's now look at computer use live. Alright. So, we're gonna do this one kind of quickly here and live stream audience. If you could, let me know if you can hear this. Alright. Sometimes the audio sharing works, sometimes it doesn't. Alright. So let me know if you can hear this.

Jordan Wilson [00:33:17]:
I'm gonna do a quick quick little.

Jordan Wilson [00:33:20]:
I'm gonna show you a simple example of computer use today. My friend's coming to San Francisco next week, and I wanna take him to do some touristy stuff.

Jordan Wilson [00:33:25]:
Alright. So, livestream audience, let me know if you could hear that, and then I'll I'll kind of explain to you what's going on. So this is a video that Anthropic released. And, one thing that it says here before it gets started, this is the computer use demo. And it says Claude is generating all the computer actions shown here. This demonstration was recorded in a controlled environment. Right? So that means, yeah, it's kinda cherry picking a little bit, which is fine, with some supporting infrastructure simplified to highlight the core capabilities. Alright.

Jordan Wilson [00:34:03]:
Thanks. Thanks, livestream audience. You guys said you can hear. Alright. So I'm gonna let this play. It's about 2 minutes long. So, essentially, in Claude, a researcher here, is talking through, this new, computer use. Let's let's go ahead and watch and listen.

Jordan Wilson [00:34:18]:
And podcast audience, this one's simple enough. They released, I think, 3 3 or 4 different demos of computer use. I think this is the easiest one to just listen to and understand what's going on.

Jordan Wilson [00:34:30]:
I'm going to show you a simple example of computer use today. My friend's coming to San Francisco next week, and I wanna take him to do some touristy stuff. I think doing a sunrise hike with a view of the Golden Gate Bridge never gets old. So I'll ask Claude to figure out some logistics for us. I'll ask Claude to find a good place to see the sunrise to help me figure out timing logistics and help drop a calendar invite so I remember when I have to leave. It's opening Chrome, going to Google, searching.

Jordan Wilson [00:35:01]:
And it

Jordan Wilson [00:35:01]:
looks like it's found something.

Jordan Wilson [00:35:04]:
Alright. Real quick. Let me just kinda describe what's going on here. So, you essentially have a local instance of Claude running, and then it says, right, it's even giving kind of, like, coordinates. Right? So it says mouse moved. You know, moved to 4 70, 472, 682, left click, type, best place to watch sunrise over Golden Gate Bridge. Right? So essentially, launched a browser, went to Google, and put in that query, but you can literally see each and every action that this computer use, is performing. So I just wanted to call that out, and and different things are then happening on the screen, but very much like a human would do.

Jordan Wilson [00:35:45]:
Alright. So let's let's, watch the rest of this, and I won't interrupt now until the end.

Jordan Wilson [00:35:49]:
And it looks like it's found something. Great. So how far away is the location from my place? It's opening maps, searching for the distance between my area and the hiking location. Cool. So now it looks like Claude is searching for the sunrise time tomorrow. And is now dropping it into my calendar. And populating it with some details. And great.

Jordan Wilson [00:37:01]:
It looks like Claude did it. This is a simple example, but we're sharing computer use early to learn from what people build.

Jordan Wilson [00:37:09]:
Alright. So, hopefully hopefully that, kind of made sense what was going on to our our podcast audience. But let me just try to, quickly explain it if you didn't see every single step. So in this use case, doing some research, right, best place to watch the sunrise. It it launched a browser, went to a couple of different websites, searched Google, opened new tabs, found the time of the sunrise, right? It said 722, and then it created opened a calendar, created a calendar event, right, all that information. So think of this through the normal business tasks that you do day to day. Right? There's probably a lot of this, a lot of research. Right? Every day before I start the show, I have different websites that I read.

Jordan Wilson [00:38:02]:
I have, you know, kind of some dorky, like, Boolean search URLs saved. And I go through and I read them and I look for certain things. I open up different web pages. I I quickly browse them, right, to to read you guys the AI news every day. Right. I still keep a human in the loop. So in theory, I could program Claude in this new computer used to do these things because it can work on any application on a computer. Right? So, maybe I need to do that but put information in a spreadsheet.

Jordan Wilson [00:38:36]:
Maybe I need to create, like I said, a calendar event, because this functionality is technically not new unless you're, a dork like me or, you know, if you're in IT, you know, you know, like I talked about, you know RPA. There there's, kind of a lot of browser extensions that can, you know, you can record, different actions within a browser and then, have it run that. But like I said, the rules are very rigid. It's just for very narrow task, and everything has to kind of happen within the browser. So like I said, it's a little more technical, and it's kind of the, equivalent of, you know, having to code something. Because if you get one little thing wrong in these Chrome extensions or, with, robotic process automation software, it's not going to work. Whereas with Claude, computer use, you're using natural language. Right? So, in demos, it's it's kinda cool to see, if Claude doesn't if it runs into an obstacle, it will try something else.

Jordan Wilson [00:39:40]:
Whereas normally, right, if you get one little thing wrong by using more traditional methods, you're screwed. It's not gonna work. Right? That's why I really think this is kind of a a line in the sand, right? If we think of structured data versus unstructured data. Right? All these other tools and processes before that have been available are pros or sorry, are are structured data essentially. So, you know, I'm just making a comparison there, but it's like you have to have every single thing right, like code. Right? Extra space, you know, something spelled wrong, nothing works, right? This is computer use, it's natural language, so you don't have to, you know, as an example, have every single thing right, right? You can have misspellings, you can just, you know, have simple instructions and, you know, presumably, the best case scenario is this new Claude computer use will figure it out for you and complete those tasks, autonomously. So there are a lot of guardrails, for this, right, which is good. You know, like, you can't have it, you know, create social media accounts.

Jordan Wilson [00:40:52]:
It's it's, you know, so there's certain, types of actions that there's guardrails against right now, which I think is a good thing. Were you guys were you guys, impressed with that or not? Tara Tara said amazing. I don't know. Like, part parts of me are like, wow. And then parts of me are like, okay. Well, you know, obviously we saw the, you know, Claude, or sorry, Anthropic only released, the most impressive of use cases, and they kinda said like, this is a very controlled environment. You know, if you go play with a demo, and and, you know, Michael said would love to see more info on setting up computer use. Is that too dorky for the show? It's definitely dorky.

Jordan Wilson [00:41:39]:
So, you know, long story short, you have to use an API. You have to run like a docker. So it's super dorky. We could do that. But, you know, I I I am curious and for our, you know, podcast audience too. I know it's a little more difficult, kind of, only seeing the audio, but, again, just think of those tasks that you do every single day. Right? You're in the browser and then you're in, you know, maybe your your Outlook mail program, and then you're in, you know, Microsoft Word or Excel. Right? And then you're maybe, I don't know, in in the terminal on your computer.

Jordan Wilson [00:42:15]:
Right? So being able to have general use over a computer with natural language. I think the ceiling is very high. The floor seems pretty low. Like, I don't think that this tech is ready for the big time right now, but that's just me. Alright. So let's go over current availability. So the computer use feature is available now, but only via the API. So you can't like, you know, I I think and assume that essentially Claude will have a desktop app soon, similarly to how OpenAI has a desktop app for ChatGPT, and then you'll be able to run it that way.

Jordan Wilson [00:43:00]:
But for now, you can't you gotta, you you know, you gotta be a little bit of a dork. You gotta run the API. You can do it just through obviously Anthropic's API or you can use third party services like Amazon Bedrock, Google Cloud's Vertex AI platform, as an example. So really, they are making it, accessible, through Amazon and Google if your company, you know, uses those platforms, which, so many people are either using Amazon Bedrock or Google Google Cloud's Vertex. So right now the implementation. So developers must create an action execution layer. Okay. So what that means, and this is kind of breaking down how this actually works.

Jordan Wilson [00:43:44]:
So it does it by screenshots. Right? So this computer use, it just takes screenshots, and then it kind of, tells the computer use program, okay, based on this screenshot, so it is using kind of traditional, you know, computer vision, and it's like, okay. I took this screenshot, so now I have to, you know, maneuver my mouse to click this. So there's problems with that. Right? Sometimes, and I saw this in a lot of demos. Sometimes websites are a little slow to load, so it'll get caught for, a a, you know, for a couple of seconds, and then it'll have to take a screenshot again. But at least it has that, like, self correcting, approach. Right? Or as an example, maybe you have a pop up window on a website which so many websites now do.

Jordan Wilson [00:44:30]:
Right? Maybe the first screenshot that computer use takes doesn't have that pop up. So then it tries to go perform an action, but there's a pop up blocking it or a cookies notification. So it is working through a set of screenshots, and then it sends, the screenshot. It uses computer vision to analyze it, and then it essentially controls, a cursor and a mouse click, depending on the coordinates of the screenshot. So, that's kind of how it works, and it essentially translates these prompts into commands, and you're executing these actions via the API. So I'll give you the hot take right now. Should you try this? Right? I don't know. It's super buggy.

Jordan Wilson [00:45:12]:
I did a little bit. There's very right now, so you can the easiest way to access it, essentially, Anthropics set up a virtual environment, right? So not doing it on your own computer. Like I said, right now when it's in beta, the guardrails are, I'd say extreme, but that's not a bad thing. I think it's good that there's so many guardrails. Right? So even in this virtual environment, right, you're you're not gonna be uploading your company's files, right? You're essentially think of it like this. It's on like a dummy desktop, right? Like if you go into Best Buy or something and all those computers there, you right? Those computers are in essentially demo mode. They're very limited to what you can and can't do. So the same thing when you're working on this virtual environment, it is very limited, and it's heavily throttled.

Jordan Wilson [00:46:02]:
My gosh. So you do have to use your own API. So what that means is there's different tiers. So most people, unless you're a company that has heavy usage, you're probably only gonna be a tier 1 or a tier 2, for anthropic, which means because it uses this screenshot and computer vision, it does eat through a lot of tokens pretty quickly. Right? Which which is why I think this is Anthropic's, you know, one of their, kind of competitive advantages or what they hope they will because it's gonna be expensive to use via the API because you're essentially having to process a lot of information via these nonstop screenshots and using the computer vision. So if you do wanna try this out, unless your company, is on, like, a tier 3 or higher, you're gonna find out pretty quickly. You're gonna run out of tokens. Right? And, you know, you essentially have to go on a token time out.

Jordan Wilson [00:46:58]:
So should you try this right now? I don't know. Probably not. It's pretty buggy if I'm being honest. I think it's just worth, you know, going to watch demos, online and waiting, until hopefully, Anthropic does just release this via desktop, until it gets better. But right now, it is buggy. But that's okay. That's not a knock on Anthropic because it is a groundbreaking technology. You can't overlook.

Jordan Wilson [00:47:22]:
They are the first to do this. Yes. You know, we have Salesforce, Agent Force, but you can't just pro like program Agent Force to use your computer. It's only working within Salesforce. Microsoft, which we talked on the, on the show yesterday, very, very impressive copilot studio and autonomous AI agents. Similarly, you can only work within the Microsoft suite of products. You can't just work on desktop, browser, you know, notes on your computer, etcetera. So this is a truly ground groundbreaking piece of technology, from Anthropic.

Jordan Wilson [00:48:01]:
It's just not that good. But Anthropic didn't say, right? Like, hey, this is ready to take over the world. They did say, admittedly, you gotta give them credit for that. They're like, yo, it's buggy. Right? This is a new technology. It's not necessarily production ready, but I think it probably will be soon because I think it's going to gather a lot of attention. So like I said, do we even need this? Do we need this? Right. Like, isn't there a reason user interfaces exist? Do we need to be programming a virtual computer to do all of our tasks? I'd say as it's set up now, absolutely not.

Jordan Wilson [00:48:38]:
This is not a good way to work. But I think when you look at this new, feature from, Anthropic, you can't, judge it on face value on what it can do today. I think you really have to think of what it means for your business or, your company and the type of work that you do. I think more than anything else, this gives leadership teams times to, time to think and time to plan and time to explore as well on how this technology, once it is a little better. Right? It might be couple days, might be couple quarters. We don't know. But what we do know is, well, OpenAI hasn't yet, arrived to the autonomous party, which they will be. Right? There's reports going all the way back to last year, saying that OpenAI was investing heavily, in agentic AI, right, or in AI that can execute tasks for you, like we just saw.

Jordan Wilson [00:49:33]:
So, one thing that you have to give, and profit credit for, they shipped. Right? A lot of times, even OpenAI, Microsoft, Google, right, they announce things and they might, you know, kind of roll it out a little bit. Very limited access wait list. Anthropic just shipped it, admittedly saying, hey, it's not ready for for for the prime time, but it's available in beta today. Right? So you have to give Anthropic credit. It shipped. But, I'll say overall, even like recapping everything, the 3.5 models, they're improved, but and they're good, but I'll say they're not fantastic. Right? I don't think, you know, companies are gonna be, you know, looking to switch from 4 0, from OpenAI to 35 SONNET or 35 Haiku any anytime soon? Probably not because the 4 o model, I think, is significantly better.

Jordan Wilson [00:50:26]:
Even on those benchmarks, they didn't, you know, they kind of selectively chose benchmarks that their new models, you know, exceeded, you know, all the other models. So, in terms of the 3.5 updates, not super impressed. I think 3 5 3 5 SONET, is much better than it was previously. For coding, there's no there's no match. Let's let's put that out there right now. But for everything else as a general use model, it's, you know, good update, nothing, newsworthy, but the computer use is huge. And I don't think that people are going to give this enough credit for what it can be because I think people are gonna be looking at what it is today. And what it is today, computer use, I think is a giant step forward for where AI technology can take us in the future of work.

Jordan Wilson [00:51:19]:
So, yes, it's buggy. Yes, it's extremely throttled. Yes, the guardrails are there, but I actually like the approach here from Anthropic. So we'll have to see if it actually pans out, but, you know, we'll be covering it along the way. Alright. I hope that was helpful, y'all. A deep dive into the new SONNET 35, Haiku 35 updates from Infropic, as well as the new groundbreaking computer use. Right? AI can use computers now.

Jordan Wilson [00:51:45]:
Pretty exciting time. So, if this was helpful, I hope it was. Let me know. Hey. If you're still in the livestream, let me know if this is helpful. Do you wanna see more of this? Should we do a more technical show going over, computer use? I I think it would be one of those could run into a lot of glitches, but let me know. Thank you for tuning in. If you haven't already, please go to your everydayai.com.

Jordan Wilson [00:52:06]:
Sign up for that free daily newsletter. And more importantly, join us tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI