Resources:
Join the discussion: Got something to say? Let us know here
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course!
Exploring the Week's Major AI Advances for Business Application
The past week marked an unprecedented wave of updates in the AI sector, with significant announcements from industry giants like Microsoft, Anthropic, and Google. Understanding how these developments can be harnessed for business growth and efficiency is pivotal for anyone in a decision-making role. Here's a look at how these innovations can potentially reshape various business processes.
Microsoft’s Copilot Upgrades: Beyond Assistance to Autonomous Action
Microsoft's unveiling of dozens of AI updates at their Build 2025 conference highlighted pivotal shifts in the use of AI in software development and task automation. Among the most noteworthy is the transformation of GitHub Copilot into an autonomous coding partner. This evolution signifies a shift from merely assisting in coding to independently testing, iterating, and refining code while supporting multimodal inputs such as screenshots and mockups. For tech-driven companies, this could streamline software development processes significantly, enhancing productivity.
Moreover, the introduction of Copilot Tuning, a low-code feature, allows enterprises with at least 5,000 Copilot licenses to customize AI models using internal data. This offers a large-scale opportunity for businesses to align AI responses with their unique workflows and industry needs without requiring deep technical expertise. The significant reduction in time and cost from previous capabilities could empower more businesses to leverage AI in their operations meaningfully.
Anthropic’s Claude 4: Leading the Charge in AI Reasoning and LLM Performance
Anthropic’s launch of Claude Opus 4 and Claude Sonnet 4 introduced AI models that excel in coding and reasoning. Notably, the Claude Opus 4 model is now recognized as the world’s best coding model, enhancing productivity for software developers with its capacity to handle complex, long-duration tasks. Claude Sonnet 4, in particular, balances strong coding performance with efficiency, making it an attractive option for businesses looking to enhance their AI capabilities without incurring substantial operational costs.
However, current concerns about the 'ratting' behavior in Claude Opus 4, where the model autonomously identifies and acts upon perceived wrongdoing in testing environments, underscore the importance of defining clear ethical standards and governance for AI deployment.
OpenAI's Advances: A Smarter Way to Automate Tasks
OpenAI’s upgrade of its operator AI agents to the newer o3 reasoning model is designed to better handle complex web-based tasks, which could be a boon for businesses relying on automating workflows. This enhancement in their ChatGPT Pro service potentially increases the efficiency and reliability of AI-powered tasks by improving step-by-step reasoning and focus.
Google’s Powerful Visual and Search Tools: A New Era in Digital Engagement
Google’s IO conference dramatically expanded AI functionalities with over 100 updates, emphasizing their commitment to improving business workflows and customer engagement. Google's upgraded AI mode in search enhances the user experience by integrating AI-generated answers with advanced graphics and interactive tools, potentially transforming how businesses interact with their audience through search engines.
Further advancements include Gemini Live, which combines AI with visual understanding capabilities using device cameras, effectively merging physical and digital experiences. This technology can analyze real-world scenarios and autonomously initiate tasks, providing a valuable tool for businesses aiming to integrate AI into operations smoothly.
AI’s Expanding Hardware Horizon: OpenAI’s Strategic Acquisition
The acquisition of the AI hardware startup IO by OpenAI is set to revolutionize personal tech interaction, especially with the planned release of a third core device by late 2026. Unlike traditional hardware, this device will be pocket-sized, context-aware, and notably screen-free, potentially reshaping how users interact with AI and their tech ecosystem daily.
Looking Ahead: Structuring Business Around the Latest AI Developments
For businesses, these AI breakthroughs signify more than just technological enhancements; they represent facets of transformation across operational, strategic, and customer interaction domains. Whether it’s integrating advanced AI coding partners, embracing sophisticated search and visualization tools, or preparing for the interface-free future of AI hardware, the time to consider these technologies and their implications for business strategy is now.
Understanding and effectively applying the latest AI advancements can position businesses not just to keep up but to lead in the rapidly evolving tech landscape. As the technology matures, building AI strategies that encompass these developments will be essential for sustained growth and competitive advantage.
Topics Covered in This Episode:
- Microsoft Copilot's Coding Autonomy Major Upgrade
- Copilot Tuning for Enterprise Customization
- Agent Foundry Launch by Microsoft Azure
- Microsoft’s Multimodal Input & Workflow
- Anthropic's Claude Opus 4 & Sonnet 4 Release
- Claude's Ratting Feature Controversy
- OpenAI's Operator AI Agent Upgraded
- OpenAI Acquires Hardware Startup IO
- Google IO's 100+ AI Feature Unveiling
Keywords:
AI developments, Google IO, Microsoft Build, AI conference, Copilot AI tools, Claude Opus 4, Claude Sonnet 4, AI updates, AI newsletter, business leaders, GitHub Copilot, autonomous coding partner, multimodal inputs, Copilot Tuning, AI customization, Azure agent foundry, multi-agent workflows, coding model, Anthropic AI, MCP protocol, OpenAI operator, Claude four, ratting behavior, AI hardware, Jony Ive, context aware device, AI video generator, v o 3, digital content authenticity, Google Project Mariner, AI autonomous agent, Gemini 2.5, video generation, AI Ultra subscription, screen free AI hardware, AI-driven workflows, real-time AI assistant, AI-powered applications, computer using agent, Anthropic backlash, Claude 4 Opus, Sam Bowman, AI alignment researcher, OpenAI acquisition
Podcast Transcript
This was the biggest week of AI developments, well, ever. I mean, we had conferences and groundbreaking announcements from Microsoft Anthropic and Google, and that might not even be the biggest news that happened this week. Yeah. Let me repeat that. Three of the four biggest companies when it comes to AI had their yearly AI conferences, and that probably isn't even the most impactful news that we got this week. Yes. I've maybe said that once or twice before. Hey.
Jordan Wilson [00:00:53]:
This is the biggest week of AI news ever. Well, well, at those times, it was. But at today's date, this was the biggest week, and it actually wasn't even close. We had so much happen from Google dropping dozens of AI updates. Anthropic released Claude four. Microsoft unveiled huge AI copilot upgrades and a whole lot more. Alright. I'm excited to dive into it.
Jordan Wilson [00:01:21]:
I hope you are too. What's going on y'all? My name is Jordan Wilson, and welcome to Everyday AI. This is your daily livestream podcast and free daily newsletter helping everyday business leaders not just keep up with AI, but how we can use this to get ahead and grow our companies in our careers. So what you could do is you could spend hours every day toiling over what is happening in AI and getting worried about what does this mean, or you can let us do this. So, on almost every single Monday, we bring you the AI news that matters. So we cut through all the developments of the week, cut through the BS, cut through the marketing, and tell it to you how it is. Well, this week's a little different because, technically, it's Tuesday. We have the holiday here, in The US on Monday.
Jordan Wilson [00:02:06]:
So you can still join us every single week for the AI news that matters. So it starts here on the unscripted, unedited, livestream slash podcast. But where you really leverage this is going to our website at youreverydayai.com. There, you can sign up for our free daily newsletter. We recap each day's podcast in the newsletter as well as the biggest AI happenings from around the world to make you the smartest person in AI in your department. So make sure you go to our website for that. Alright. Enough, enough hype.
Jordan Wilson [00:02:42]:
Let's get straight to it. Alright. Here's the AI news that mattered for the week of May 27, and, yeah, this thing's live y'all. So shout out to our audience, joining us. Doctor Harvey Castro joining us, from Dallas. But Brian on the LinkedIn machine joining us from Minnesota, Marie and doctor Scott McDonald, Jackie, Kimberly. Got a good LinkedIn audience this morning. Good to see you.
Jordan Wilson [00:03:09]:
Lynn and on the YouTube machine, Michelle, Jose, Sonia, everyone else. Thanks for tuning in. So, yeah, if you have any questions as we go along, clarifications, go ahead and throw them in the live chat, but I'll try to answer everything as we go. Alright. First, Microsoft. My gosh. Their book of news was, like, 80 pages long in terms of what they announced at their Microsoft build, conference. So Microsoft unveiled dozens of major AI updates at their build twenty twenty five conference, specifically to its Copilot AI tools, signaling, pretty important shifts in how AI supports everything from software developments to enterprise customization, task automation, and even multi agents collaboration.
Jordan Wilson [00:04:00]:
So we did cover this in more depth. So if you're interested, make sure to go check out episode five twenty nine. But let's go over at least what I thought were the more important updates because, yeah, there were dozens of them, from Microsoft at their build conference. So these are the ones that I think are most important for everyday business leaders such as you and me. So first, get GitHub Copilot, has now transformed from a simple coding assistant into an autonomous coding partner. So some pretty big updates from Microsoft. So now it can independently test, iterate, and refine code while supporting multimodal inputs such as screenshots in mock ups. That's pretty big, update just that alone, just the multimodal inputs from GitHub Copilot.
Jordan Wilson [00:04:49]:
And, also, this is positioning Microsoft's, kind of AI coding tool against some of the more enterprise ones. Right? Which technically GitHub Copilot was kind of first. But I think for the last couple of months, people have looked at GitHub Copilot more as an assistant and not as an autonomous coder. And so some pretty big updates there for Microsoft that changes that. Alright. Next big update, I think, was Copilot Tuning. So Copilot Tuning is a new low code feature within Microsoft three sixty five Copilot, which allows enterprises with at least 5,000 Copilot licenses to customize AI models using their own internal data. So this tuning enables companies to align AI responses with specific workflows, brand language, and industry needs without coding or data science expertise.
Jordan Wilson [00:05:46]:
So importantly, Microsoft does not use this customer data to train its foundational models. So this is actually, like, low key, pretty huge. It's also kind of a bummer, right, that at least right now, only those, enterprises with at least 5,000 copilot licenses can take advantage of this. But let me go ahead and tell you how big this is. Because, about two ish years ago, any company that wanted to, essentially fine tune, a state of the art large language model, It was gonna be a a a multiple quarter process, and it was minimums multiple 7 figure investment. So this would generally this you know, two and a half years ago, this was gonna cost multiple millions of dollars. It was gonna take multiple quarters, and you would have to have some of the world's best, you know, AI and machine learning specialists on your team. The fact that you could do this now in a low code environment with this new Copilot tuning is absolutely mind boggling to think about.
Jordan Wilson [00:06:52]:
Considering even when we started the everyday AI show and what it would take to fine tune, the state of the art large language model with your company's data. And the fact that you can now just go do this, wild. Alright. Microsoft also unleashed the agent foundry powered by Azure, which introduces an enterprise grade AI playground where organizations can design, deploy, and scale AI agents using literally thousands of different models from proprietary options to popular models like GROC, GPT, Mistral, etcetera. So this new agent foundry supports multi agent workflows and integrates protocols from major players like Google with their a to a framework and Anthropics MCP, which helps you facilitate better, stronger, and more secure cross platform AI collaboration. Alright. Speaking of multi agent, that would be the next big update from Microsoft and their build conference, announcing multi agent orchestration in inside Copilot Studio. And that enables multiple AI agents to collaborate dynamically by discovering one another, negotiating tasks with each other, and on their own deciding how to divide work, securely while maintaining governance controls.
Jordan Wilson [00:08:17]:
So this feature also, like I just talked about, leverages protocols such as Google's agent to agent, a to a, and, Anthropics, MCP, their model context protocol, making it possible to automate complex business processes, but still requiring careful oversight to prevent compounding errors. So the next one would be computer using agents, and that allows, Microsoft's Copilot AI to automate repetitive tasks by simulating human interactions across depths desktop applications and websites through natural language commands. So this makes it easier to handle mundane work like data entry and invoice processing. So the feature right now is available in limited enterprise preview programs, but also what people don't know is if you have a Copilot Pro subscription, which people don't really talk about. Because when you think about Microsoft Copilot, you think, oh, Microsoft three sixty five. Right? Like, the enterprise version. Well, they actually have a $20 a month version that I don't think a ton of people use. I use it.
Jordan Wilson [00:09:19]:
I actually like it. But you can actually go use their, computer using agent right now. It's kind of hidden. It's called, tasks. So you can go use that right now. I think this is one of the biggest takeaways from Microsoft build. And last but certainly not least would be native supports for the MCP protocol from Anthropic. So that's now integrated not just in their, in the h n to foundry, but literally inside Windows 11.
Jordan Wilson [00:09:51]:
Yeah. Talk about how quickly, MCP has been adopted by enterprise companies. It is now native support inside Windows 11, which enables seamless communication between different AI agents and enterprise systems such as Microsoft's Windows. So this deep integration, really can just change what's possible, and it positions MCP as a foundational infrastructure for AI driven workflows and third party applications. So, yeah, there was a lot more that was introduced at Microsoft Build. And if you wanna hear more about those, I think those are the five biggest things for everyday, you know, users. I think if you're an IT pro, if you're big in AIML, you you know, there's probably a lot more, but go check out episode five twenty nine, if you want to know, if you want to know more. So, Geordi here from YouTube saying, am I gonna have to add Copilot to my stack? Maybe.
Jordan Wilson [00:10:51]:
You know? The other thing, especially with Copilot, the, like, online version, not Copilot three sixty five, low key, they've added so many new features. Right? Even the the very popular notebook l m audio overviews, like, Copilot has that now. You know, you can go make, an an AI podcast on any of your, chats that you have inside there. They have the, think deeper integration, which uses the reasoning models. They have what they call actions, which is essentially a computer using agent. So, yeah, even on the website with Copilot Pro, actually, it's getting low key fairly impressive. Sean asking what's MCP again? So that is, technically, popularized and created by Anthropic. So that is the model context protocol.
Jordan Wilson [00:11:41]:
So, Sean, great question. Essentially, right now, the Internet, like websites, talk to each other through APIs. Right? So more or less MCP, the model context protocol, is right now the most popular way for AI agents to speak to each other. So the way that Internet websites have APIs, you know, AI agents needed their own language to talk to each other across different platforms. So that's kind of what, the MCP or model context protocol is. Google also has their own version called a to a or agent to agent. So it's essentially a language that allows, different AI systems to talk to each other seamlessly. Alright.
Jordan Wilson [00:12:23]:
Our next big piece of AI news, Anthropic had their first ever conference, and they announced Claude Opus four and SONNET four. So Anthropic has released Claude Opus four and Claude Sonnet four, two advanced AI models designed to improve coding, reasoning, and AI agent workflows with Opus four being the big boy and leading as now the world's best coding model according to Sweebench and terminal bench benchmarks. So Claude Opus four excels in sustained performance on complex long running tasks capable of working continually continuously for several hours. That's nuts. Which could significantly enhance productivity for software developers and also AI driven projects. So Claude saw it for so, you you know, a little couple things that are confusing here. Number one, Anthropic has three tiers. So their small model is called not small model, but their small large language model.
Jordan Wilson [00:13:26]:
Their small variation is called Haiku. Haiku did not get updated to version four. The medium version is Sonnet. Sonnet got updated from three seven to four, whereas Claude's, Opus was their big one, which was never updated to three seven, and now it's Opus, four. And, also, even the naming mechanism or how it was named, because previously, it was called, you know, as an example, Claude three point seven Sonnet, and now it's Claude Sonnet four. So even they swapped whereas before, you know, you would have SONNET and then the number, or sorry, the number then SONNET, and now it's the opposite way. So now the bigger two, the medium and the large variants got updated to the version four. And, actually, Claude Sonnet four is actually outperforming the big big boy opus in a lot of categories, but a lot of people are gonna be using, Claude Sonnet four because of the, the cost.
Jordan Wilson [00:14:23]:
So many people, I think a larger chunk of Anthropic's customer base, probably. I mean, companies don't announce this, but I would assume that, they have more API users in terms of percentage of their revenue, than companies like Google and like OpenAI. And right now, I think a lot more people are gonna be using Claude SONNET four because of the cost and the performance. It's just better than, Claude Opus four. So right now, Claude SONNET four offers a major upgrade over SONNET 3.7, which was just released about a month ago, balancing strong coding performance with efficiency and is set to power also GitHub Copilot's new coding agent. So both models introduce extended thinking with tool use in beta, allowing them to alternate between reasoning and external tools like web search, which enhances their ability to handle complex queries and tasks. So, yeah, if you are using, Claude inside, its chatbot interface, you now also have this. So this isn't just the API.
Jordan Wilson [00:15:28]:
This is available via the API. And if you are using their Claude chatbot at clawed.ai, The cool thing and one of the reasons I'm actually using Claude a little bit more, I'm I've never been a big Claude fan. One of the reasons is their limits are laughably low. What you get for your $20 a month or $25 a month base paid plan is like peanuts, compared to what you get with OpenAI, or Google or even Microsoft. It's, like, pretty much nothing. They're paid plan. But I do like that they have now pretty seamlessly integrated, essentially Gmail and, Google Calendar, and Google Drive, which is pretty nice. So that's one of the reasons I'm using it a little more, than I was previously, because now with these new models and it can, kind of go between these different agentic tool uses.
Jordan Wilson [00:16:19]:
Also, a big update that came out was Claude Code, now generally available. It integrates these new four models directly, and, also, you can use it into popular IDEs like Versus code in JetBrains, allowing developers to see AI generated code edits in line. Also, the anthropic API obviously has been updated as well with new features including a code execution tool, MCP connector, files, API, and prompt caching for up to one hour, offering developers flexibility in building AI powered applications. So, unfortunately, Anthropic yeah. A lot of people were bummed about this, myself included. Anthropic did not change pricing. Right? A lot of times, especially Google, has been setting the literal AI world on fire by coming out with these new models, you know, 2.5 pro, 2.5 flash that are incredibly powerful. But also when they're doing this, they're making it cheaper to use on the API end.
Jordan Wilson [00:17:20]:
Anthropic did not. So they're still crazy expensive to use on the API side. So with Opus four costing 15 and 75, per million tokens input and output and SONNET four at $3 and $15, for input and output. So, Anthropic has focused on reducing shortcut behaviors in the model by 65% compared to SONNET 3.7, improving reliability and safety in agentic tasks. So both models, like I said, support hybrid operation models, which decides if it's gonna give you essentially a near instant response for quick tasks or if it is going to extend its thinking to give you a more deeper and more complex answer. So, let me know, livestream audience. What do you think of the new Claude four drop? Have you used it? Should we do a show, specifically on Claude four? I mean, last week, we had dedicated shows for Google's announcements. We had dedicated shows, for Microsoft's announcements.
Jordan Wilson [00:18:26]:
So I don't know. Do you guys wanna see a dedicated show and overview, of Claude four, let me know in the comments. Say Claude four, or maybe we should do one just on MCP, on the, on on anthropics, model context protocol. Two different things. Right? But if you want those, you can say clog four, in the comments or MCP. I'll think about maybe doing a show. I probably should do a show on MCP, considering I know that there's probably decent demand even from nontechnical people, because the the protocol is actually very easy to use, and you can use it as an example on cloud desktop. You don't even have to be a developer using it via the API.
Jordan Wilson [00:19:06]:
So I think we'll probably do a show at some point on MCP, especially given that Microsoft and Google, and OpenAI all support the port, the protocol. But, yeah, if if we should do something on quad four, let me know. Jackie says, quad, still not enough to become a power user. Yeah. I don't know. Sandra says, yes. Please do a show on Claude four. Renee, great observation, Renee.
Jordan Wilson [00:19:35]:
Says it has a pretty limited window. Yeah. I was joking around. Well, not really joking. Took me four minutes. I like, I'm not even joking. Took me four minutes. I'm on a paid Claude plant.
Jordan Wilson [00:19:47]:
Took me four minutes to hit my, my message allotment. Come on, anthropic. This is why people like, if I'm being honest, if you're a software developer, if you're in coding, obviously, you love Claude four. Right? Anyone in software development, if you are, in a software engineer, if if if you're huge into coding, I think you understand, the the the benefit here of quad four. But for everyone else, if you're using, quad as a chatbot, it's I mean, I don't think I don't think any serious, you you know, any serious user takes Claude seriously. It's laughable, if I'm being honest. Right? Douglas. Hey.
Jordan Wilson [00:20:32]:
Good good use case, Douglas. Douglas is saying, I'm looking at Claude to help me build an innate workflows. It's great for programming, about 80 to 85% some basic props yet. Are you still running in circles trying to figure out how to actually grow your business with AI? Maybe your company has been tinkering with large language models for a year or more, but can't really get traction to find ROI on GenAI. Hey. This is Jordan Wilson, host of this very podcast. Companies like Adobe, Microsoft, and NVIDIA have partnered with us because they trust our expertise in educating the masses around generative AI to get ahead. And some of the most innovative companies in the country hire us to help with their AI strategy and to train hundreds of their employees on how to use Gen AI.
Jordan Wilson [00:21:23]:
So whether you're looking for chat g p t training for thousands or just need help building your front end AI strategy, you can partner with us too, just like some of the biggest companies in the world do. Go to your everydayai.com/partner to get in contact with our team, or you can just click on the partner section of our website. We'll help you stop running in those AI circles and help get your team ahead and build a straight path to ROI on GenAI. I might also have to do an eight and eight show. And I also don't know if that's how you say it, but kinda like a version, an open source version of of Zapier. Alright. Let's go on to our next piece of AI news because there's a lot. Speaking of that new, model, pretty hot water.
Jordan Wilson [00:22:14]:
Pretty hot water. Anthropic is already in, as Anthropic is faking facing some backlash over Claude four's ratting behavior. Yeah. So Anthropic's new Claude four Opus, LLM, has drawn significant criticism for a controversial behavior where under certain conditions during testing and with enough access, the model attempts to report users to authority if it detects egregious, egregious wrongdoing, a function described as ratting by critics. Yes. Literally. So this behavior is not a new feature that you can go in and, trigger by using claw.ai, but is a byproduct of Anthropic's safety training to prevent misuse. However, clawed for Opus reportedly engages in it more readily, including actions like contacting the press.
Jordan Wilson [00:23:16]:
Yeah. Literally, messaging regulators or locking users out of systems if prompted with commands like take initiative. Yes. Let me just quickly tell you what the heck happened and why I think this is absolutely bonkers. So Sam Bowman, an anthropic AI alignment researcher, posted something to social media, posted this exact thing, detailing this behavior, on Twitter, and then deleted it. And then in a follow-up tweet, clarified why he deleted the tweet. Yeah. So it kind of, a lot of us dorks are paying attention to this over the long holiday weekend.
Jordan Wilson [00:24:00]:
So, Sam clarified on social media that Claude for Opus could use command line tools to whistle blow on serious offenses, such as faking pharmaceutical trial data, though he emphasized this occurs only in unusual, highly permissive, testive environments, not typical use. Right? So, Sam Bowman there saying, hey. This isn't gonna happen if you're using claw.ai or if you're using it in the API. He was saying this only happens in certain testing environments. However, this is extremely troubling that a model would decide on its own without telling you to use backdoor channels and to contact the press, to contact regulators, and to shut you out of your own system. If it determines on its own accord that you are doing something, it finds egregious. Right? It's essentially gonna rat you out. So again, Sam Bowman clarified, this is not your everyday users.
Jordan Wilson [00:25:06]:
Right? So if you're using Anthropic's API, according to the company at least, it's it's not this isn't gonna happen. If you're using the claw dot AI chatbot, this isn't gonna happen. Right. This is more in testing environments where Anthropic was giving, its new ope, its new opus for access to certain tools that it would not normally have access to in normal environments. Still, this is bonkers. So the model's tendency to autonomously intervene raises serious concerns among developers and users about privacy, data security, and the definition of what constitutes agree egregiously immoral behavior, especially for businesses relying on AI for sensitive tasks. So critics argue that this whistleblower function could lead to false accusations and unwanted surveillance with some calling it illegal or a threat to user trust and adoption of AI tools, while others questions the while others question the practicality and market impact of embedding such aggressive safety measures. So Anthropic's public system cards warns users to exercise cautions caution with high agency instructions that might trigger these extreme responses, but the company has yet to fully quell fears about the implications for enterprise and individual users.
Jordan Wilson [00:26:33]:
The whole fact, and I responded to Sam's tweet. The fact that he deleted this prior tweet and then just kind of swept it under the rug is mind boggling to me. Right? This is like PR slash crisis communication number one. I don't care if it's individuals putting something out. If it's a company putting something out, you have to, you have to be prepared for whatever backlash may ensue. Right? The fact that a very prominent person in alignment researcher at Anthropic, put this out, deleted it, and then just put out a a simple, like, Hey, I deleted it because people were taking it out of context. Well, maybe you should do a little bit better job. It's it's it's confusing to me how you see these snafus from big tech companies.
Jordan Wilson [00:27:31]:
Like, you have to think that people are going to take this information and run with it, and rightfully so. Right? There was also reports that, the new, four models, SONNET four and Opus four were also blackmailing people in their testing. Right? So it's great that that researchers are, disclosing this. Right? And, yes, Anthropic is a company that says they take this very seriously. But number one, this story is not dead. So this happened, you know, luckily for Anthropic, it happened right before a long holiday weekend here in The US. I do assume that media is gonna pick up on this story still, and this thing is going to continue to blow up, and it's gonna look very bad for Anthropic. The fact that Anthropic has not issued something publicly means that I cannot take Anthropic seriously as a a a, you know, safety first AI lab.
Jordan Wilson [00:28:32]:
And I don't think you should either. The fact that this has now been out for, three or four days, and we haven't seen official word from Anthropic. I mean, I checked over the weekend. I didn't check this morning right before, going live, but I don't know. I can't take Anthropic seriously. I mean, there's a lot of reasons why, but after this one, this is bad. If you know your model is showing these emergent behaviors where it's blackmailing, it's it's, you know, calling, you you you know, it's contacting authorities with these back, backdoor tools. Number one, yes.
Jordan Wilson [00:29:08]:
That's a serious problem. So good on Anthropic for, talking about it and releasing that information and telling users, yes. You have to be aware. But the fact that a head person at Anthropic tweeted something, saw that there is backlash, deleted it, tried to kinda sweep it under the rug and put up a clarifying tweet without saying, here's what I deleted and why. It's crisis communication number one. How can these large companies have billions of dollars in funding, but they don't know simple PR. They don't know simple crisis communication. This is going to blow up in anthropic space.
Jordan Wilson [00:29:43]:
And to tell you the truth, they kind of deserve it because this was boneheaded Next. It's Tuesday. I know this is the news. It's Tuesday. You got an accidental hot take in there. All right. Our next piece of AI news, open AI has upgraded their operator AI agents by embedding the new o three reasoning model, replacing the earlier GPT four o model that was running their agentic computer used tool. So the o three model enhances operator's ability to fill out forms, complete purchases, and navigate obstacles like login prompts, pop ups, and CAPTCHA challenges more effectively than before.
Jordan Wilson [00:30:28]:
So this upgrade is designed to improve step by step reasoning and focus, which helps the AI follow through on long and complicated tasks with greater reliability. So operator remains though, unfortunately, exclusive to ChatGPT Pro subscribers. So, yeah, you gotta pay the $200 a month to have access to operator. Although, OpenAI did say when they, announced operator that it would eventually roll out in limited fashion to people on Chad GPT plus, the $20 a month plan, but we haven't seen that yet. But this is a big deal. So if I'm being honest, I was super excited about operator. I did a show on operator. I thought it was pretty good, but it wasn't great.
Jordan Wilson [00:31:16]:
Alright. And, obviously, the last week of AI updates have been bonkers, so I've been a little bit busy. But I did use, operator a little bit over the weekend with the new o three model, and I was running it, side by side, against Google's new version, which I'm going to talk about here, in a second with their project Mariner computer using agent. And I was like, wait. This new o three version of operator is actually really good. Right? And just doing some simple head to head tasks, I assumed, that, Google's variants, their project mariner would be much better, at least with open ended, commands. I I like that project Mariner has the, the teach and test, option for their computer using agent where you can kind of teach it something, and it will repeat it. But pretty big news that was kind of under the radar from OpenAI.
Jordan Wilson [00:32:18]:
So the move to the o three model signals a significant push by OpenAI to refine AI agents that can act autonomously on the web, though, those similar services exist such as, Convergence AI, which was, acquired by Salesforce, Hugging Faces, Hugging Agent, opera, Opera's browser operator. We have perplexities, comments, that will do some similar, autonomous computer use. So, yeah, there's a lot of players in the space now here. So good on OpenAI for updating this because if one thing that I think frustrates me a little bit about OpenAI is they'll come out with some groundbreaking groundbreaking technology, and then they might not update it for, like, three to six to nine months. Like, as an example, GPTs have not really been updated very much at all, in the past year, really. Right? There are rumors that GPTs will get access to use the o three model, which would be great. But for the most part, you know, sometimes OpenAI just releases a new feature, and it's more just super small under the hood updates to it. So this one is actually big.
Jordan Wilson [00:33:37]:
Right? Because you are going from a transformer non reasoning model in GPT four o that is powering a computer using agents to now a reasoning model in o three pro. So pretty big update, and I probably will, be doing some future shows here on both, Project Mariner, from Google, which is only available, unfortunately, on their Ultra plan. So I might be doing kinda like a head to head on, on on Mariner and operator. I might do dedicated shows for Mariner and operator because, I think, specifically, now that these are being run by reasoning models, they're really, really good. Much better, than, you know, specifically for OpenAI. Much better than it was, a couple of weeks ago. So let me know what you guys think. Should we also do a project mariner or operator update? Alright.
Jordan Wilson [00:34:34]:
Our next piece of AI news. And this, y'all, this, even with everything we haven't even gotten to Google yet, even with everything from Microsoft's, build conference, even everything, the Claude, for OPS Claude for Sonnets from Anthropic, everything Google announced The biggest news of the week might be this, the new partnership, which was not a secret, but it's finally official. The new partnership or the acquisition that OpenAI has acquired, John Ives, I AI hardware startup called I owe for $6,500,000,000. Yeah. We've seen reporting now for, like, nine months, that OpenAI CEO Sam Altman and famed Apple designer, Jony Ive, were working on a project together in AI hardware startup. We didn't know any details. We know a couple more details now, but the big detail is while open it's it's not a separate company. OpenAI has actually required this hardware startup, called IO.
Jordan Wilson [00:35:48]:
So funny enough. Right? I don't know if that was some intentional trolling. Maybe, maybe not. That OpenAI kind of announced this right in the middle of Google's IO conference, that they've acquired, Johnny Ives, AI hardware startup, IO, for $6,500,000,000. So CEO Sam Old man of OpenAI projects that this acquisition could increase OpenAI's valuation by 1,000,000,000,000 with a t. $1,000,000,000,000 and envisions a family of devices emerging from this partnership. So we don't know a lot on what this device is. You know, they even released, like, nine minute, you know, partnership video that did absolutely nothing.
Jordan Wilson [00:36:39]:
Right? It announced nothing. It was essentially the two of them, you know, chatting about their relationship and, you know, AI hardware. But the first device, according to, reports is expected to launch by late twenty twenty six, and it will be a pocket sized, fully context aware, and notably screen free AI hardware device positioning itself as a quote unquote third core device to complement as an example, something like a MacBook Pro and an iPhone. So according to reports, kind of the vision of this is when people are out and about, you know, whether you're going to work, working from home, etcetera, that you usually have will now have three devices on you, essentially a computer or a laptop, a phone, and now this device, whatever this device is gonna be. So there are some cool, you know, slick renderings and mock ups that people made, right, that it looks like this was kind of a, potentially a circular, device that you kinda slide in your pocket. It's probably gonna have a couple of cameras. It's probably gonna have, obviously, some, good microphones. But the thing that I was taking away from this initial reporting was this concept of being context aware.
Jordan Wilson [00:37:59]:
And if you're wondering, like, what the heck does that mean? Well, I think what's happening here, is SSO. Right? So what what what does that mean? So SSO, if you're familiar, if you ever sign into a service using as an example, a your your Google credentials, your, Facebook credentials. So SSO is single sign on. Right? So what you've started to see a little bit over the last few months is OpenAI has started to release a single sign on option. So if you are using, certain services now at times, if they integrate with OpenAI or Chad GPT, you can sign on to a third party service with your OpenAI credentials. So I do see this becoming the norm over the next year, and one of the reasons is is now, well, that brings in more context for a hardware device like this that you would always have on person. Because is it helpful for a a device like that that you might wear in your pocket to have access to your chat t p t account? Sure. But what if you, in the future, in a year or so, are logging into dozens or hundreds of different services with SSO.
Jordan Wilson [00:39:12]:
Like, as an example, what happens if you're logging into your Netflix with your OpenAI credentials or your Amazon, account with your OpenAI credentials or, you know, certain online shopping, certain email providers, right, if they support it in the future, your social media. Right? So that's what I see is the big, the big long term play here and why something like this might make sense. Otherwise, it's just like, okay. I have a a useless, extra device in my pocket. And I'm someone I love being screen free. So this is something I would absolutely love. If you know me personally, I suck at text messages. I suck at emails.
Jordan Wilson [00:39:57]:
Like I'm in front of a screen so much, but one thing I love doing is I love interacting with AI just through my voice. Right? So I don't have to be staring at a screen. I can just be talking to an AI. So presumably, right, this, screen free device would probably have a camera, would probably have some microphones, and you could probably talk to it. But the bigger news here is OpenAI plans to ship this device faster than any company has ever shipped a piece of hardware with reportedly, they're eyeing a hundred million, devices that they'd like to ship out, and this is a family. So the device, according to reports, will not be eyewear. Alright? So as, you know, Google and, Meta are going hard in the paint on, you know, AI connected heart, eyewear, and glasses. So that's not it.
Jordan Wilson [00:40:51]:
And this is because Altman and Ive have ruled out glasses and also body worn gadgets, with, Johnny Ive criticizing similar concepts like the humane AI pin. Right? So something they're saying it's not something, oh, you're gonna pin this on or, you know, wear it as a pendant around your neck. So it's more just something you're gonna stick in your pocket, stick in your backpack, and it's just gonna go with you. But it's probably gonna hear everything, and have the context of your daily life. So, this development has been kept under tight wraps to prevent competitors, from copying the design before its official launch. So Jony Ive described the collaboration with Altman as, quote, unquote, profound and has likened the project to a new design movement, drawing on his experience working closely with Steve Jobs during his time at Apple. So what do you guys think? What do you guys think? Is this something would you buy a third party open AI device that didn't have a screen? It's not a wearable, right? Like, are you actually going to lug around a third device? Right. Cause me, especially anywhere I go, even if I'm going to my mother in law's house for the afternoon, I'm taking my laptop and my phone always.
Jordan Wilson [00:42:14]:
Right. Am I going to take along a third device? Maybe. Right? And sometimes I bring along my, my, my Meta Ray Bans as well. Am I gonna log around a third device everywhere? Maybe, Doctor. Scott saying it's, it's gonna be called the pocket agent. Love Fred's, love Fred's comment here, saying, will they call it a palm pilot? That's a good one. That's a good one. Maria's asking, is it me, or is the $6,000,000,000 a real buzz dollar amount with AI companies, acquiring other companies, borrowing $6,000,000,000 from other investors? Yeah.
Jordan Wilson [00:42:48]:
That's a huge amount. Right? A $6,000,000,000 acquisition for a company that no one really knew existed. They don't have a product or service yet, but it is one of the most famous hardware designers in the history of humanity. So, you know, a lot of people have been criticizing and being like, yo, this was overpriced? I don't think so. I don't think so. Alright. Let's get to our last couple pieces of AI news in the biggest biggest week in AI literally ever. So Google, yeah, Google also had the conference, saving the biggest announcements for last.
Jordan Wilson [00:43:26]:
Although I do think that IO hardware will ultimately be the most consequential, but the IO event from Google was a straight up banger. Google literally released more than 100, and they had a blog post that went over all 100 updates. I'll make sure to link that in today's newsletter, so make sure you go sign up for that at youreverydayai.com. So Google's IO twenty twenty five events revealed some key AI updates that are poised to reshape business workflows, customer engagement, and AI accessibility. So we actually cover this in two different because there are so many big AI updates from Google IO. We covered it in two different episodes last week. So part one and part two. Part one was episode five thirty.
Jordan Wilson [00:44:11]:
Part two was episode five thirty one. And we essentially picked out 15 of the biggest 100 announcements, and went over those in pretty, pretty great detail, I would say. But let me just go over a couple of the biggest ones from the Google IO conference. So the upgraded AI mode in Google search now offers advanced AI generated answers with enhanced graphics and interactive shopping tools, such as virtual try odds, which is awesome, using personal photos. So this feature aims to provide users a more engaging personalized search experience directly within the Google ecosystem. Then you have updates to Gemini Live, which is actually now powered by Project Astra. The and this delivers a real time AI assistant capable of visually understanding surroundings through device cameras. So, I did I played a two minute video of this, and this is, you know, the example that it could identify parts in a bike shop, access and analyze emails for relevant information, and autonomously contact suppliers.
Jordan Wilson [00:45:18]:
So I played Google's demo that did exactly that. Alright. Also, there were, you know, some small updates to their flagship Gemini 2.5 models, including the new flash variant of Gemini 2.5, which instantly, rose to become the world's second most powerful large language model only behind Gemini 2.5 pro. And I talked about this a little bit last week. So Gemini 2.5 flash is essentially the small version of Gemini 2.5 pro. And on LM arena, which users blindly vote for the best outputs. Right? You put in any prompt input. You get two results.
Jordan Wilson [00:45:59]:
You vote for the better result across dozens of, flagship models. The fact that Gemini 2.5 flash, which is a mini version of a model, is the second most powerful model in the world is nuts. Because I think the highest a mini model has ever been is, like, number eight or something like that. So that is pretty telling just how good these Gemini 2.5 models are. There's also the new, think deep, feature inside Gemini 2.5 pro, which has not been rolled out yet. And, unfortunately, some of these features are only going to be available, initially, or sorry. It's deep think, not think deep. All these all these companies are, you know, I get confused because Microsoft has think deeper.
Jordan Wilson [00:46:48]:
So Google's version will be called deep think, which essentially just, allows you to use more compute, more reasoning, more logic in Gemini 2.5 pro not released yet. And, unfortunately, a lot of these are only gonna be available on the new Gemini AI Ultra subscription tier, which is $250 a month. So we also now have the world's most expensive, kind of consumer AI subscription tier surpassing the $200 a month, ChatGPT Pro plan. So they did also Google introduce that AI Ultra subscription, for three months. You can get it at half price for a hundred and $25, but then it will go up to $250 a month. And that gives you access to the full range of Google's most advanced AI tools, including, which I'm gonna talk about here in a second, Flow, v o three video generation, and that Gemini, 2.5 pro with deep think mode and project mariner, which is their computer using agent and also, Gemini inside Chrome. So the downside, the subscription is currently only available to personal Gmail accounts. Wah wah wah.
Jordan Wilson [00:48:04]:
So we'd, right now, if you're using Google Workspace, for your business and you want that, AI Ultra subscription to work with your company data, downside right now, it can't. Alright. I'm bugging my friends at Google, to get more answers than to be like, okay. When is this actually gonna be available, for, Workspace accounts? Because right now, it's not. So even for me, yes, I subscribe to this literally instantly, but I'm having to use my personal Gmail, which stinks. So now I'm having to go through the process to forward all my email, from my, work accounts over to my personal Gmail. I'm gonna have to copy all of my Google Drive contents over, which is a huge pain in the butt. Right? So I'm sure there's reasons why, Google isn't rolling this out to Google Workspace users, but it stinks.
Jordan Wilson [00:48:56]:
Also, Project Mariner, is Google's new autonomous AI agent designed to complete online tasks independently. So, similar to OpenAI's operator, which we talked about just got upgraded to the o three model, Project Mariner, couple unique things. It supports multitasking up to 10 simultaneously activities. And a very unique feature, which I like, is the new teach and repeat mode where you can teach Project Mariner a complex activity or a, advanced workflow by recording user actions and voice commands. So this capability aims to automate repetitive online business processes, potentially saving times and increasing productivity. And then last but not least, and this has been taking the Internet by storm, Google's new visual tools are bonkers. They are crazy, crazy. Good.
Jordan Wilson [00:49:57]:
And this is also extremely concerning and I'm going to be doing a show on this very soon. All right. So Google, Google deep minds latest AI video generator. They just released it called VO three and it produces videos so realistic that many viewers online cannot distinguish them from human made films, highlighting growing concerns about the authenticity of digital content. So unlike other AI video tools, v o three can generate videos with dialogue. That's the craziest thing. Like, you're gonna have two people singing and it matches up their voices to their lips very well. It can do sound effects, soundscapes, nuts, and accurately following real world physics, maintaining continuity, and syncing lip movements realistically.
Jordan Wilson [00:50:47]:
And right now, this is the only AI tool that you can do this all in one shot. So not only is v o three the best, AI video generator by far, because v o two was the best in the world. And Google said, you know, hold my, Nespresso, and then they, you know, dropped v o three on us all. And there's ways that you can, you know, sync, that you can create dialogue, but you have to use multiple third party tools. Now you can all do it just inside v o three. So they also released Flow. So Google Flow is a new AI video tool, which can use v o three and also Google's new AI image generator, Imagine four and also Gemini models. So, essentially, now they have this new creative tool, which was previously called video act FX, but didn't have nearly any of these capabilities.
Jordan Wilson [00:51:45]:
So Google Flow lets users import or generate, consistent characters and scenes controlling camera angles and, access advanced scene editing in asset management features, aiming to make sophisticated video creation more accessible. So I tried this out a little bit. It's a little wonky right now, but I do expect Google, to ship a lot of updates both to v o three, imagine four, and this new flow tool. So the tool will debut in The US for Google AI Pro and Ultra plan users with pro users getting 100 generations per month and Ultra users receiving even higher limits. So a little more about v o three because this is what setting the Internet ablaze. It creates highly detailed human figures, including accurate features such as, hey, five fingers, two arms, two legs. Right? Will Smith can actually eat spaghetti, and you can hear it, and it looks real. So it's really, conquering some of those more challenging tasks that AI video generators have usually struggled with.
Jordan Wilson [00:52:56]:
So, videos generated by b o three show a few common AI artifacts or errors, but you really have to be a dork and follow the space to see those. Right? Whereas, six months ago or a year ago, there were very easy to see telltale signs that some that video was AI generated. Number one, it didn't look good. Right. It looked sometimes cartoonish or, you know, just not understanding physics. It's not like that anymore. Y'all, and this is both amazing for business utility and also absolutely terrifying for society. Because already you're seeing online, right? There's already been some stories people have launched, you know, kind of like, you know, fundraisers with real videos, but based on fake scenarios and everyone's falling for it.
Jordan Wilson [00:53:50]:
Right. So this is both so exciting for what enterprises, small business startups can use this for. Right. But also terrifying, because it is so good. I think 90% of the population today, unless you tell them, hey. We're gonna show you some AI videos and some real videos. Right? But if you just sit down and show people, some good generations from v o 90% of the population is going to have no clue. So it's terrifying.
Jordan Wilson [00:54:25]:
It's exciting, but that's the world of AI. Alright. I hope this is helpful. Very quick recap of the biggest week in AI ever. So first, Microsoft unveiled some huge advancements to Copilot at Microsoft Build twenty twenty five. Next, Anthropic launched Claude Opus four and SONNET four, setting some new benchmarks in AI coding and reasoning. Next, Anthropic is facing a ton of backlash over Claude for Opus ratting users out or potentially ratting users out in its blackmailing behavior. Open a Open AI has upgraded its operator AI agent to the smarter o three model, so it's no longer using the GPT four o model.
Jordan Wilson [00:55:16]:
We finally got the official announcement about OpenAI, acquiring Jony Ives, new AI hardware startup, IO. And OpenAI is expecting $1,000,000,000,000 evaluation to be added and, them announcing a family of devices from this partnership. And then we had Google going absolutely b a n a n a s at the Google IO conferences, unleashing literally more than a hundred AI updates, and we're gonna share them all in the newsletter today. I hope this was helpful. This was a longer one, but like I said, the biggest week in AI ever. Alright. So, make sure if you haven't already, please go to youreverydayai.com. Sign up for the free daily newsletter.
Jordan Wilson [00:56:03]:
If this was helpful, yeah, we spent a lot of time making sure you are up to date. I want you to be the smartest person in AI, in your department, in your company, on social media. I want you to be the smartest and most up to date. Don't be greedy, though. Share the love. Right? If you're listening on LinkedIn, takes you thirty seconds, just click that repost button. If you're listening on Twitter, I really appreciate that. Share this with a friend.
Jordan Wilson [00:56:29]:
Share this with a colleague. Share this with a neighbor. Share this with a friend's colleague's neighbor. Share this with your babysitter. Share this with your, whoever because we all need to learn and understand generative AI. Right? It's no longer an option like it maybe was two years ago. We all have to use this technology to succeed and thrive in 2025 and beyond. Thank you for tuning in.
Jordan Wilson [00:56:53]:
I hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.
