Ep 754: Anthropic’s ‘scary’ new model, Microsoft Copilot’s ‘Code Red,’ OpenAI’s Superinteligence New Deal and more

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


AI's Competitive Heatwave: Business-Specific Insights from the Latest in Anthropic, OpenAI, Microsoft Copilot, and Meta

The AI sector’s landscape shifted notably in the past week, with advancements and strategic plays that carry clear implications for business leaders. This comprehensive analysis breaks down the developments across Anthropic, OpenAI, Microsoft Copilot, and Meta, translating the headlines into practical meaning for organizations making investment and technology adoption decisions.

Anthropic’s Claude Managed Agents: Raising the Bar for AI Deployment

Anthropic’s recent introduction of Claude Managed Agents in public beta reflects a decisive move toward accessible, production-grade AI for technical teams. The platform lets developers build, host, and scale AI agents without infrastructure management, offering sandboxed code execution, detailed credential management, scoped permissions, and persistent session capabilities—including support for multi-agent coordination.

Specific value for organizations:

  • Enables internal teams to deploy AI-driven automations—such as real-time system monitoring or responsive customer support—without needing to understand or manage backend infrastructure.

  • Facilitates rapid prototyping through direct chat-based agent creation, reducing time from concept to deployment without requiring advanced coding skills.

  • Enterprise-ready orchestration minimizes operational friction while maintaining governance with checkpointing and context management.

  • Clear cost structure: Usage billing is token-based, meaning budgeting for AI workloads can be precise but needs to be actively managed, especially as parallel agents are spun up for complex tasks.

This deployment model is currently technical-first (requiring backend access, not available to public Claude AI users) but signals the larger trend: managed AI operations are now a reality for non-hyperscale businesses.

OpenAI’s Codex Expansion and Super App Convergence

OpenAI’s internal signals and product testing suggest a strategic move to consolidate ChatGPT, Atlas browser, and agentic engineering tools into a unified desktop experience. The newly-leaked Scratchpad feature for Codex points to running parallel tasks from user-generated to-do lists, turning the Codex app into an actionable hub beyond code-writing alone.

Key implications:

  • Potential standardization around Codex as the central business desktop AI, reducing tool sprawl and workflow fragmentation.

  • The introduction of persistent, long-running agent tasks (‘heartbeat system’) means businesses can tackle processes that require continuous oversight or multi-step logic, previously a pain point with transient chat sessions.

  • Pending model upgrades (GPT-5.5/‘Glacier’) could raise baseline performance for those investing in OpenAI’s SaaS ecosystem. Close attention should be paid to rollout announcements—early adoption may confer competitive advantage in automation and efficiency.

Microsoft Copilot’s Strategic Code Red and Impacts on Enterprise AI

Recent reporting highlights Microsoft’s ‘Code Red’ urgency for Copilot, with CEO-level involvement driving a near-term sprint to improve performance and user experience. Despite being an early enterprise leader, Copilot’s adoption lagged amid rising competition from Anthropic, OpenAI, and Meta.

Insights for the enterprise:

  • Upcoming releases with new features (including the e7 suite and high-frequency updates) aim to restore Copilot’s value proposition; inefficiency today does not predict stagnation tomorrow.

  • The Copilot platform is already integrating both OpenAI and Anthropic’s models, offering users access to multiple leading AI engines—potentially giving Microsoft’s product suite a breadth and reliability advantage, provided internal training and change management are addressed.

  • Microsoft’s position as a core investor in two competing AI vendors signals a unique hedge: even as some users ‘jump ship’, Copilot’s underlying infrastructure benefits from the enterprise’s own investments elsewhere.

For buyers, this suggests that integrating or maintaining Copilot could offer a backdoor to future innovation without constant vendor switching, as long as organizational readiness and user training gaps are proactively managed.

OpenAI’s “Superintelligence New Deal”: Early Tax and Workforce Policy Proposals

OpenAI’s 13-page industrial policy proposal explicitly calls for government action on taxing AI-driven labor, creating national public wealth funds, and enacting a four-day/32-hour workweek tethered to AI productivity gains.

Considerations for management teams:

  • Whether or not these proposals become law, signals are clear: regulatory and tax impacts on automation are entering mainstream conversations. Workforce planning and automation strategies should incorporate exposure testing for robot taxes and/or new compliance requirements.

  • The emphasis on treating AI access as a basic right for workplaces and schools is significant. Refining upskilling programs for employees will be critical as “access” may soon be part of baseline regulatory compliance or competitive hiring.

  • Proactive scenario analysis—how sudden changes to corporate tax structure, labor regulations, or AI usage requirements would impact the bottom line—is now prudent, not theoretical.

Meta’s New MUSE-Spark Model: Performance and Access Shifts

Meta’s launch of the MUSE-Spark model (currently closed-source, not yet API-accessible) surprised observers with near-parity to Google, OpenAI, and Anthropic’s top models in language and visual tasks, though it trails in coding and abstract reasoning.

Actionable outcomes:

  • For organizations with research or development teams, the soon-to-come API represents access to frontier large language model performance without the costs typically associated with competitor models.

  • Since the model is closed and proprietary, adoption strategies must emphasize risk assessment on vendor lock-in.

  • Early benchmarking suggests that for certain workloads—document analysis, writing, visual comprehension—MUSE-Spark can be competitive with flagship products, broadening sourcing options for AI capabilities.

Anthropic Mythos: Security and Democratization Divide

Anthropic’s Mythos preview model, intentionally withheld from public release because of its unprecedented ability to find critical software vulnerabilities, is being distributed selectively to major industry players. This marks the first tangible split in AI democratization, as foundational cybersecurity tools become restricted to an elite group (Project Glasswing).

Strategic takeaways:

  • Security-conscious organizations outside Fortune 100 circles must reassess threat models and vendor reliance, as the latest defensive tools may be inaccessible for months or years compared to big tech peers.

  • The era of instant, universal parity in AI tooling is over. Partnerships, preferred vendor status, and rapid information-sharing with select AI labs now matter for security competitiveness.

  • Public narrative around ‘too powerful to release’ models may shape customer and regulator expectations—transparency and multi-stakeholder review will be on the table for future technology assessments.

Additional AI Sector Moves: Keeping Pace with Multinational Benchmarks

Several additional points merit attention:

  • OpenAI’s new $100/month ChatGPT Pro tier offers expanded Codex access, filling a vital gap between basic and enterprise pricing.

  • Alibaba’s Happy Horse 1 leads in video AI, and ZAI’s GLM 5.1 debuts as the highest-benchmarked open-source model for software engineering (surpassing GPT-4, Claude, and Gemini in SweeBench).

  • Regulatory developments—including possible Pentagon blacklisting of Anthropic in the US and early warning from Goldman Sachs on AI-displaced worker risks—underscore the need for agile legal and HR policy frameworks.

Summary: Realigning Business AI Strategies for 2024

The past week’s movement shows that AI strategy cannot be outsourced to procurement or R&D alone. Enterprise IT, compliance, operations, and leadership teams have new opportunities—with parallel agent platforms for workflow orchestration, rising competitive performance from unexpected vendors, and early signals of policy and economic impacts now at the forefront. The window for “wait and see” is closing, and differentiated competitive advantage will emerge from active, informed management of AI adoption, cost, compliance, and vendor partnership strategies.


Topics Covered in This Episode:

  1. Anthropic’s Claude Desktop Power User Redesign
  2. OpenAI Codex Super App Expansion Leaks
  3. Anthropic Managed Agents Beta Launch Details
  4. OpenAI’s Superintelligence New Deal Policy
  5. Meta Muse Spark Model Benchmark Results
  6. Microsoft Copilot Code Red Performance Push
  7. OpenAI CEO Sam Altman Targeted Incidents
  8. Anthropic Mythos Cybersecurity Model Withheld
  9. ZAI GLM 5.1 Open Source Model Outbench
  10. Alibaba Happy Horse One Video Model Ranking




Episode Transcript 



Jordan Wilson [00:00:16]:
Anthropic has a new scary model. Very scary things are happening to OpenAI CEO, Sam Altman. Microsoft is reportedly panicking about Copilot's performance, and somehow Meta's newest AI model is crushing it. Yet that's not even half of what's moving and important in the AI world right now. That's because OpenAI and Anthropic both confirmed big desktop releases this week. Oh, and open AI wants to tax robots and let humans only work for days. Yay. All right.

Jordan Wilson [00:00:53]:
If you missed any of that or everything else that we're gonna go over today, don't worry. That's what we're here for. This is the AI news that matters. If you're new here, welcome to Everyday AI. We do this well every day. It's your daily livestream podcast and free daily newsletter helping everyday business leaders like you and me not just keep up with everything that's happening in the AI world, but how to make sense of it and to get ahead and grow your company and career. So if that's what you're trying to do, awesome. Starts here with the unedited, unscripted livestream podcast.

Jordan Wilson [00:01:23]:
But to be the smartest person in AI at your company, make sure to go to our website at youreverydayai.com. Alright. There, you can sign up for the free daily newsletter, and we'll tell you everything else that's happening today. But on Mondays, we give you the AI news that matters. So don't spend all week, you know, spending multiple hours a day reading things and being like, oh my gosh. Is this real? Is this fake? No. Just join us on Mondays. It's a great way to kick off the week.

Jordan Wilson [00:01:51]:
I do this literally nonstop twenty four seven to help you. Alright. So let's start off with our first piece of AI news, which is actually a big one, and it hasn't happened yet, but the companies have confirmed that it is. That's because we have, well, things are heating up just like the weather here in April in Chicago. It's like, my gosh. It's finally more than 20 degrees. But as the weather is heating up, so is the AI competition. So we are getting new releases this week from Anthropic and from OpenAI.

Jordan Wilson [00:02:24]:
So here's what we have. BIA, reporting from testing catalog. So Anthropic is preparing epitaxy? Apotixy. I don't know. And what are these code words? Can we get some easier ones? Apotaxy, a major power user redesign of Claude. So according to details uncovered from Claude Code, the internal project is codenamed Epitaxy, and it signals a big shift toward a more professional power user focused desktop experience that could ship this week. So the redesigned interface introduces a single window layout with dedicated panels for plans, tasks handled by sub agents, and diffs, plus live coded previews and support for working across multiple repos. So Anthropic is also developing a new coordinator mode, which would allow Claude to manage and delegate work across multiple parallel sub agents while concentrating on higher level planning in synthesis.

Jordan Wilson [00:03:27]:
Users will also be able to reportedly create agents directly inside the app on the fly. Alright. Now with OpenAI, we've seen a lot of rumors swirling lately, but it seems like they're all but confirmed as multiple members of OpenAI did say that they're gonna be shipping some major updates this week in codex. And, well, that's because they're quietly starting to test a new scratch pad feature for codex that would let users run multiple tasks in parallel, a move that points to a big expansion of codex beyond just coding and moving into a central hub for AI driven work. So, yeah, essentially, this Scratchpad, you type out a bunch of things. Right? It can be notes, and then they turn into chats, and then codex just does them. So think if you were to, you know, leave yourself to dos, and then codex just does them. Right? So pretty cool.

Jordan Wilson [00:04:24]:
So these references suggest that OpenAI is also consolidating ChatGPT, the Atlas browser, and the different software in agentic engineering tools into the single super app, which we've been talking about over the last couple of weeks. But it does, appear anyways, that codex may be kind of the final landing platform. I'm not sure if that's, ultimately gonna be true if they're gonna name it something else, but it does look like a lot of these rumored features of bringing, the, the different functionalities all into one are at least being debuted in codex. So we have seen reports that codex may be the ultimate winner here, at least when it comes to, desktop software. So, yeah, if you don't know, you could use Chat GPT on the web, but also the desktop software. You have the Chat GPT app, and then you have the, the codex app. So, in the new, kind of, leaks here from OpenAI, one of the most telling discoveries is a heartbeat system designed to maintain persistent connections with long running tasks. And that well, if it sounds like OpenLaw, yeah, that's because that approach does close the mirror systems already used by OpenLaw and Anthropic's, managed agent Project Conway, which we talked about last week, making OpenAI's move a clear competitive response as a desktop play.

Jordan Wilson [00:05:51]:
Separately, social media posts from OpenAI employees featuring snowflake emojis. Yeah. We're talking about snowflake emojis here on the show. Have sparked speculation about a new model release codenamed Glacier, that some believe to be GPT 5.5. Raising the possibility that OpenAI could pair a major platform launch with a model upgrade in the coming days. So, yeah, maybe we'll see the, the rumored GPT 5.5. Maybe we'll see the full super app, or maybe this week, we'll just see a codex, release with some of the, other features kind of baked into codex. My guess would be the latter, but on this one, my guess is as good as yours.

Jordan Wilson [00:06:36]:
Don't have any, inside intel on this one at least. Alright. Next, Anthropic has launched their new Claude managed agents to make AI agents easier to build, run, and scale. So it is in public beta, and it's offering a full production stack that lets developers build and deploy cloud hosted AI agents without managing infrastructure themselves, which makes this a notable step toward, more practical enterprise ready agentic AI. So if that sounds super confusing, well, it might be. So you do have to use this on the back end in Anthropic's platform. So you're not using this in Claude AI, FYI. Right? So you're not gonna go to claude.ai.

Jordan Wilson [00:07:21]:
You're gonna be using Anthropic's platform. So the good thing is, well, you don't have to have a paid cloud account to do this. You just have to, at least have a credit card on file because you will be charged for usage. So, yes, you can do this more on the technical side, but the cool thing is if you've used the, the GPT builder in ChattGPT, it's kind of like a sort of like a version of that. Right? So you can simply chat with Claude, to help you build agents in this new Claude, managed agents, but you can also go a little bit more technical and under the hood. Right? And the cool thing is it can connect to basically any MCP server. It can connect to anything. Right? So in the same way that you might use cloud code, and you might not know how to do any of this coding, but it's using the terminal and connecting to all these API services and doing all these magical things.

Jordan Wilson [00:08:12]:
That's kind of what, Claude managed agents looks like inside of INTROPICS platform. So I did get to play with it for a little bit, I think on Friday. So I haven't spent multiple hours, but it does seem like a pretty simple way to build agents, but then to have them contained in Anthropic's, kind of sandbox. You don't have to worry about deploying it out on your own. So the platform handles sandbox code execution, credential management, scope, scoped permissions, checkpointing, and end to end tracing. Meaning teams can focus on defining the task, tools, and guardrails, while Infoprix orchestration system manages the tool use context in error recovery. So cloud managed agents also support long running autonomous sessions that persist through disconnections and include multi agent coordination, allowing one agent to spin up and manage others to paralyze, sorry, payroll allies. That's a hard word to say.

Jordan Wilson [00:09:09]:
Right? Complex work. So sounds great in theory. Right? And it is. However, I will warn you, running parallel agents is great. Right? Especially if you're using cloud code or if you're using codex. Just keep in mind, if you are using cloud managed agents, yeah, all those spinning up of sub agents is, yeah, gonna cost you because you're paying via usage. You're paying via tokens. You pay paying via the API.

Jordan Wilson [00:09:40]:
So, keep that in mind. Sounds great, and it is. Right? I've I've tested it and, you know, I instantly had an agent that connected to, you know, my email newsletter and, you know, all these other services that had, MCP, data. Right? Which is great. So you can just say, hey. I have all these services. You know, go connect to them. It'll bring up an an authorization page.

Jordan Wilson [00:10:02]:
You click a couple of things, and all of a sudden you have an agent. Right? Let's say there's five pieces of software that you use all the time. Right? And you're like, okay. I could, you know, try to piecemeal this together or, well, this is where, this new release from Anthropic really, really works because not only will it just kind of build it for you. Yes. You do have to off, you know, authorize, the agent, but then you can just run it in the sandbox. But like I said, the cost will add up fairly quickly. Alright.

Jordan Wilson [00:10:32]:
A new deal for superintelligence. That is our next story because OpenAI published a 13 page policy document titled industrial policy for the intelligence age, ideas to keep people first. So, this is what a lot of people are calling the superintelligence new deal, and it outlined how governments should tax, regulate, and redistribute wealth from AI as the technology rapidly reshapes the economy. So the blueprint argues that AI progress is accelerating so quickly that The US may need a new social contract, right, comparable to the progressive era or the new deal, to address risks like mass job displacement, cyberattacks, and social instability. So OpenAI proposes bold new ideas, including a national public wealth fund funded partly by AI companies, taxes on automated labor to replace shrinking payroll taxes, and a four day thirty two hour work week that shares AI productivity gains with workers. So the document also calls for treating AI access as a basic right for workers and schools, creating containment plans for dangerous autonomous systems, and triggering automatic expansions of unemployment and wage supports when AI driven disruption hits preset levels. So parts of this, I think were really good. Right? And if you didn't get a chance to read this, we shared this in our newsletter last week, but that's why you should be subscribing to our newsletter.

Jordan Wilson [00:12:06]:
So I think parts of this are great in theory. Many of these things will never see the light of day, because many of them require the government to act in some official capacity. And this is coming from someone that used to cover, the government as a journalist. The government doesn't work like that, especially today's, federal government. I don't think anything of this magnitude, we'll see the, the light of legislation, in the next, I don't know, three to five years. Right. So what we should really be following is the states, and we will see if states, you know, adopt anything like this. Obviously, I would keep an eye on California, which is where all the big tech companies mostly all the big tech companies are headquartered.

Jordan Wilson [00:12:55]:
So a couple of things I kinda wanted to point out. Right? Like, the robot tax, very popular. A lot of people have talked about that. That makes sense. And you do have to, I guess, tip your hat to OpenAI for saying, like, okay. Yeah. Like, if AI takes all these jobs, we need to have money to help all the humans, and we should be taxing the robots. You know? That makes sense.

Jordan Wilson [00:13:16]:
But then on the other hand, you know, they're essentially saying that AI needs to be deemed a basic right. So, you know, on one hand, they're like, okay. Well, this thing that we're selling, you know, we need to call it a basic human right. But at the same time, we're like, we know it's probably gonna take a lot of jobs, and so we need to do something about that. So, I've talked about this a lot over the course of the last three years. I'm not gonna bore you with my, hot takes, but, you know, overall, I do think, AI is going to change what full time employment means in The US. I think, ultimately, AI will replace, more full time jobs than it will create. But I do think the future of work as well, a lot of people that aren't even entrepreneurs, they're gonna have multiple knowledge working side hustles.

Jordan Wilson [00:14:03]:
Right? So I don't know if you're, a lawyer, maybe you get laid off from your law firm instead of being a full time employee. You might just have 10, you know, very niche lawyer side gigs. Right? Yeah. It's kind of the way I see things shaking out, but, yeah. We'll see. Alright. Next piece of AI news. This one was kind of shocking.

Jordan Wilson [00:14:27]:
Yeah. Metta has a new model and it's actually pretty good. Yeah. So Metta announced their new muse spark. That's their new AI model after investing a ton of money and a ton of time. We're talking billions of dollars and more than a year. So like I talked about on our Friday features, I did get to sneak this one in on our, new Friday features. But, you know, I said, it's been a year since Meta released llama four.

Jordan Wilson [00:15:00]:
In in AI time, that feels like a decade. Right? It seems like almost everyone wrote meta llama off or sorry. Meta off because they didn't really come up with anything after Llama, but we knew that they had some big shifts internally. And it looks like their first model anyways, fairly impressive. So, yeah, the company offered, you you know, they had an acquihire of more than $14,000,000,000 for Scale AI and its CEO, Alexander Wang. Then the company reportedly offered some engineers paid pay, pay packages worth hundreds of millions of dollars to staff the new, MSL or the Meta Super Intelligence team. So the, Muse Spark is the first model in a series, that was known internally first as Avocado. So that was the codename.

Jordan Wilson [00:15:48]:
So if you've been listening to the show, we've been talking about that. And right now, it's initially available only on Meta's AI app and their website. And they do have plans to essentially replace the llama models anywhere with the new, Muse Spark. The other thing to keep in mind, well, unlike previous open releases via the llama series, the new Muse Spark is not open source. So it is closed. It is proprietary. Right now, it's only for free. Right? So presumably, that will change.

Jordan Wilson [00:16:23]:
And right now, it is not available via the API, although the team at Meta did say that they will be rolling out, the API soon. So according to independent evaluations from artificial analysis, Muse Spark already matches top models from Google, OpenAI, and Anthropic in language in visual task, but falls behind in coding and abstract reasoning, tying for fourth place in the broad AI test. Yeah. I was actually, fairly shocked. Right? So, if talk about artificial analysis a lot on the show. It's essentially it's kind of like an aggregator. Right? So it takes all these different benchmarks and all these different scores from all these different places and gives all the models a score. Right? So right now, Google and OpenAI are tied, with their respective models.

Jordan Wilson [00:17:16]:
And then in technically second place, you have Claude with, Opus four six. And now in third, technically, you have MuseSpark. Right? Which is pretty impressive. The other thing you have to think, I think people are looking at this a little bit differently because Meta did say that they've rebuilt this model from the ground up. Right? So this is not, according to Meta, just a new version of llama that's been improved upon. This is what Meta says, a built from scratch new model. And the fact that it's doing that well already, a, just one point behind clawed at Opus 4.6 on the artificial analysis. And what is maybe even crazier, on arena.

Jordan Wilson [00:18:04]:
Right? So we talk about arena formally, LM Arena. So this is the blind taste test, and it's also third right now on LM Arena. Although that could change at any second because it's only by, like, one point. But, regardless, it's a top five model, by benchmarks and by user preference, which if I was, putting money on this beforehand, I would have said they were probably gonna be in more of the five to eight range. So fairly impressive. And a lot of people were kind of dragging Meta, right, because they released their benchmarks, and they're like, okay. Well, Meta released all these benchmarks, and they're not even really top on any of them. But when you think about it, this is technically their first model in this series, and it's, you know, top two, three, four depending on what you look at.

Jordan Wilson [00:18:51]:
I don't know. I'm impressed. I've used it. My actual usage is mixed, Right? Because I'm a very heavy g p t five four pro user, and I was giving it very complex tasks. You know, there is also a new con what is it called? It's called contemplating mode, right, which kind of runs these, multiple agents simultaneously. So that was the thing I was, like, really looking forward to because I'm like, oh my gosh. This thing runs, you know, 16 agents at a time or something like that, and it's supposed to be comparable to, you know, Gemini Deep Thinker, OpenAI's GPD five four pro. To me, I wasn't as impressed with the, the new contemplating mode, but I was maybe more impressed actually with its coding abilities, its writing abilities.

Jordan Wilson [00:19:35]:
So, yeah, you have a new model to try out at least. Alright. Going from a impressive model to a company that is maybe not impressed with its current AI outputs. That is because according to reports, Microsoft is under a Copilot code red. Alright. So, and this is according to BNP, Paribus analyst, Stefan Slowinski, who reported that Microsoft CEO, Satya Nadella, has declared a Copilot's code red inside of Microsoft, signaling an all out push to enhance Copilot's performance and user experience. So the urgency comes as investors express frustration over Copilot's limited traction despite Microsoft's leadership in software in general. So Nadella's initiative reportedly includes the upcoming launch of the e seven suite, which, I believe should be here around the May, with ongoing updates and new features planned throughout the year to accelerate Copilot's adoption and usefulness.

Jordan Wilson [00:20:42]:
So according to Slowinski, the initial feedback on Copilot is improving, suggesting Microsoft's renewed focus could pay off as it leads to better user satisfaction and market perception. But the competitive, threat from rivals such as Anthropic is a major reason behind the Code Red strategy as Microsoft aims to stay ahead in enterprise AI tools. So, Slowinski also noted that Azure could still outperform expectations due to growing demand for tokens and higher GPU pricing, even if internal usage increases further. Here's the thing with Microsoft. Right? It's no secret that the enterprise has been rather frustrated with Copilot. Right? They were, one of the first out of the gate. Right? You technically had Chat GPT first, but, I mean, Copilot was the first, like, serious enterprise business AI tool. And I think a lot of enterprises who adopted early and invested heavily, right, in 2023 and 2024, maybe they've been disappointed in the last, you know, two years or so as you've seen Google anthropic in, Google anthropic and OpenAI really just take off.

Jordan Wilson [00:22:00]:
However, if I'm Microsoft, I'm not exactly worried. Right? They're the only company that has the, you know, the green flag at the for the most part across the entire enterprise. Right? It's much easier, for Copilot for Microsoft Copilot, to break its way through the enterprise. Although, obviously, Google, OpenAI, and others have been really cracking that space. But in the end, I'm not super concerned. Top level, if I'm Microsoft, yes. You gotta make Copilot better. Yes.

Jordan Wilson [00:22:32]:
A lot of people don't enjoy using it. Yes. A lot of Copilot users, are jumping ship, specifically to OpenAI, into Google. But I don't know. Microsoft's a big investor in anthropic. Microsoft is the biggest single investor in OpenAI. So yes, it's bad if they're losing, if they're losing users to Microsoft or, or sorry, if they're using loose my gosh, I can't speak today. If they are losing users, to OpenAI or Anthropic, but in the end, they're still just making money off that anyways.

Jordan Wilson [00:23:11]:
So, we'll see if this, Copilot code red leads to anything. We did see similar, stories, earlier this year that, you know, Saturday was going full PM mode. Right? Like, product manager, he's rolling up his sleeves, sitting down with the product team. So, I'm actually and I'd like I told some people this, had an in person event last week in Chicago, and I told people this. Like, I'm actually bullish on Microsoft. I've seen a lot of what they've released the last couple of weeks. Right? They're essentially what they're doing. I I'm not gonna say they're white labeling a lot of products.

Jordan Wilson [00:23:48]:
Right? But they came out with a version of Copilot, their Copilot co work, which is very similar to Anthropic co work. It's really good. Right? They have their new task feature, which is really good. Similar to, some features on Anthropic and OpenAI, just schedule tasks. So, I think Microsoft has actually been shipping a lot. I think those companies that maybe haven't found, that utility in Microsoft Copilot, it's actually more of a training and education problem, versus a model problem because now you get the best of both OpenAI and Anthropic when you're using Microsoft. Alright. Let's get to some scary stuff happening to OpenAI CEO Sam Altman.

Jordan Wilson [00:24:28]:
Yeah. This was shocking about to read about over the weekend. So OpenAI CEO Sam Altman's San Francisco home was targeted twice over the past, four days, raising concerns about the risks facing tech leaders in the AI sector. So the latest incident happened early Sunday morning when suspects in a car allegedly fired a round of shots at Altman's property before fleeing the scene. So police quickly traced the vehicle using surveillance footage and arrested two suspects later that morning. So officers searching the suspect's residence found three firearms, and both individually were booked for negligent disarm discharge of a firearm. So this attack followed a Friday morning incidents in which a 20 year old man from Texas allegedly threw a Molotov cocktail at Altman's home. So security at Altman's property extinguished the fire from the Molotov cocktail, and no injuries were reported by either or in either incident.

Jordan Wilson [00:25:34]:
But the two attacks come as Altman has publicly voiced concerns about the societal impact and anxiety surrounding AI, calling it the largest change to society in a long time. So the rapid succession of attacks underscores the growing tensions and security risks for leaders at the forefront of AI development. So Altman did respond in a blog post after the, incident on Friday, and he was also critical of a New Yorker article that questioned his trustworthiness acknowledging the impact of those negative narratives. So Altman did admit past mistakes, including, being, you know, conflict diverse and mishandling issues with the OpenAI board, but emphasized his commitment to improving OpenAI's mission. He called for less dramatic rhetoric in the AI industry, advocating for broad technology sharing and urging constructive debate to avoid further real world harm. Here's here's the harsh reality. Right? I'm I'm gonna say this is someone that lives in Chicago. And that's important because I think maybe the majority of our listeners are not from Silicon Valley.

Jordan Wilson [00:26:47]:
Right? But I know, you know, there's other, you know, popular tech publications where the majority of people are from Silicon Valley. Silicon Valley is a bubble in a bubble. Right? I don't quite think that Silicon Valley and all the big AI frontier labs really understand what the rest of truly understand. Right? Because I don't think you truly understand unless you live it, what the rest of the world or what the rest of The US feels about AI. And the reality is most people don't want it. Most people don't like it. Most people view AI as a threat. So unfortunately, this is a in an extremely unfortunate incident that happened.

Jordan Wilson [00:27:31]:
But I think that we're gonna continue to see AI leaders from all the big companies. I think this is gonna be unfortunately an ongoing issue, their literal safety. Right? Because as people start losing their jobs to AI, right, you can't just get mad at the cloud. Right? Unfortunately, it's people like Sam Altman, people like Dario Modi, from Anthropic, people like Sundar Pichai, you know, people like Satya Nadella, it's the faces of these big, you know, four or five companies, you know, Mark Zuckerberg at Metta as well. These are the people that people are going to be angry at. Right. Because unlike, you know, the internet, there was really no face of the internet, I guess you could say maybe Bill Gates. But ultimately, the Internet was a very slow change to jobs.

Jordan Wilson [00:28:28]:
It was a slower change to the economy. Yes. You had the .com boom and bust, but things with AI are moving much, much faster. And I don't think that people, in Silicon Valley necessarily, largely understand how the rest of The US really feels about AI. And, yeah, I think that, unfortunately, we're gonna see ugly incidents, and I don't want it to happen. Right? And I I hope all the, you know, leaders of these AI companies stay safe because ultimately, I I am very optimistic about AI's future and doing, more good than bad. You know, hopefully, it's able to cure diseases and do all of these great things, but, yes, it's going to cause a lot of unemployment at the same time, and people are going to be mad. So this is terrible.

Jordan Wilson [00:29:18]:
I hope it doesn't happen again. But, unfortunately, I do, think that the leaders of AI tech companies are gonna have to be, you know, doubling up their security, as the rest of the, kind of US finally sees what AI is capable of in terms of job displacement. Alright. Last but not least, more scary stuff, a model so scary, Anthropic can't release it. So Anthropic has announced its new Mythos preview model, which they say is so powerful at finding software vulnerabilities that the company is keeping it private, raising concerns about both cybersecurity and access to advanced technology. So Anthropic said its new Mythos preview model has found thousands of critical vulnerabilities, across major operating systems and web browsers, including a twenty seven year old flaw in OpenBSD and a sixteen year old bug in FFmpeg, all that have previously gone undetected. So, essentially, they're saying that, their new Mythos model is a cybersecurity whiz, and it's able to find thousands of these, you know, zero day bugs that, you know, millions of human researchers could never find. But the company is not, at least for now, releasing the model publicly.

Jordan Wilson [00:30:46]:
Instead, they have their new project Glasswing, which is essentially a group of companies that they're giving access, to mythos, to these companies. And they're essentially saying use this to harden up, your software to make, you know, your software better because when a model like this kind of hits the streets, right, we want these, you know, big tech companies to be safe, and we want the technology that people use, to not be exploited by a model like Mythos or similar. Right? So the company right now is sharing it only with partners such as Apple, AWS, Google, NVIDIA, Microsoft, and 40 other organizations as part of project Glasswing, and that is kind of their defensive cybersecurity initiative. But this move marks the first time in the modern AI era that a major model is being withheld from the general public, but is being released privately due to concerns over its potential misuse, creating a significant knowledge and technology gap between elite companies and the broader public. So, yes, there was times early on. Right? Like, even I remember OpenAI way back in the day because I was using their, you know, their early, GBT. I I forgot if I was using GBT two or GBT three, technology. Right? Like, back in 2020.

Jordan Wilson [00:32:06]:
I remember there was a time they were like, oh, we're not gonna release this model because, you know, it could, you know, write lies about people and, you know, they eventually released it. It wasn't that they just released it to 40 companies. So there has been time in the past when companies have said something like, oh my gosh. Our model's too good. We're not gonna release it. But they eventually did release it. Right? This one with INTROPIC, presumably, they'll eventually release a version of Mythos. Maybe it's a stripped down version, but it seems like at least for the short or medium term, for the first time, there's a huge tech divide.

Jordan Wilson [00:32:39]:
Right? There is, you know, the the democratization of AI may no longer be a thing anymore. Right? So it's like, oh, we had a great run for the last, you know, four or five years when, you know, the Fortune 100 companies, you know, were using the same thing as, you know, small mom and pop shops. So that time may not be gone, with this new Mythos model. So the company claims that Mythos was not intentionally trained to be a cyber threat, but its advanced coding abilities led to the discovery of vulnerabilities that the even the top human experts and previous AI tools missed. So what's my take on this? I mean, I did a whole episode, so you can go listen to that seven fifty two. So I'm not going to spend too long on it. I mean, pardon me, I think Anthropic made the right move here. Right? If they are truly actually concerned, about this, being a a model, that could be a cyber threat.

Jordan Wilson [00:33:38]:
Okay. That's great. To me, I don't know. You know, Anthropic has had, in its CEO, I've had a lot of, I won't say boy that cried wolf, but they've had a lot of, instances in the past where they're really hyping things up. Right? They're like, oh, you know, AI is gonna, you know, take all coding jobs. And then they said, AI is gonna take, you know, half of white collar jobs. And, yeah, those things might ultimately come to fruition. I don't know.

Jordan Wilson [00:34:07]:
But to me, it seems like this was a strategic play with Glasswing. Right? You get everyone talking, you know, about how that's this new dangerous model. Right? And, you know, I I don't know. To me, I think Anthropic had a huge and embarrassing data leak, right, a couple of weeks ago. And they know that they're gonna be going for an IPO here, presumably in quarter three or quarter four. And they need something. Right? They need something in between that. You can't have your last big, you know, international, splash on the news radar to be, oh, that time you accidentally leaked your source code to your most popular product on the Internet.

Jordan Wilson [00:34:49]:
Right? That's not a good time. And then, like, four months later, you know, at least outside of the AI scene. Right? Like, we're talking about AI, you know, or we're talking about anthropic every day, but I'm saying the entire world, right, the entire world was talking about anthropic at that code leak. Anthropic needed something, I think, to divert the attention from, oh my gosh. We accidentally just, leaked the source code to our most popular product called code. Right? And we're getting ready for an IPO here. We need to start to spin up a new narrative. So, you know, now it seems like this new narrative, whether it's a 100% true, 50% true.

Jordan Wilson [00:35:25]:
I don't know. Right. I don't know. I would say it's 50% true. They are actually, concerned about, you know, releasing this publicly because, yeah, it could create a lot of, a lot of, bad bad actors. We'll just say that. Right? With all the software that we use, yet at the same time, I do think this is a little bit of pre IPO marketing and, you know, just trying to, flex on everyone and saying, yeah, look at how good our models are. Alright.

Jordan Wilson [00:35:51]:
So that is it for the big stories of the day, but we're gonna end or sorry, of the week. We're gonna end with our what's new and what's next. So this is a combination of, you know, some leaks, some rumors, and just, you know, some pieces of news, some updates that came out this week, that, you know, we just didn't give, we didn't have enough time to give a full full attention to all of these. So, we're gonna go quick here. So, we're starting Google Ads notebooks inside of Gemini with full bidirectional sync to notebook l m. Yeah. So it's kinda like projects, but it also works with notebook l m. Pretty cool.

Jordan Wilson [00:36:29]:
So leaks show that Anthropic may be building a lovable ask full staff software building program. That would be crazy. Brad Gerstner of Ultimater Capital said that companies are already using OpenAI's spud model, and it rivals Claude's mythos. OpenAI launched a $100 a month chat GPT pro tier with 10 x codex access. So, yeah, if you didn't wanna pay the $200 a month, but you wanted more than the $20 a month pro plan, now you have the mid tier, $100 a month pro tier. Apple is testing for premium material smart glasses design, powered by AI, and paired to your iPhone targeting a launch in 2027. Spotify now creates podcast playlists from natural language prompts. Just I don't know.

Jordan Wilson [00:37:20]:
Maybe you ask for the best everyday AI episodes. Try that. Alright. Alibaba's happy horse one unexpectedly took the top spot in the AI video arena. Yeah. It's looks really good. Better than CDANS, better than v o three, sorry, v o three one at least for now. Speaking of models that made a splash, z a I released their GLM five one, which is not only the new, state of the art model for open source, but it also outbenched top frontier models on SuiteBench like GPT five four, Opus four six, and Gemini three one Pro.

Jordan Wilson [00:37:57]:
That's huge. Right? An open source model. I mean, you gotta have, like, a supercomputer thing to actually download this thing. Right? But it's open source. And it outperformed the big three on Swee bench, which is one of the most, popular benchmarks for software engineering. Microsoft updated its copilot terms, because previously, they had that copilot was for inter, entertainment purposes only. So, yeah, there is some criticism on that, so they changed it. Next, a DC court allowed the Pentagon to blacklist Anthropic, but other agencies can still contract with Anthropic.

Jordan Wilson [00:38:35]:
So, yeah, the ongoing, kind of battle, might be closer to being closed. We'll see, if that gets, appealed. Next, according to an Axios report, OpenAI is projecting a $100,000,000,000 in ad revenue by 2030. Nebius is in talks to acquire AI twenty one Labs for up to $3,000,000,000 according to reports. OpenAI is partnered with Upwork so users can hire freelancers directly in Chattopty. Goldman Sachs, Goldman Sachs came out with a new report that said AI displaced workers face lower earnings and higher unemployment risk for a decade. Ella Marina has released the full history of its AI leaderboards as a public dataset. OpenAI is testing a new image generation model on Chad GPT, AB test, and LM arena.

Jordan Wilson [00:39:30]:
We talked about them testing it on LM arena, but now they're also testing it inside of Chad GPT on AB test. So that would presumably be their new v two of their image model. Google Workspace launched a feature where Gemini suggests the best meeting times for everyone. I I mean, we've been needing that for, like, twenty years. Alright. Next, Pico launched an AI self video chat beta where your AI agent talks, remembers, and acts in real time. Google quietly dropped this one. It's called the AI Edge Eloquent.

Jordan Wilson [00:40:02]:
It is a free offline dictation app on iOS using Gemma models. That is it's super impressive too, FYI. Speaking of Google, they're preparing a Jules v two coding agent that can set goals and drive improvements without prompting. The Gemini app now lets you create interactive three d simulations and models inside of the actual chat, which is really cool just to visually explain things. Google also expanded its finance tools globally with new AI capabilities. QuadcoWork hit general availability for all paid plans, and Meta signed a $21,000,000,000 deal with CoreWeave to expand AI cloud capacity. That was a lot. Alright.

Jordan Wilson [00:40:44]:
I hope this was helpful. Sorry. I got a little tongue tied there. It just happens. Right? There's so much going on. Even I struggle to talk about it. So don't spend hours every single day trying to keep up. Join us on Mondays as we bring you the AI news that matters.

Jordan Wilson [00:41:00]:
If you are newer here, right, on Wednesdays, we go hands on. We usually do a deep dive on one, tool. So make sure to check out today's newsletter. We'll probably do a poll on that. So what do you wanna see? And then on Fridays, we do our AI feature Fridays, which is where we usually do a handful of new features that you can start using now. And on Tuesdays and Thursdays, we kind of rotate our shows. So I hope this was helpful. If you're listening on the podcast, do me a favor.

Jordan Wilson [00:41:28]:
Leave a review for us. I'd really appreciate that after you subscribe to the podcast. So thank you for tuning in. If you haven't already, please go to your everydayai.com. Sign up for the free daily newsletter. Thanks for tuning in. Hope to see you back tomorrow and everyday for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI