Resources:
Join the discussion: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Unlocking Practical AI Shifts: What This Week’s Major AI Developments Mean for Business Leaders
This week’s AI news cycle brought several key developments that carry immediate and tangible implications for business owners and decision makers. From billion-dollar partnerships to evolving regulatory strategies and advanced AI tool integrations, the signals from top industry players provide a strategic roadmap—if you know where to look.
Disney’s $1 Billion OpenAI Partnership: AI-Powered Content Licenses and Distribution
Disney has taken a decisive leap by investing $1 billion in OpenAI. The entertainment giant’s move provides OpenAI with access to Disney’s high-value intellectual property—including characters from Star Wars, Marvel, and more—for integration into OpenAI’s Sora, an AI-driven video generator. What’s critical here is the business architecture: more than 200 animated characters, costumes, and props will be licensed for use in minute-long AI-generated videos slated for Disney-owned platforms.
This agreement signals the normalization of AI-generated, licensed IP in mainstream entertainment. From a business perspective, it sets a playbook for maximizing IP value and establishes a model for legacy media partnerships with AI firms. This arrangement also positions OpenAI as a core technology provider inside Disney’s operations, with ChatGPT’s productivity suite expected to become standard for employee workflows.
Moreover, Disney is asserting its IP position on multiple fronts—partnering with OpenAI for exclusive access while simultaneously sending Google a cease-and-desist letter for alleged copyright infringement related to GenAI models. This “partner, protect, and pursue” strategy will likely shape how other IP-rich companies engage.
AI Federal Regulation: Centralized Over State-by-State Patchwork
An executive order from President Trump this week is set to override state-level AI regulatory powers, aiming for federally centralized oversight. This shift comes when the U.S. lacks comprehensive federal AI laws but faces more than a thousand state bills in play.
Tech sector lobbying for this approach is directly related to operational efficiency. Inconsistent state-by-state regulations could slow down development and diminish the U.S. competitive stance, notably against China. For decision-makers, this move hints at imminent nationwide harmonization in compliance requirements—especially relevant for companies with operations or clients across multiple states. The caveat: without Congressional action, current rules may remain vague, with few federal “guardrails” protecting consumers or businesses in the short term.
Meta’s Strategic Shift: Open Source to Proprietary AI
Meta’s move towards closing the source on its next-generation AI models marks the potential end of open access for external developers. Where past Llama models could be freely downloaded, fine-tuned, and forked by any sufficiently-equipped organization, future “avocado” variants are projected to be accessible only under proprietary terms.
This is more than a tech trend—it’s a market signal. Meta’s leadership, after lackluster Llama 4 uptake and concerns about IP leakage to competitors, is opting to contain future advanced models. The company is investing up to $72 billion in 2025 for talent and infrastructure, recently onboarding Scale AI’s CEO to drive next-generation development. Businesses relying on open-source models need to reassess their roadmaps, as open architectures may dwindle among Big Tech providers.
OpenAI’s GPT-5.2 Release: Enhanced Enterprise Productivity and Usability
GPT-5.2 stands out in both capability and strategic timing. Benchmarks show this model now surpasses human professionals in over 70% of tasks tested—ranging across 44 mainstream occupations—and does so at speeds 11x faster than before, as measured by OpenAI’s GDP-val standard. Notably, GPT-5.2 also resolves former shortcomings in spreadsheet and PowerPoint generation, raising the bar for day-to-day business automation.
A technical leap in context window retention means GPT-5.2 can extract “needle-in-the-haystack” insights from enormous data sets—nearly 100% accuracy at 256K tokens—allowing for more reliable analysis across lengthy documents or conversations. For enterprises handling compliance, research, contracts, and intricate knowledge work, this reliability in long-form recall and precision should recalibrate expectations of what AI can automate or accelerate.
Enterprise Workflow Integrations: Adobe Apps Inside ChatGPT
Adobe and OpenAI’s integration deal now allows Photoshop, Acrobat, and Express to operate natively within the ChatGPT interface. End users can edit photos, design collateral, and manage PDFs simply by describing tasks in natural language, then switching context across multiple apps without losing workflow history.
The significance is operational: creative and document management tasks that previously required standalone, resource-heavy applications can now be performed interactively within a single AI-driven platform, with context and history preserved. This not only boosts efficiency but also paves the way for business teams to execute multi-step projects in one interface—reducing time to publication and lowering hardware requirements.
Industry-Wide AI Agent Framework: The Agentic AI Foundation (AAIF)
A newly launched Agentic AI Foundation, organized under the Linux Foundation, brings together otherwise competing organizations—OpenAI, Google, Microsoft, Amazon, Bloomberg, IBM, and more—to collectively define open standards and interoperability protocols for AI agents. Projects like Anthropic’s Model Context Protocol (MCP), OpenAI's Agents MD, and Block's Goose provide the skeleton for inter-agent communication and standardized tool access.
For enterprises, this means the near future will likely feature agent-based systems that natively cooperate across vendors, platforms, and internal tech stacks. Standardization mitigates vendor lock-in and forms a common denominator for enterprise adoption of AI-powered automation, multi-agent workflows, and cross-platform digital solutions.
Next-Level Agentic Research Tools: Google Deep Research Upgrades
Google’s release of its Gemini Deep Research agent as an embeddable tool for third-party developers expands enterprise options in specialized research automation. This agent can process massive volumes of context—already deployed in high-stakes domains like financial due diligence and drug research—and is slated for integration into Google Search, Google Finance, Gemini apps, and NotebookLM.
This shift signals imminent verticalization: expect new domain-specific research and analytic agents to emerge for industry niches, unlocking time and cost savings for high-complexity, information-rich business functions.
Key Takeaways for Decision Makers
IP Licensing-as-Strategy: The Disney-OpenAI model sets a precedent for monetizing licensed assets via AI, while asserting legal boundaries against unlicensed use elsewhere.
Proprietary AI Model Access: Strategic pivots away from open source among major platforms will reshape accessibility for external developers and enterprises; review dependencies accordingly.
Centralized AI Policy: Businesses should anticipate a move toward unified compliance frameworks for AI, reducing uncertainty for interstate operations in the US.
AI as Workflow Hub: Integrations between foundational business tools and AI platforms drive next-level operational efficiency, reducing context-switching and turnaround times for end users.
Standardized Inter-Agent Protocols: Adoption of open standards lowers the barrier to deploying complex, agent-driven enterprise systems—an inflection point for scalable AI adoption.
Research Automation: Expanded API and agent options for industry-specific research will directly impact knowledge-intensive verticals.
These developments highlight an operational and policy environment in flux—rapid adjustment will be necessary to capture business value as the AI landscape continues to evolve.
Topics Covered in This Episode:
- Disney Invests $1B in OpenAI Sora Partnership
- Disney Grants OpenAI Access to Copyrighted Characters
- Trump Executive Order Blocks State AI Laws
- AI Regulation: Federal vs. State Controversy
- Meta Shifts Llama Model Open Source Strategy
- OpenAI CEO Teases Upcoming ChatGPT Features
- Google Releases Gemini Deep Research Agent
- Adobe Photoshop Integrates with ChatGPT Apps
- Agentic AI Foundation Unites Top AI Rivals
- OpenAI Launches GPT-5.2 Model Benchmarks
- GPT-5.2 Performance: Spreadsheets & Presentations
- Long-Context AI Model Retrieval Improvements
- Anthropic Claude Agents and Accenture Partnership
- OpenAI, Google, Anthropic Update Agent Frameworks
Episode Transcript
Jordan Wilson [00:00:46]:
What a weird week in AI. I mean, we got arguably one of the most powerful AI models in the world dropped in a tweet. Meta, the open source AI company might be going closed source. President Donald Trump is trying to take away the state's powers when it comes to AI. Disney is going all in on AI and the biggest AI competitors are all teaming up for an AI project. It's like backwards week in AI, and it might leave you scratching your head with a couple of questions. Well, we're gonna be giving you a lot of answers today on Everyday AI. What's going on, y'all? Welcome.
Jordan Wilson [00:01:26]:
My name is Jordan Wilson, and this is Everyday AI. This is for you. If you've ever felt confused by the onslaught of AI news, well, Everyday AI is for you, but specifically our Monday segment AI news that matters. So if you're new here, everyday AI is a daily livestream podcast and free daily newsletter helping everyday business leaders like you and me make sense of all this AI and how we can leverage the practical and, actual stuff to grow our companies and our careers. So if that's what you're trying to do, awesome. It starts here with the unedited, unscripted livestream podcast, but take it to the next level. Make sure you go to our website at youreverydayai.com. In our free daily newsletter, which you've gotta go grab on our website, we're gonna be recapping the highlights from today's episode, as well as everything else happening in the AI world today.
Jordan Wilson [00:02:14]:
And just a little announcement, make sure you join us tomorrow and Wednesday. So that is December 16 and December 17 for our 2025 AI road map rewind. So back in January, we made 25 kind of crazy AI predictions. And here we are. The test is due. Did we fail? Did we lead you astray? Right? Because if you go back and look at it, literally go back and look in the comments. People were like, this is crazy. This isn't gonna happen.
Jordan Wilson [00:02:48]:
Right? I think some of it kinda did, but, you gotta make sure to tune in. We did it originally. It was like five really quick shows. We're not gonna do five recap shows. We're just gonna jam it all into, two shows, so make sure to join us for that. Alright. Let's get into the AI news that matters for the week of December 15. And first one yeah.
Jordan Wilson [00:03:09]:
This one's a big head scratcher. Disney and OpenAI shaking hands. So the Walt Disney Company announced a $1,000,000,000 investment in OpenAI, making the entertainment giant one of the most prominent legacy media companies to formally partner with a leading AI firm. So this deal grants OpenAI access to Disney's copyrighted characters from franchises, including Star Wars, Marvel, and other properties for use in its AI powered short form video generator, Sora. Disney plans to allow more than 200 animated characters along with costumes, props, and accessories to appear in Sora generated videos that can run up to a minute long. And Disney is gonna reportedly be playing these on its platforms. Weird. Something I predicted back in January, but no one believed.
Jordan Wilson [00:04:08]:
Not specifically this, but that you see AI generations on the big screen, and it would look real. Disney CEO Bob Iger said the agreement reflects a push to extend Disney's storytelling through generative AI while also protecting creators' rights and safeguarding original works. As part of the deal, Disney will become a major OpenAI customer and plans to integrate AI tools across its business operations, including making ChatGPT available to employees. So Sora and ChatGPT are expected to begin offering licensed Disney character content early next year, marking one of the first large scale uses of licensed Hollywood IP in generative video. This one's important because on the same day, Disney also issued a cease and desist letter to Google accusing the tech giant of massive copyright infringement through its GenAI models. You know, I I said this many years ago, like, for the most part, yeah, these big media companies, small media companies, everyone, you're gonna have a couple of choices when it comes to generative AI. You either partner, you sue, or you lose money and maybe go out of business. So Disney chose one of each here.
Jordan Wilson [00:05:27]:
They partnered with OpenAI. They invested a billion dollars for an equity stake in the company and also granted some exclusive access to OpenAI. Yet on the same day, they sued Google and said, hey. You're they sorry. They didn't sue Google. They sent a cease and desist letter. They sued Midjourney correction on that a couple of months ago. But, you know, ultimately, this is what's going to happen in the long run, especially with copyright and IP that's much easier to decipher.
Jordan Wilson [00:05:56]:
Right? The written word, little hard. Right? If you see, a Marvel character or Mickey Mouse, a little easier to say, maybe, copyright has been infringed here. Alright. Our next piece of AI news, little about face, from the Trump administration, at least when it comes to their normal policy, but not at all when it comes to AI. So president Donald Trump signed an executive order aimed at blocking US states from enforcing their own AI regulations, instead calling for one central source of approval at the federal level. So The United States currently has no comprehensive federal law governing AI even as there's more than a thousand AI related bills that have been introduced across state legislatures. So, yeah, essentially, US president Trump saying, yeah, we should govern this at the federal level, but there's no, official laws or regulations at the federal level at least congressionally. So White House AI adviser David Sacks said the order will give the administration authority to push back against what it sees as the most onerous state rules while still allowing regulations related to children's safety.
Jordan Wilson [00:07:15]:
So tech companies have lobbied for nationwide rules instead of state rules because they said a patchwork of state laws could slow innovation and weaken The US's position against China as firms invest billions of dollars in AI development. And right now, I mean, state laws do vary widely. And critics say the executive order undermines state authority to protect residents with advocacy groups arguing that state level rules fill critical gaps in the absence of meaningful federal guardrails. So California governor Gavin Newsom accused president Trump of siding with tech allies, allies saying the order seeks to preempt laws designed to protect Americans from unregulated AI systems. Yeah. This one, I mean, it's really like three or four states, that you would actually need to worry about how they respond to this. One would be, well, California because that's where a lot of the big AI companies, like Anthropic, OpenAI, and Google are headquartered. You also might wanna keep an eye, on Washington where, Microsoft is headquartered as well as maybe just New York as well.
Jordan Wilson [00:08:28]:
I'm guessing there's gonna be a lot of, a lot of pushback on this. It's probably if you're looking at it just from a US versus China innovation, it's probably technically the right move just from that angle. But from the safety consumer protection, you're not gonna have a lot. Right? Essentially, companies for the most part, they're gonna write like, if if if something happens with AI, there's not gonna be a federal, congressionally passed law to hold this up to. Right? A lot of people don't talk about this or don't know. I used to be a journalist back in the day. Right? So much of what is governing social media right now from a federal legislation standpoint was things passed in the nineties, early Internet, early Internet era legislation. Right? To get actual, congressionally passed laws across in The US takes a very long time, especially anything meaningful.
Jordan Wilson [00:09:26]:
Right? If Meta had not been pushing open source large language models, for the last few years, I don't think China, would have made open source what it is today because, yes, you do have open source in some fields competing really closely with proprietary models, and maybe that did even force OpenAI's hand, at producing a very impressive GPT open source model. Google has been pushing their, Gemini three models very, very good. So, you know, even if Meta's, strategy of open source models didn't work in the long run or maybe they'll still continue it, I don't know. It has pushed the space forward, which has helped innovation in general, and it has ultimately caused consumers to win. So even if we're not gonna get any more MetaLama open source, I hope they still keep the Llama name though. Nothing against avocados. But, yeah, maybe they'll continue to do the open source and have a proprietary, we'll see. And we'll obviously continue to cover.
Jordan Wilson [00:10:30]:
And unlike previous llama models, the avocado variant, will be expected to be proprietary, meaning external developers will now will not have free access to the core technology. Yeah. So right now, if you have a powerful enough computer, you can go download, you know, llama four, and you can fine tune it and fork it and do it kinda do what you want with it, but that may not be the case anymore for Meta's future models. So the change in direction follows the kind of lackluster reception of Llama four when it was released in April and concerns inside of Meta reportedly about rivals using their open source architecture without restrictions. So CEO Mark Zuckerberg has also been spending billions of dollars in the same way like I drink ounces of water just like haphazardly. Right? Spent countless billions of dollars to recruit and, retain leading AI talent, including the acquihire of Scale AI and its CEO. Now they're Meta's AI leader, Alexander Wang, to help Meta catch up with OpenAI, Google, and Anthropic. Meta increased its 2025 capital spending forecast to as much as $72,000,000,000 on AI, highlighting the scale of its AI ambitions.
Jordan Wilson [00:11:55]:
So we'll see if this actually plays out. Is Meta going to become a closed source proprietary AI company? I think regardless, we found out that their approach maybe didn't work. Right? If you rewind back to, around February, March, of this year, before Meta released Llama four, there was a lot of people, not myself, who said, oh, yes. This next version of Llama is going to be state of the art. It is gonna be the best model. Open source is going to take over proprietary. I don't think so. Although, you know, you do have to see, their impact overall of open source.
Jordan Wilson [00:12:39]:
Right? If Meta had not been pushing open source large language models, for the last few years, I don't think China, would have made open source what it is today because, yes, you do have open source in some fields competing really closely with proprietary models, and maybe that did even force OpenAI's hand, at producing a very impressive GPT OSS open source model. Google has been pushing their, Gemma three models very, very good. So, you know, even if Meta's, strategy of open source models didn't work in the long run or maybe they'll still continue it, I don't know. It has pushed the space forward, which has which has helped innovation in general, and it has ultimately caused consumers to win. So even if we're not gonna get any more MetaLama open source, I hope they still keep the Llama name though. Nothing against avocados. But, yeah, maybe they'll continue to do the open source and have a proprietary, we'll see. And we'll obviously continue to cover.
Jordan Wilson [00:13:43]:
All right. Next one little tweet that people haven't been talking about, but if you're a ChatGPT user, like 900,000,000 of you out there, and you like new things, which is a lot of people, this tweet didn't really get a lot of notice. So, CEO Sam Altman, right after the release of its new GPT 5.2 model, which we'll be talking about here in a couple of minutes, OpenAI CEO, Sam Altman, took to Twitter and teased that OpenAI is not done. So in a tweet late Thursday afternoon, Altman said that OpenAI had a, quote, few little Christmas presents for you next week. So if you're confused what that might mean, well, it might mean that ship miss is upon us. So last year, during what they call the twelve days of ship miss, or shipping features, if you don't speak geeky AI talk, OpenAI released a lot of new features. So they released finally Sora after teasing it months earlier. They introduced and released their O one reasoning model, the ChatGPT pro tier projects upgraded, advanced voice mode and more.
Jordan Wilson [00:14:57]:
However, Google kind of grinch them. They stole the show, because Google, although OpenAI opted for some heavy marketing and daily live streams and really leaned into it, Google just quietly woke up and chose ship miss violence Because I think most close observers noted that Google probably won, the late year AI shipping season because they rolled out their, family of Gemini 2.0 models, their v, v 2.0 video upgrade, their imagine three. They upgraded NotebookLM, and they announced and released deep research, which set off the entire category. So that's led many to wonder, are we gonna get the same end of the year flurry this year, and what could be left from OpenAI? So, we do know that probably Google is gonna be releasing, you know, maybe, Gemini 3 Flash, in a smaller version of Gemini, their nano banana. But what does OpenAI have left? Well, we'll have to see, and we'll obviously be covering it in this week's newsletter. We do know that they've been testing a new image model, and some, different versions of coding models that they haven't released yet, but not a lot is known outside of that. So, could be interesting to see what they release. Speaking of releases and competition, well, Google actually updated and released their deep research and made it an agent on the same day that OpenAI released GPT 5.2.
Jordan Wilson [00:16:39]:
So Google has released a reimagined and improved version of its Gemini Deep Research agent, a major update powered by its newest foundation model, Gemini 3 pro, marking a shift from research reports to embeddable AI research tools for developers. Well, you might be saying like, why does this matter? Well, Google is opening up its deep research, which is extremely good to third party apps now through this new interactions API in the Gemini deep research agent. So what this means, if you've ever used, Gemini's Deep Research or OpenAI's Deep Research, and if you haven't, by the way, you're you're you're just burning money, right, in the in the, the nature of wasted time. If your company is not using OpenAI's Deep Research or Google Gemini's Deep Research, you you seriously, stop. We'll start using it and stop not using it. Anyways, what this means by Google releasing this as an agent, and opening this up is I think we're gonna get a lot of agentic research tools for specific domains. Right? You know, it's it's actually not a saturated market. Right? I I I thought that we would see dozens of domain specific deep research tools, but maybe with this new update, we finally will.
Jordan Wilson [00:18:01]:
So, essentially, Google's very impressive, deep research agentic tech is now open for the masses. So, you know, let's say you work in, I don't know, construction management or something, right? There's probably going to be a construction management, deep research agent that is going to be released and likely built on Google's new tech. So Google says the updated deep research agent can handle massive amounts of context and synthesize large volumes of information and is already being used for high stakes tasks like financial due diligence and drug toxicity safety research. The company, this part is important, also plans to integrate Deep Research, the new version directly into Google search, which is interesting, Google Finance, the Gemini app, so updating it there, and also NotebookLM.
Midroll [00:18:59]:
This podcast is sponsored by Google. Hey, folks. I'm Amar, product and design lead at Google DeepMind. Have you ever wanted to build an app for yourself, your friends, or finally launch that side project you've been dreaming about? Now you can bring any idea to life. No coding background required with Gemini three in Google AI Studio. It's called vibe coding and we're making it dead simple. Just describe your app and Gemini will wire up the right models for you so you can focus on your creative vision. Head to ai.studio/build to create your first app.
Jordan Wilson [00:19:33]:
Alright. This one's awesome. Alright? If I'm being honest, Photoshop might be the app that I've used for the longest. I feel I've been using Photoshop for, let me do the math here, twenty five years at least, and now they've teamed up with ChatGPT. So Adobe and OpenAI have joined forces to make Photoshop, Acrobat, and Adobe Express available within ChatGPT. So this integration means now, like, 900,000,000 weekly active ChatGPT users can now edit photos, design invitations, and even manage PDFs for free directly inside of the ChatGPT interface. So users can connect Adobe apps by going into their settings, going to the app and connectors inside ChatGPT, then selecting and linking their preferred Adobe tool. So once connected, users can simply upload their images or documents, select the Adobe app, and just describe in natural language, the edits that they want.
Jordan Wilson [00:20:39]:
Right. And this is really, really cool, because it, it also gives you a slimmed down version of the editor if you want. If I'm being honest, I love Photoshop, but maybe I need like a new computer because sometimes when I open Photoshop, it's like some of these like resource heavy desktop applications, they just crush my computer. Right? So now being able to bring in kind of the best of Photoshop into ChatGPT is going to be really impressive in what creatives can do. And this goes back to what I've been saying for multiple years about how large language models, front end large language models are AI operating systems. Right? And this just goes to that point. Right? Being able to work with the tools, the software that you use every day. Right? When ChatGPT launched their apps, I think they launched with seven.
Jordan Wilson [00:21:36]:
Now they're up to more than 30, including these three from Adobe. And I do assume, even though the, the GPT store never really took off, I do assume that this will just because of the way that it handles, and shares data, and it keeps that context active within the conversation. So not even just being able to, you know, naturally edit, with natural language, with the, you know, the the Adobe Photoshop app. But, obviously, then you can keep the conversation going and chat with a different app, and you can still have that, going. So imagine I don't know. Let's just say there's, you know, five different creative tools. In the future, you probably be able to work with all of those and keep the context going. And projects creative projects that used to take, you know, multiple days, you might be able to get done in, I don't know, a couple of minutes just by being able to direct it all with natural language.
Jordan Wilson [00:22:31]:
All right. Our next piece of AI news, some big AI rivals coming together. So, the Linux foundation has launched the Agentic AI Foundation or the AAIF bringing together nearly every major tech company, AI company, and are setting open standards for autonomous AI agents. So this means that for the first time, direct competitors like OpenAI, Anthropic, Google, and Microsoft are working side by side to create interchangeable protocols for AI agent development. So the foundation centers on three main open source projects for now. Anthropic's MCP or their model context protocol, Block's Goose framework, and OpenAI's agents MD standard. So Anthropic's MCP, probably one of the most well known ones, was introduced last November and is now the core standard for connecting AI models with external tools, data, and applications. Bloxgoosed, was released in early twenty twenty five, and it gives developers a structured MCP integrated framework for building agent workflows, combining language models and extensible tools.
Jordan Wilson [00:23:50]:
And then you have OpenAI's agents MD, which is a markdown based instruction format for coding agents, and that's been adopted by more than 60,000 open source projects and is also supported by major frameworks like GitHub Copilot and Gemini CLI. So the AAF, again, that is the Agentic AI Foundation. The member list, it's like a who's who of tech and enterprise. So aside from the companies already mentioned, OpenAI, Anthropic, Google, Microsoft, and Block, you also have Amazon Web Services, Bloomberg, CloudFlare, IBM, Oracle, Salesforce, SAP, Snowflake, Hugging Face, Uber, and a lot more. So, yeah, essentially, everyone has decided, hey. And this goes to, you know, more the geopolitical side, I believe. Right? Most of these companies based here in the US and saying, well, if you want to standardize, AI adoption in general and if the US wants to keep pace with China, you need these different companies' frameworks to be able to talk to each other. Right? And you need enterprises to be able to come in and, interchangeably be able to move around these systems and to have agents from different platforms be able to talk to each other and understand each other.
Jordan Wilson [00:25:07]:
So in in the same way, that the Linux Foundation, has set up other, kind of open standards across the web, this is another big step forward, and I think a much needed one. Right? Because there's so many now agentic frameworks. So I think it's good to have an overarching umbrella foundation in the Agentic AI Foundation, to kind of hopefully help create a little bit more cohesion, among competitors. Alright. In our last big AI news story, OpenAI has released GPT five two, and it's pretty good. So OpenAI's new GPT five two model delivers significant performance gains in writing, coding, reasoning, and more. So the launch, though, may be more noteworthy than the actual benchmarks and specs of the model. So the launch follows the internal code red, reportedly declared by CEO Sam Altman a few weeks ago, which was a company wide push to prioritize improving OpenAI's models in response to intense competition, specifically from Google and Anthropic.
Jordan Wilson [00:26:24]:
Mainly, Google Gemini's three. So essentially, Google has been gaining on OpenAI like crazy. So Google is now up to 650,000,000 monthly active users compared to OpenAI's 800 million weekly active users. I did see that number, get reported at 900,000,000, so we'll have to see, when and if OpenAI confirms that. But, essentially, over the last few quarters, Google is catching up to OpenAI in terms of just number of users. So, it seems like, according to reports at least, that CEO Sam Altman, said that they needed to focus just on model performance, that maybe they were getting a little too diluted focusing on, you know, things like shopping experiences and, potential ads and agents and browsers and all these other things. So reportedly, OpenAI had a little bit of a code red. You know, their last two models were not, overwhelmingly, positively received in GPT five and GPT five one.
Jordan Wilson [00:27:27]:
So with GPT five two, at least the sentiment so far seems good. Right? The vibes overall seem good. The benchmarks are great. So GPD five two is available in three versions. You have your instant, designed for fast information retrieval, thinking, which is optimized for coding, math, and planning, and pro, which is the most powerful tier for complex tasks. So here's one that I thought is really telling. Some benchmarks, I'm just gonna actually focus on two. They're kinda specific.
Jordan Wilson [00:27:59]:
They're a little dorky, but I think we really gotta talk about them. So, OpenAI says that GPT five two thinking outperformed human professionals in over 70% of tasks on g d on GDP valve benchmark and also completed them 11 times faster. So if you don't know what the GDP val benchmark is, it's actually pretty important. It is an open AI created benchmark that measures model performance on economically valuable real world tasks across 44 occupations. So essentially, the work that most of us do across 44 applications. And here's why I think it's important to talk about. You might be saying, okay. Well, okay.
Jordan Wilson [00:28:44]:
OpenAI's own benchmark. Well, yeah, of course, they're gonna have the best, you you know, score on it now. But OpenAI actually released GPT back in September. And guess what? When they announced it, they were not the top model. Anthropic actually was. Right? Which I think said something about OpenAI, you know, doing good research. Right? And putting something out is kinda gutsy to put out a benchmark that says this is showing how well models can do economically valuable work, and then to have one of your biggest competitors be, they actually got the best score. But now GPT five two thinking wipes away the competition on this specific benchmark, and I think that's important.
Jordan Wilson [00:29:28]:
And I will say, anecdotally, it seems it's really based on two things. GPT five one was not good at creating spreadsheets, right? A lot of errors, and they're also not good at creating PowerPoints. They just weren't very good. They weren't effective. They weren't great at storytelling. And I think that's two things in my experience so far. GPT five two is actually really, really good at. Right? That's one of the, we did a show a couple months ago on, like, hey.
Jordan Wilson [00:29:57]:
This is like the secret cheat code of Claude. It's the best, front end large language model at creating presentations and spreadsheets, which is what so many of us do. Right? That is economically valuable work. That is consultants in a nutshell. Right? This is what consultants do. Right? They they go find data. You know, they find these trend lines, connect data, and then put it in spreadsheets and put it in PowerPoints. Right? And Claude's been really good at that, but they've maybe been the only front end large language model that's truly excelled at that.
Jordan Wilson [00:30:29]:
Right. Google's good in concept, but not good at creating the actual files. And, GPT five one, not really good at it. GPT five two, really, really good. Impressive. Alright. One other quick benchmark to talk about, with GPT five two. This one, I'm only going to say the acronym once because it's pretty long.
Jordan Wilson [00:30:48]:
It is the MRCRB two, but let's just call it the four needles benchmark. So this essentially measures how well models retrieve specific needles or kind of like hidden information, kind of like a needle in a haystack. Right? So it measures how well models can find these four needles across a long context window. So if you look at GPT five one, right, so when you start at the, lowest context, so it measures it across eight, eight k tokens to 256 k tokens. So GPT five one, well, you know, 100%. Right? About a 100%, in the beginning. But as it went on, as the context window, grew at 256 k, it only got 40% right. And that's a huge drop off.
Jordan Wilson [00:31:38]:
Right? And this is something that even people that are, you know, quote, unquote good at using large language models, they don't keep context window in mind. Right? I do. I I watch it like a hawk. I have token counters. I I I'm always doing internal needle in the haystack. I created a needle in the haystack, you know, kind of internal benchmark three and a half years ago before there was an official benchmark. It's extremely important because these models, it's it's it's it's like they're like any of us on a Friday. On a Friday, your brain is fried.
Jordan Wilson [00:32:07]:
You forget things that you would normally remember on a Tuesday morning or, you know, Monday midday. Large language models are kind of the same. You know, there's a lot of things that most of us don't understand about how they work. Right? But over the longer context, they start to forget things, hallucinate, and just lack in performance. So GPT five two on this new form needles benchmark, nearly a 100% retrieval at 256 k tokens, which right. I I was trying to read and understand because this just happened over the weekend. You know, essentially, I I saw an OpenAI employee just say like, oh, they multiplied something in a different way. So I have to, like, go and, like, do more research or see if I can get someone at OpenAI to just be like, what the heck do you mean? Right? But presumably, they've figured something out internally that allows models to retain, essentially their power to, retain their knowledge and retention over a longer period, over a longer context window, which is huge.
Jordan Wilson [00:33:09]:
Right? Not just for OpenAI, but, you know, presumably, this is something that will get figured out by the other labs as well. So, yeah, it's it's it's something that a lot of people don't pay attention to. Right? If if you run the same, kind of test early on, but you extend the context window and then you run it later, it's usually gonna be a lot worse later on. So pretty big update there. Small dorky thing from GPT five two. Alright. That is all of our big news stories for the day, but let's wrap with our what's new and what's next. So some of these are rumors.
Jordan Wilson [00:33:43]:
Some of these are little updates. Some of these maybe are bigger updates that just didn't make our, you know, top eight or top 10 stories for the week. So here we go in bullet point fashion, what's new and what's next for this week. Alright. Here we go. IBM will acquire data streaming specialists, Confluent, for $11,000,000,000. Next, ChatGPT will soon support apps in custom GPTs. Not yet, but it is coming.
Jordan Wilson [00:34:10]:
OpenAI quietly started supporting their own version of Anthropic's skills feature. So some tipsters online found that out. We'll share that in today's newsletter. Google released an ultra version of notebook large model for Gemini Ultra subscribers. Just way better, way better rates. And it may be, the version of notebook LLM that gets Gemini three first because Gemini, 2.5 pro is still powering notebook LLM. Instacart and OpenAI partnered on grocery checkout in ChatGPT. Mistral released in, their Devstral two coding model.
Jordan Wilson [00:34:48]:
ChatGPT added an extended thinking feature on its pro model, which didn't exist before. So now when you go to ChatGPT pro, there's actually an extended. So you can, like, stack more thinking on the biggest thinking model. So it's like thinking x x x l. Then you have OpenAI released its state of enterprise AI report, which we shared about in our newsletter last week. Google teased GenTabs, a feature that builds apps out of your open tabs. This one looks awesome. So it's on a sign up right now, unfortunately.
Jordan Wilson [00:35:23]:
But, yeah, if you have, like, 10 tabs open, you know, researching a vacation or something like that, GenTabs, part of disco, will just create an app for you based on that. Really cool. Next, we have Thinking Machines released to Tinker in LLM fine tuning APIs. We finally know one thing Thinking Machines is working on. Claude hinted at full agent mode, aside from a chat mode, so maybe we'll see that. Infropic is partnering with Accenture to train 30,000 consultants on Claude. This comes right after Accenture had a similar announcement with OpenAI. Like I alluded to later, we might see a Gemini three flash or a nano banana two flash, around the corner from Google.
Jordan Wilson [00:36:08]:
OpenAI is testing a new image model themselves, which looks much better than their current image model. Time named the architects of AI as person of the year. OpenAI named former Slack CEO as its chief revenue officer. Google added VO powered video abilities to its Pomelli marketing tool. Pomelli is a great like, I I think it was just made for small businesses, but it's amazing, by the way. OpenAI launched their OpenAI certifications courses with select partners. But, yeah, if you're part of the general public, you can't go get OpenAI certified just yet. And last but not least, Cursor released a new visual editor.
Jordan Wilson [00:36:50]:
Alright. That is it for the AI news that matters this week. So remember, we do this almost every single Monday. We cut through all the marketing, all the hype, all the confusion, and hopefully, for our busy leaders out there, we say here's what matters, here's what it means, and here's how you can leverage it. So I hope this was helpful. And as a reminder, make sure please tune in tomorrow and Wednesday for our 2025 AI roadmap rewind. So, we did our 25 AI predictions back in January, the culmination of hundreds of hours of research from the year before. So we're going to grade ourselves.
Jordan Wilson [00:37:30]:
We're going to go over it and think of this almost as like a state of AI for 2025 and end of the year recap, but also we got to hold ourselves accountable. Did we lead you astray? Right. A lot of people come up with these crazy AI predictions and none of them come true. I think we're gonna do okay, but I mean, we're gonna ramp it up even more in 2026. So, make sure to tune in for that special two part series. All right. Thank you for listening. I hope this was helpful.
Jordan Wilson [00:38:00]:
If so, tell someone about it. Make sure to share this, online. If you're listening on the livestream, if you're on the podcast, please make sure click that follow button. If you could leave us a rating, we'd appreciate that. Then go to our website, youreverydayai.com. Sign up for the free daily newsletter. Thanks for tuning in. We'll see you back tomorrow and everyday for more everyday AI.
Jordan Wilson [00:38:20]:
Thanks y'all.
