EP 498: Meta drops Llama 4, Microsoft Copilot levels up its AI game, GPT-5 roadmap hits snag and more AI News That Matters

Resources:

Join the discussion: Got something to say? Let us know here


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Major Developments in AI: What This Means for Your Business

The latest updates in AI are not only expanding the capabilities of businesses but reshaping how organizations think about integrating artificial intelligence into their operations. Here’s a breakdown of the most significant advancements, directly from recent AI developments.


Microsoft's Copilot Enhancements: A Toolbox for Efficiency

Microsoft has rolled out a substantial update to its AI assistant, Copilot, catering to everyday business operations. Key features include memory and personalization capabilities, allowing Copilot to remember user preferences and tailor assistance accordingly. For businesses, this means greater efficiency in daily tasks as Copilot can now track user-specific details and offer personalized suggestions. Additionally, Microsoft's introduction of web-based actions enhances Copilot's capabilities, enabling it to perform tasks such as booking tickets and shopping online directly within the browser. For decision-makers, this offers a streamlined way to handle day-to-day operational tasks without the need for multiple apps or software.

MidJourney v7: Revolutionizing Creative Processes

The release of MidJourney v7 brings voice inputs and a draft mode that accelerates the image creation process. Businesses involved in creative industries will find the rapid generation of visuals particularly useful. Draft mode allows teams to create quick, low-quality drafts for review before refining into high-quality visuals, fostering more dynamic and iterative creative workflows. As companies increasingly rely on AI for creative solutions, integrating such features can enhance creativity and output speed.

OpenAI's Internal Knowledge Access: A Game-Changer for Collaboration

OpenAI's latest feature allows ChatGPT team users to access dynamic data stored in Google Drive, opening doors for improved collaboration and real-time information retrieval. This is particularly beneficial for businesses that rely on extensive documentation and internal data sharing. For example, employees can now generate summaries and tailored reports, enhancing the efficiency in handling large volumes of data stored across platforms. As this feature expands to other tools, it's poised to transform how organizations approach data handling and decision-making.

The Emergence of GPT-5 and Its Implications

OpenAI's announcement regarding the delay of GPT-5, yet introducing new 04 models, highlights a shift in the AI landscape. This delay is attributed to challenges in integrating advanced functionalities while maintaining performance. However, the introduction of 04 mini models indicates an ongoing evolution in reasoning capabilities. For businesses relying on AI for complex problem-solving, keeping an eye on these developments can be pivotal. The ability to leverage more advanced models ensures that companies remain at the forefront of innovative solutions and competitive edge.

Amazon's NovaAct: Advancing Autonomy in Business Operations

Amazon's unveiling of NovaAct introduces a toolkit designed for creating autonomous web agents capable of executing various tasks independently. This innovation underscores a significant step towards increasing automation in business processes. By enabling sophisticated task management directly through web browsers, companies can significantly enhance workflow efficiency and reduce manual intervention.

Conclusion: Navigating AI for Strategic Advantage

These advancements highlight an exciting phase for AI integration across industries. Businesses are encouraged to stay informed on AI developments to leverage technology for strategic initiatives effectively. As AI continues to evolve, organizations can capitalize on these innovations to enhance productivity, streamline operations, and maintain a competitive advantage in their respective markets. Understanding and integrating these technological enhancements will be crucial for forward-thinking business leaders seeking to harness AI's full potential.


Topics Covered in This Episode:

  1. Microsoft Copilot Update Features
  2. Midjourney v7 Image Generator Release
  3. ChatGPT Teams' Google Drive Integration
  4. OpenAI's GPT-5 Delay and New O-Series Models
  5. GPT-4.5 Passing the Turing Test
  6. Amazon's NovaAct AI Toolkit Launch
  7. Meta's Llama 4 Release and Features



Episode Keywords:

AI developments, trillion dollar companies, AI models, AI image generator, Chat GPT updates, GPT 5, Open AI, Turing test, Jordan Wilson, Everyday AI, AI news, daily livestream, free newsletter, career growth, companies, Monday updates, Inception game, NVIDIA, Google Cloud next conference, Las Vegas, META, Llama four, Microsoft Copilot, personalization, MidJourney, GPT 4, AI agents, Amazon, NovaAct, Gemini, multi-modal AI.



Podcast Transcript


Jordan Wilson [00:00:17]:
I hate saying this each Monday, but my gosh, it's been another crazy week in AI developments. I mean, think about it. We've had multiple trillion dollar companies release and update their best AI models and features. We finally have an AI image generator release that we've been waiting on for more than a year that's still a leader in the pack. We got a bunch of Chad GPT updates and news on GPT five and another model we weren't expecting from open AI and apparently an AI model has passed the Turing test. Yeah. Can't argue. It's been another crazy week in AI development, and I don't blame you if you can't keep up.

Jordan Wilson [00:01:09]:
I do this every single day and it's hard for me to keep up, but that's why on most Mondays, we bring you the AI news that matters. So what's going on y'all? My name is Jordan Wilson, and I'm the host of Everyday AI. And this is your daily livestream podcast and free daily newsletter, helping everyday people like you and me not just learn what's happening in the world of AI, but how we can all take advantage of it to grow our companies and to grow our careers. Is that what you're trying to do? Trying to make sense of all this AI? Are you trying to learn it and then leverage it in your day to day? Well, it starts here. This is where you learn what's going on, but how you leverage it, that happens on our website. So go to your everyday a I Com. Sign up for our free daily newsletter there. Every single day, we recap each day's podcast live stream, as well as keeping you up to date with everything else that you need to not just keep up, but to get ahead in AI.

Jordan Wilson [00:02:00]:
So if you haven't done that already, make sure you go do that, and you can go listen to now almost 500, episodes. Yeah. I think we're on episode four ninety eight or something like that. So I gotta cook up something special for number 500 running out of time. So hope hope y'all can join, for that. I believe that's on Wednesday. So before we get into, the AI news yeah. Because there's a lot, and like I said, we do the AI news that matters segment almost every single Monday.

Jordan Wilson [00:02:29]:
A couple things, we extended voting for the inception game. So that was our partnership with NVIDIA and their inception program highlighting some of the best AI startups, in the NVIDIA inception program. So, make sure, both in the show notes and on our website, you can go vote. There's two different ways to vote, so we shared that. Voting ends Tuesday, April 8 at 11:59PM Central Standard Time. So if you haven't already voted, make sure you go do that. One other, kind of housekeeping thing for us, We will be out in Las Vegas for the Google Cloud next twenty twenty five conference. So, in partnership with Google, looking forward to that.

Jordan Wilson [00:03:11]:
There should be a lot of updates coming out of, that show. So, hey, if you're gonna be at the, Google cloud next conference, make sure you holler at me. You know, whether it's on LinkedIn or, email, I always put that information into the, show notes as well. Alright. Enough chit chat. Like I said, so much AI news this week. Huge releases from Meta with Llama four. Microsoft, essentially said copy and paste to every single other cool AI feature that's out there that they didn't have yet.

Jordan Wilson [00:03:45]:
We have GPT five news. We have chat GPT updates. Midjourney seven is finally here. So much to get to. Let's dive in. But hey. What's up, livestream audience? Hey, Graham from Ireland. How you doing? Big bogey joining us, on, YouTube.

Jordan Wilson [00:04:01]:
Thanks. Doctor Scott on LinkedIn says congrats on number 500. Doctor Harvey Castro, great to see you. Sandra and Kyle, from, YouTube. Thank you all for tuning in. Alright. Let's get to it. Let's get to the AI news that matters.

Jordan Wilson [00:04:18]:
There's a lot y'all. First, Microsoft essentially said, oh, there's all these cool new features out there. Let's develop them all and release them all at once. Alright. So Microsoft had a celebration of their fiftieth anniversary, and they released so much. Alright. So Microsoft has rolled out a massive update to its AI assistant copilot, introducing features like memory, personalization, web based actions, and a lot more. Alright.

Jordan Wilson [00:04:50]:
So here's just some of the new updates. I couldn't even include them all because it would take an entire show. But here's I think the ones that are gonna, probably most impact everyday users. So Copilot can now remember users' preferences, interest in details to tailor advice and suggestions in their new memory features. So users retain control over what Copilot remembers or can opt out entirely. Also, there's some new personalization options. So micro Microsoft plans to offer personalize personalized appearances for Copilot, including the option to bring back Clippy. If you've missed Clippy over the last, I don't know, twenty five years, you know, the iconic assistant from earlier Windows versions is making its AI, return.

Jordan Wilson [00:05:39]:
Also, there's actions. This one's pretty big. So, Copilot can now perform tasks, directly via its web browser. Yeah. Microsoft just silently rolled out, Agentic AI, in browser. Yeah. So you don't have to download, you you know, a program. It's just working in the browser with their new actions features so you can do things such as booking tickets, reserving restaurants, and even making purchases.

Jordan Wilson [00:06:10]:
So combined with new shopping tools, Copilot can research products, find discounts, and streamline online transactions. So, yeah, big agent play, there for Microsoft as well, as well as, a huge expansion of their Microsoft Copilot vision feature, which was previously available in web tools, and it's now rolling out to Windows and mobile apps, which is a wildly useful feature. So I use that all the time on the Edge browser. It's kind of like Google has something like this in AI Studio, which it's great when it's when it works. Google AI Studio, their stream in real time was being a little finicky for me this weekend. So, you know, I might be using the Copilot vision a little bit more where you can literally just tap one button. Copilot sees everything that's on, that screen, and you can talk to it in real time. So pretty exciting there.

Jordan Wilson [00:07:03]:
Also, deep research. Yeah. Like I said, like, Microsoft literally just rolled out every single feature that they didn't have, already. So, now Copilot can analyze extensive documents in online sources for complex projects integrating with Bing for AI powered search responses. So it can generate also, you know, that whole oh my gosh. So much so much from, Copilot here. I should have teased teased y'all with this in the beginning. You know that, like, notebook l m thing that is, absolutely amazing and how you can generate a podcast, on any of your information? Well, now you can do that inside Copilot as well, with audio summaries to explain detailed topics.

Jordan Wilson [00:07:48]:
Also, new updates to its pages. So the new functionality inside, pages enables Copilot to organize notes and research across multiple documents into a single workspace simplifying project management and collaboration. And that's not even all y'all. I couldn't do, like, you know, thirty minutes of Microsoft news, but, everything's rolling out at kind of a different time. So many of these features are launching in initial versions already, with improvements expected in the coming weeks. So, availability will vary by market and platform. So we will continue in our newsletter, to keep you informed when these all come out. So much.

Jordan Wilson [00:08:32]:
Yeah. Yachell says, amazing. Kimberly says, gotta try that out. Big Bogey is loving co pilot. Joe says, perhaps a co pilot update walkthrough episode. Joe, you know what? Maybe. Alright. You know, at the end for our live stream audience, I'm gonna ask you what we should cover maybe tomorrow or later this week because there's so much.

Jordan Wilson [00:08:57]:
And I do wanna do a dedicated show on one of these new updates. So I'm gonna let you all, choose which one that is. Alright? So maybe I'll have you vote at the end. Alright. Next piece of AI news. The king is back. The king has returned to quote, one of the best nineties movies ever. So after a year plus of wait, Midjourney has released v seven of its image generator, bringing new features like voice inputs in its faster draft mode, that allows you to work more in natural language, versus more, you know, I'll say Midjourney promptees.

Jordan Wilson [00:09:39]:
So, now with MidJourney v seven, voice input is now available, letting users speak prompts directly to the model, which then converts audio descriptions into text and then generates images. Also draft mode, I think, will be pretty popular, as it offers rapid image creation producing lower quality image though in just seconds, whereas sometimes MidJourney, can take a little bit longer. So users can also refine drafts by enhancing or varying them into high quality outputs. So I think that's, what draft mode is really gonna be, most used for. Yes. It is faster than the normal full full mode, inside Midjourney b seven, but I think it's more for iterating on images and using more natural language in draft mode where the full mode, I I I think is, more if you are, really good at prompting Midjourney. Right? I'm a big midjourney fan. I always have, but, you know, I don't know.

Jordan Wilson [00:10:39]:
I think over the last year or so, it seemed like the interest, for AI image generator, at least for our audience, had gone down a little bit. So I don't know. Maybe I should ramp it back up now, especially, with the new GPT four o, ImageGen that has gone absolutely viral over the last couple of weeks. Also with Google Gemini's, new Gemini two point o Flash, that does, image generation very, very well in line multimodal. So yeah. Maybe. I don't know y'all. Do do you guys care? Podcast audience, let me know as well.

Jordan Wilson [00:11:11]:
Should we do more, AI image generation? I I think now, obviously, the quality's fantastic. The quality is fantastic, and it's really good. So, let's talk a little bit. So now there is a personalization feature that's actually mandatory, for v seven users. So before using the models, users must rate 200 image pairs to create a tailored style for generations. So older v six personalization styles can still be used, but mood boards remain unavailable for now. So there are two modes available in v seven. There's turbo mode, which doubles generation cost for high performance, while draft mode costs half as much and is much faster.

Jordan Wilson [00:11:58]:
So some features still use v six v six technology, including upscaling in painting and retexturing, though those will transition to v seven in upcoming updates. So so far user feedback is mixed actually with some praising improved realism and artistic quality, while others criticizing ongoing issues like human anatomy errors and text rendering accuracy. But many feel the update is incremental rather than groundbreaking. So I'll say the same. I think midjourney in terms of visuals in aesthetics has always been number one. Right? Even as we got the new, update from, chat g b t four o ImageGen, even as we got the imagined three model from Google, that you can use inside Google Gemini two point o Flash. You know, and obviously, like dozens of other AI image generators. Midjourney has always been the king when it comes to style, when it comes to aesthetics.

Jordan Wilson [00:13:01]:
However, where it has lacked, oh, I also gotta, mention ideogram because I think ideogram, b three that just came out is really, really good as well. But where Midjourney has always thrived is in visuals. Right? It is the most aesthetic, but it struggles in other areas. It still can't use text. Right? So if you want text incorporated at all, MidJourney is not really your thing. Also, prompt adherence, I think, in the very little testing I did, actually got a little worse in v seven. So, you know, if you do have much more complex, prompts, I do think even something like g p t four o ImageGen is a little better. So just depends on ultimately what you want.

Jordan Wilson [00:13:43]:
But, you know, as an example, if you are creating, or if your company is, you know, trying to create better, multimedia with videos and things like that, Midjourney might be best for that. Right? Because, I think it's still probably, the best starting point if you are trying to create AI video, and you are going text to video or or sorry, image to video. I still think Midjourney v seven is probably the best for most use cases. But for everything else, especially when it comes to prompt adherence, when it comes to iterating on an original image, when it comes to text, Midjourney is still not it y'all. Kimberly says over, underwhelmed. Underwhelmed. Alright. Next.

Jordan Wilson [00:14:28]:
I don't know why no one talked about this. We covered it in the newsletter and I put it out on the Twitter machine. This is actually huge. We have, like, mini rag now inside ChatGPT. And I'll explain what that is after I tell you what's new. So, OpenAI is starting to roll out its internal knowledge access for chat g p t teams users. So right now, it is only available, for teams users. And right now, the only thing available is Google Drive.

Jordan Wilson [00:15:05]:
So the new feature, and this is in your connectors settings if you are on a team plan. It it, it just started rolling out this past week, and it allows ChatGPT to retrieve real time information from internal files anywhere in your Google Drive, and it can summarize content and create tailored outputs like demo scripts or summary. So Google Drive is the first platform supported with access gradually rolling out over the next few weeks. Let me just say this. Scary good. Scary good. And you might be wondering like, oh, Jordan, when should you just use Google Gemini? It connects to Google Drive as well. It does ish.

Jordan Wilson [00:15:50]:
So if I'm being honest, this is one area that Google Gemini still struggles in. I think, you know, even though Gemini 2.5 pro might very quickly become my most used model, over GPT four o, because y'all inside let me just put this out there. Inside Google AI Studio, 2.5 pro million token context window, the world's most powerful model, and it's available for free with a million token context window. On the front end of Google Gemini chat, it doesn't have that million, token context window. So also inside Google AI Studio, you can't turn off, data sharing. So, you know, definitely don't use it with anything, you know, sensitive or proprietary. So but it struggles. It really struggles.

Jordan Wilson [00:16:43]:
Google struggles for whatever reason, accurately pulling information, from Google's, own Google Drive. Chat GPT teams does a way better job, and it is extremely impressive. So, if you do have a team's account, you need to be logged in as the team's admin. Go into, your workspace settings and look for connectors. So it takes, you you know, I don't actually don't know how long it takes. I just let it kind of sit there in the background. So it might take anywhere from, I don't know, five, ten, couple hours to fully sync everything. But then, essentially, you can click a new button that says internal knowledge in anything in your Google Drive.

Jordan Wilson [00:17:25]:
Instant access. Extremely impressive. And the reason why is because it's all dynamic. Right? So, yes, inside Claude, even inside Gemini, there's certain instances that were great when you can upload files individually, but it's not dynamic. Right? And this is why I do think this might be, the first consumer, you know, true mini RAG system. So what that means is anytime you're using a large language model, the thing you always have to keep in mind is, well, your data, recency, and just basic prompt engineering one zero one. So having this feature inside, ChatTBT teams is huge. So OpenAI does plan to expand support to other tools such as CRMs, project management systems, and data analytics platforms soon.

Jordan Wilson [00:18:14]:
But right now, it is just Google Drive. So you do have to be on a Teams plan, which is $25 per person per month. I still believe you have to have a minimum of two users, to have a Teams plan. But if I'm being honest, even if you're a solopreneur or even if you're the only one using it, it it it's probably just worth it to just pay that extra license just to use this feature alone, especially if you are a power user of chat GPT. Yeah. You gotta worry about security as big bogey face says. Yeah. Don't definitely don't just throw in, docs in there, haphazardly.

Jordan Wilson [00:18:53]:
Also, a good point, if you have that connected and if you're using it, you know, you really have to increase your personal responsibility as the expert in the loop. Right? I think I'm gonna stop saying human in the loop, FYI, because I really think it's about expertise in the loop, but you have to be much more vigilant to see what Chat GPT is using and what it's not. Alright. More Chat GPT news. This month might be some of the biggest news that snuck under the radar. So OpenAI has announced that they're kind of delaying their plans for the much anticipated GPT five, but also slipped in that, okay, we're actually gonna be releasing, two new o series models in the meantime. So OpenAI has unveiled updates to its AI roadmap, including a new o four mini model and also details about the now delayed rollout of GPT five. So OpenAI plans to release a new o four minutei model alongside the previously announced full version of the o three reasoning model within, quote, unquote, a couple of weeks, according to CEO Sam Altman's Twitter post.

Jordan Wilson [00:20:08]:
So the o four mini model is expected to serve as a next generation successor, to a reasoning model that we'd right now have o one and o three. So, yeah, I'm really interested to see what they're going to do. Are they gonna have three versions of their o thinking models available? Because for some, for some instances, I love o three mini high. That's actually been one of my workhorse models recently. But are we still gonna have o one, o three, and o four available? Because I still use and prefer o one pro, for certain instances, which you do have to be on the $200 a month, Chat GPT Pro plan. But o one Pro is the most powerful model I've ever used. I think even for certain, tasks, it's better than Google Gemini 2.5 Pro. But, I mean, we'll see how much we actually get to keep.

Jordan Wilson [00:21:04]:
So what's with this GPT five delay? Well, GPT five has been described as more of a unified model, incorporating all of the other models, you know, so advanced reasoning, voice functionality, canvas, search, deep research, tools, and everything. So, at least what we've been told is g p t five won't be a new model per se. Right? Like, g p t four, g p t four five, g p t four o. It's more going to be a system. And OpenAI has said that they will offer g p t five with tiered access, standard intelligence settings for unlimited use, higher intelligence levels for chat g p t plus subscribers, and even higher settings for chat g p t pro users. So OpenAI also, yeah, in addition to that, but also, let me mention why it's delayed. Well, at least according to, Sam Altman, he noted that the company has found it harder than expected to integrate all the features smoothly while maintaining performance, but improvements in GPT five designs have exceeded initial expectations. So kind of telling both sides of the story like, oh, it's actually going way better than we initially thought, but also at the hard time or also at the same time, we're finding it difficult more difficult than we thought to fully incorporate everything.

Jordan Wilson [00:22:23]:
So previously, you know, essentially, OpenAI said, yeah. We're not gonna release any new models, before, GPT five comes out, but change of plans here. So we're gonna be getting in o three full, and we're going to be getting in o '4 mini. Personally, I'm not looking forward to this new g b t five system, and I don't think power users should be looking forward to it either. That's just me. I don't know. It's not out yet. I would much prefer to not have a system decide what model, to use.

Jordan Wilson [00:22:58]:
If I'm being honest, I know better. Right? If you are a power user that has, you know, use every single model, thousands of prompts, you know which, model to use for which scenario. Right? I I know it like the back of my hand. I don't want necessarily a system deciding which model to send it to. I often use three or four models, in the same, project, but going back and forth in model switching. So, I mean, hopefully, g b d five is smart enough to do an adequate job. I don't have a ton of hope if I'm being honest. Alright.

Jordan Wilson [00:23:33]:
More OpenAI news. Just some bullet points here. So, Sam Altman also tweeted that OpenAI is officially developing an open weights model, so they might actually go back to being open in the OpenAI, allowing businesses to customize AI without retraining, but stopping short of full open source similar to Llama or DeepSeq, and then other, Chad GPT updates. So the very viral and extremely impressive g p t four o image, was updated. So there's a new version that rolled out. It didn't say a lot about it except it takes more time to essentially think about creating the image before it, you know, gives you the image. Also they roll the image gen out to free users, which was previously delayed. And last but not least, they are giving chat g b t plus away for free to university students.

Jordan Wilson [00:24:31]:
All right. Through May. So, essentially if you are a college student, you can get chat GPT plus normally $20 a month for free through May. You know, so we can delve into, writing our final papers together with way too many emojis. Have we passed the Turing test? Apparently, a new study, says, from UC San Diego's language and cognition lab has says that open AI's GPT 4.5 model has convincingly passed the Turing test, sparking debate about artificial intelligence's ability to mimic human intelligence and its potential societal impact. So, in the study, g p t five 4.5 was mistaken for a human in seventy three percent of cases during a three party touring test, significantly surpassing the random chance of fifty percent. So this marks a major, literally major milestone in AI's ability to simulate human like behavior. So in this study, participants engage in text based conversations.

Jordan Wilson [00:25:50]:
Alright. So this wasn't, real time. It was text based with a human and an AI. Then, the participants had to try to identify which was human and which was AI. So GPT 4.5, when adopting a specific persona, outperformed actual humans in being judged as a human. That's a lot of y'all. If you follow AI that, like, the the the touring test has kind of been this, you know, unofficial gold standard of AI development, and now we might have it. So persona prompts though were key to GPT four point five's success with instructions to act like a young person knowledgeable about Internet culture, boosting its win rate to 73%.

Jordan Wilson [00:26:39]:
Without those persona prompts, its success dropped to only 36%. So, GPT 4.5 with personas, at least according to this study, passed the Turing test, which is a huge deal. So OpenAI's GPT four o model, which powers the default version of chat g p t achieved a much lower win rate of 21%. But maybe the most shocking thing of all this, the decades old, the original Eliza chatbot that is, like, 50 years old. Right? It was technically the first chatbot. I believe it was from the what was it? The sixties. It it performed at a 23% success rate. So, actually, Eliza outperformed, GBT four o, by a couple percentage points.

Jordan Wilson [00:27:34]:
But undoubtedly, GPT 4.5 crushed the touring test. Right? A 73% win rate, it's extremely impressive. And and and we've been saying this all along. So when GPT 4.5 came out, a lot of people were confused. And they were like, okay. This thing didn't crush every single benchmark ever. So why is it important? Empathy. EQ off the charts.

Jordan Wilson [00:28:01]:
I also think this goes to show how a little bit of best practice prompt engineering goes a long way. Right? Having chat g b t with, such a simple or sorry, g b t 4.5 act as a young person knowledgeable about Internet culture, having it act under that persona, increase its win rates exponentially. So the implications of this study are significant with the study's lead author noting that AI's ability to convincingly mimic humans could lead to automation of jobs, enhance social engineering attacks, and broader social disruption. Yeah. I don't think this is necessarily, like, a great thing for AI. It's actually a little concerning. Right? Because all those, you know, scams are about to get a lot better, with GPT 4.5. I guess, luckily, in that regard, GPT 4.5 using it via the API.

Jordan Wilson [00:29:00]:
Right? So if it were to be used in a bad way, generally, you'd use it via the API because you wanna do it in mass. It's extremely expensive still. But I do think that we're gonna see a wave of new models in 2025 and 2026 like GPD 4.5, that are more tailored for emotional intelligence, versus I like, standard IQ. And that's what's really gonna trick humans, and that's where it gets, both useful in many regards. Right? Because then all of a sudden, you know, your AI powered customer support can be a little empathetic and emotionally intelligent. But at the same time, the other side of the coin can be extremely ugly. Alright. Amazon.

Jordan Wilson [00:29:45]:
Don't forget about them. They've unveiled NovaAct, a new AI toolkit for autonomous web agents. So NovaAcc is designed to create autonomous agents capable of performing tasks in web browsers. So this move signals Amazon's intensified competition in the race to commercialize AI agents and enhance their functionality beyond simple chatbots. Yeah. I think people kinda forget about Amazon even though in the same way that, OpenAI and Microsoft had this kind of relationship, right, with initially Microsoft being the biggest investor in OpenAI. Hey. Amazon is the biggest investor in Anthropic.

Jordan Wilson [00:30:30]:
So you can't sleep on Amazon, but NovaAct, their new agentic AI, is part of Nova's AI initiative, which focuses on developing foundation models for various media and input types, including text, images, and video. So the new toolkit allows developers to build AI agents that can complete step by step tasks in web browsers, such as submitting time off requests or placing recurring online orders without relying on APIs. So Amazon claims NovaACT excels in handling complex interface elements like drop down menus, date pickers, and pop up dialogues, which are challenging for other systems. So the software package available in Python enables agents to follow natural language instructions and operate in the behind the scenes mode for advanced business use. Developers can run multiple agents simultaneously to handle larger workflows, boosting efficiency for interplies, for enterprise work. So Amazon's internal testing, this hasn't been verified via third parties, have shown improved reliability compared to existing systems, but the company will monitor real world performance closely. So NovaAct positions Amazon among competitors like OpenAI, Microsoft, Google, and Anthropic in the race to develop autonomous AI systems capable of completing real world tasks. So, yeah, if you don't follow the agent space closely, I'm probably gonna do another dedicated, agent show or two, in the coming, weeks because the agent space has obviously been on fire, this week.

Jordan Wilson [00:32:06]:
But I think a lot of people are also confused, like, what the heck is an AI agent? What's different than an AI agent versus, you know, using a large language model that has tool and Internet access? Essentially, an agent, is usually powered by a certain large language model, and an agent can autonomous, autonomously make decisions on your behalf without your approval. Right? Essentially, you are giving, an agent agency. Right? That's why they call them agents. Right? You're giving it decision making powers, and it can go off and complete multiple tasks in a single sequence without human intervention connected to the Internet, connected to tools. Right? It's a very oversimplified, version of agents, but we'll probably do it in dedicated agents, show soon just because there's so much new in the space. I don't know if you guys want it. Should I do that show as well? Let me know. But also Amazon is launching a website to let developers and everyday users experiment with nova foundation models, which were announced back in December.

Jordan Wilson [00:33:11]:
Alright. Our last big piece of AI news on a Saturday, Meta unveiled llama four, its highly anticipated successor, in its open weight open source large language model lineup. So the release of llama four features new open weights models designed to push the boundaries of multimodal AI capabilities. So, yes, now, llama four multimodal by default, an open source multimodal large language model that benches actually very well. So there's four new models. Two of them are available now. So that is Llama Scout, which is the smallest model, and Llama Maverick. So yeah.

Jordan Wilson [00:34:02]:
Apparently, we went, Top Gun there. So they're available now while two more are still in training. So that is the llama four reasoning model and then llama behemoth. So they are slated for release soon. So yeah. Llama sticking with their previous, kind of llama three point, 3.2, three point three releases, having a small, medium, and large variation, with now scout, maverick, and behemoth, but then also adding that reasoning model. So the drop from Llama, the surprise drop because we had reports coming out that Llama was facing, some issues internally catching up with other open, open source models in terms of benchmarks, but I don't know. Looks pretty good to me.

Jordan Wilson [00:34:53]:
But the drop has sparked widespread excitement particularly due to a 10,000,000 token context window in the small scout model, setting a new industry standard. Alright. So yeah. 10,000,000 tokens. So, we're not sure yet how well it's going to perform. Right? In the same way, you know, oh, sorry. Google Gemini 2.5, pro with 1,000,000 context window, it is wildly, useful. But also, there's always going to be drop off with these, larger context windows because it takes longer.

Jordan Wilson [00:35:33]:
If you're using it via an API, it eats up more compute. Right? So the 10,000,000, token context window, I think has been probably the most popular piece of what was announced. But we will have to see in actuality how that plays out because, what a lot of people aren't talking about is it was trained on a 256 k context window. So, you know, I would say we really have to wait to see until benchmarks show how well this small model can take advantage of that 10,000,000 token context window. So a lot of people are, you know, shouting out right away, oh, rag is dead. Retrieval augmented generation is dead. I don't think it is. But I have been saying now for many months that I think in the future, retrieval augmented generation as we know it today will be less important than it has been in 2023, '20 '20 '4, and so far in 2025, because of these longer context windows.

Jordan Wilson [00:36:30]:
Also, I do believe that most, that most AI usage will become agentic, and reasoning as well. Reasoning models, they eat up more tokens, hybrid models as well because they're reasoning under the hood. And then when you talk about multi agentic, setups, I do think the, rag becomes a little less important, but I do think that we're gonna see a new and improved version of rag that's more applicable for hybrid, reasoning and multi agentic models. But I think, obviously, larger context windows have something to do with that, but there is an offset to that. Right? So you can't just think, oh my gosh. 10,000,000,000, you know, 10,000,000, token context window. Right? Which is, what is that? Like, more than 7,000,000 words? 7,000,500 words roughly. Right? So you're like, okay.

Jordan Wilson [00:37:22]:
I can just throw in, you know, dozens of books and, you know, like, countless hours of transcribed videos, and it's gonna remember it every single time a %. No. Remember, expertise in the loop is still important. So, Meta CEO, Mark Zuckerberg, emphasized the company's focus on open source AI, in his announcement video stating that it aims to build the world's leading AI and make it universally accessible. So he expressed confidence that open source AI will dominate the field with LAMA four marking a significant step in that direction. So LAMA four models are expected to power AI agents capable of advanced reasoning and action. So these agents will be able to surf the web and perform tasks useful for both consumers and businesses, potentially revolutionizing productivity tools. So Meta plans to host its first LlamaCon AI conference later this month on April 29, showcasing AI advancements, of Llama four.

Jordan Wilson [00:38:29]:
Alright. So we need to talk about the benchmarks. There's a lot of rumors swirling around. You know, people are doubting, LAMA's own internal benchmarks. I won't say that, because here's why. Humans have confirmed it. Third party, benchmarking services have confirmed it as well. So as an example, the third party benchmarking service, which is a great resource artificial analysis, looked at non reasoning models.

Jordan Wilson [00:38:58]:
Okay. So non reasoning, so you you know, none of the OpenAI o three, o one Pro, Google Gemini two point five Pro, etcetera. So among non reasoning models, llama four maverick is in third place and pretty close behind, GPT four o and deep seek v three, which were both just updated a couple of weeks ago. So, actually, if not for those updates from GPT four o and DeepSeek v three that just happened, llama four maverick probably would have been number one, non reasoning model on third party benchmarks. Right? So, yeah, a lot of people, if you read read the Hoopla online, right, because, originally, there were reports saying that Meta was facing delays. They weren't able to get the benchmark that they wanted. But, I mean, here you go for an open source, model that is safe to use. That's the other thing.

Jordan Wilson [00:39:49]:
You know, if you wanna use, a model from China via the API or the web, I would highly advise against it. You know, it's different if you're downloading it, fine tuning it yourself, for safety reasons or, using a tune that or or sorry, using a version of DeepSeek or other Chinese models that have already been scrubbed, by a company like, you know, Perplexity or, Microsoft Azure, etcetera, then it's obviously safe to use. But, you know, llama for Maverick, pretty impressive on the third party benches. Also, if we look at the Elo scores from the LM Arena, so this is human preference. So instantly, right away, llama four maverick, which again is that medium model, in testing is now the second most preferred model in the world. And I do think a lot of times, you know, benchmarks are important. Right? But I also think maybe just as important is human preferences. Right? Because models can essentially be overfit, to perform well on certain, benchmarks, but, you you know, it humans might not find the same utility, that you might expect based on just benchmarks alone because of that overfitting problem.

Jordan Wilson [00:41:06]:
Right? So I think Elo scores like the LM arena where you put in a prompt, you get two responses. You don't know which is which and you vote for it. Right? After millions of votes, you start to get some clear winners in terms of which models are best for humans, and that's what matters. And llama four maverick, very impressively, leapfrogs a bunch of very capable proprietary, models. Right? That's the other thing. So, yes, Llama is not true open source. It's not like an MIT, license or something like that via deep seek. It's a little different.

Jordan Wilson [00:41:43]:
They have some restrictions. Llama does, so it's under a, more of an open weights, open source ask, Llama license, but still, an open source model immediately goes to the second best model in the world. Right? With a fourteen seventeen score, on the, Elo, the LM Arena. So, Gemini 2.5 pro, so good by the way, 1439. Llama4Maverick, 14 17. And then the updated GPT four o, 1410. Alright. So extremely impressive.

Jordan Wilson [00:42:23]:
You you know, instant reactions to the new, models from Llama. Alright. That's a lot. What do you guys want? What do you guys want? Alright. I don't know if I can do it tomorrow. Maybe I can, but let me know what you wanna hear more of. There was a lot. There was a lot there.

Jordan Wilson [00:42:43]:
So live stream audience, let me know. I'll probably put something in the newsletter as well. So let's do a quick recap. In livestream audience, let me know what you care most about, what we should cover next. So here is a very quick recap of all the AI news that matters for the week of April 7. So like we said, Microsoft unveiled just about everything new, inside Copilot rolling out handfuls of new powerful AI features. Midjourney finally released its v seven after more than a year of waiting for its AI image generation model. OpenAI sneakily is rolling out what I think is mini rag with its internal knowledge access for chat g p t plus users connecting to, dynamic data in, Google Drive.

Jordan Wilson [00:43:34]:
OpenAI also announced it's kind of delayed plans for g p t five. Bummer. But on the good side, did say that they're gonna be rolling out the full version of o three and a new o four mini thinking models in the coming weeks. Next, OpenAI's GPG 4.5 in a study has passed the touring test rather convincingly. Amazon has unveiled NovaAct, its new autonomous web agent. And then last but not least, Meta unveiled, multiple llama four models. Two are already out. Two should be released soon.

Jordan Wilson [00:44:14]:
So a lot to cover, this week. Let me know what you wanna see more of. Also, please don't forget, if you're going to be at Google next twenty twenty five in Las Vegas, let me know. I think I should actually have time to go and, you you know, talk to a lot of, you know, different providers that are there, at this, Google Next conference. Maybe attend a session or two. Right? So I'm excited about this conference, in partnership with Google. And then don't forget Inception Games. We're going to have yeah.

Jordan Wilson [00:44:47]:
The the madness of March, might be just about over. I believe the championship game, is tonight. But our AI startup madness continues. We need your vote. It's actually very close. You know, we're gonna have our final two back on, for a show, and, you know, we'll announce the prize and some of those other things, live on that championship show of the inception games. So if you have not voted already, make sure you go back and listen to episode four ninety seven, where we had our awesome eight group of AI startups in the inception games pitch their service to you all. So if you haven't voted yet, make sure to go do that.

Jordan Wilson [00:45:27]:
Alright. That was a lot. I appreciate y'all. I would also appreciate you going to youreverydayai.com, signing up for our free daily newsletter. Thank you for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI