EP 494: Gemini 2.5 Pro Unlocked: Inside the world’s most powerful AI model

Resources:

Join the discussion: Got something to say? Let us know here


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Unveiling Gemini 2.5 Pro: The AI Model Business Leaders Aren’t Talking About

In the rapidly evolving world of AI, significant advancements can sometimes go unnoticed amidst a plethora of innovations. The Gemini 2.5 Pro model from Google is a case in point—a powerful AI tool that, despite its capabilities, hasn't received the attention it arguably deserves.


Breaking Down the Gemini 2.5 Pro’s Modern Approach

The Gemini 2.5 Pro is a technical hybrid model, utilizing a built-in "thinking" mechanism known as chain of thought reasoning. This model assesses the complexity of tasks and decides the computational resources required, bringing a nuanced and intelligent approach to AI processes. For business leaders, understanding the depth of such capabilities can be key to leveraging AI for intricate problem-solving and strategic planning.

Context Handling Redefined

One of Gemini 2.5 Pro’s standout features is its 1,000,000-token context window, which equates to around 750,000 words. This ability to handle extensive streams of data without losing track makes it an ideal tool for companies managing large document pools or needing comprehensive analytical capabilities. For decision-makers, this feature translates to maintaining a holistic view of complex projects, enhancing both operational efficiency and strategic foresight.

Advancements in Coding Capabilities

Gemini 2.5 Pro also excels in coding, scoring highly in benchmarks for complex code generation. This means businesses can rely on it not just for coding efficiency but also for innovation, as the model can create sophisticated applications from scratch or enhance existing tools. For tech-driven enterprises, it opens up opportunities to experiment with new solutions and streamline software development processes.

Human Preference and Accessibility

Remarkably, Gemini 2.5 Pro leads in human preference scores, suggesting it can create outputs that resonate more naturally with users. Moreover, it's available for free to users via Google's Gemini app, democratizing access to advanced AI capabilities and inviting diverse organizations to integrate it into their technology stack.

A New Era of Multimodal Understanding

With its default multimodal capability, Gemini 2.5 Pro understands and processes text, images, audio, video, code inputs, and more, seamlessly integrating them in analysis and communication tasks. This ability to cross-link different modes of input can be especially beneficial for businesses in media, marketing, or industries where comprehensive content integration is key to success.

The Business Implications of Gemini 2.5 Pro’s Innovative Model

The deployment of Gemini 2.5 Pro within organizations allows for enhanced reasoning, a wider context understanding, and improved task execution. Its capabilities can elevate the approach to project management, customer service, and product development. This model represents a step towards more effective use of AI in high-level strategic roles, suggesting that ignoring such advancements might mean missing out on vital business tool opportunities.

Conclusion: Navigating the Future with Gemini 2.5 Pro

While other AI advancements may initially overshadow Gemini 2.5 Pro, the potential it holds for redefining business operations is immense. Decision-makers who recognize and harness its capabilities may find themselves at the forefront of innovation, utilizing a tool that not only meets current technological standards but also anticipates future needs. By incorporating Gemini 2.5 Pro into their strategic toolkit, businesses stand to gain a significant competitive edge in the evolving digital landscape.


Topics Covered in This Episode:

    • Gemini 2.5 Pro AI Model Update
    • Hybrid Thinking Models with Chain of Thought
    • Massive 1,000,000 Token Context Window
    • Gemini 2.5 Pro's Advanced Coding Abilities
    • Human Preference and Benchmark Scores
    • Free Access to Google Gemini 2.5 Pro
    • Multimodal Capabilities and Use Cases
    • Future Updates and Enterprise Integration


Episode Keywords:

Gemini 2.5, Gemini 2.5 Pro, Powerful AI model, Google Gemini, Large Language Model, Chain of Thought, Reasoning model, AI capabilities, AI context window, 1,000,000 tokens, Benchmark scores, Human preference, Multimodal AI, AI Studio, Advanced coding, Coding benchmarks, Context handling, Thinking model, Transformational AI, OpenAI competition, Elo score, AI tools, API, Vertex AI, Google AI ecosystem, Multimodal by default, Creative writing, Coding software, Image analysis, AI integration, MCP support, Model context protocol, Model updates, Google's AI strategy, AI future outlook.


Podcast Transcript


Jordan Wilson [00:00:16]:
In the two plus years that I've been doing the Everyday AI Show, I don't know if there's ever been an instance where an AI update this big, especially a large language model update, has been talked about so little. I think that there's a reason for it, but we're gonna talk about it today because I think the new Gemini 2.5 pro model from Google is probably the best single large language model I've ever used. And I don't think I'm alone in that because it has not just broken just about every single benchmark, but in terms of human preference, it is quite literally off the charts. So, today, we're gonna be going over Gemini Gemini two point five pro unlocked inside the world's most powerful AI model. Alright. I'm excited for today's conversation. I hope you are too. What's going on y'all? My name is Jordan Wilson, and welcome to Everyday AI.

Jordan Wilson [00:01:20]:
This is your daily livestream podcast and free daily newsletter helping us all not just keep up with AI, but how we can use all these advancements to get ahead to grow our companies and our careers. If that's what you're trying to do, welcome. This is where you learn on the podcast or the livestream, but this is only half the battle. You need to leverage what we talk about today and where you do that is our website. So if you haven't already, please go to youreverydayai.com. Sign up for the free daily newsletter. Each day in our newsletter, we recap each day's, you know, podcast or live stream as well as keeping you up to date with literally everything everything else in the world of AI. So it is your one stop shop, to stay ahead just like this podcast.

Jordan Wilson [00:01:58]:
I always like to remind people this is unedited, unscripted, trying to bring you all something real in the world of artificial intelligence. Alright. So I am excited to get into, today's topic and talk about Gemini 2.5 pro, by far the most powerful AI model I've used. But before we do, let's first start off as we do some days well, most days with going over the, bullet points of the AI news. Alright. So first, Runway has unveiled gen four, their newest AI powered video generator capable of creating consistent characters, locations, and scenes with realistic motion and physics. So the new model allows users to generate videos using reference images and textual descriptions, offering superior prompt adherence compared to previous models in style consistency without additional training. So backed by investors like Google and NVIDIA, Runway does now face some legal challenges over copyright concerns while aiming for 300,000,000 in annual recurring revenue and a $4,000,000,000 valuation.

Jordan Wilson [00:03:04]:
So a study warns that AI tools like gen four from runway could disrupt more than a hundred thousand US entertainment jobs by 2026, raising concerns about the future of the film and TV industry. Yeah. That just goes to show how good, these new models, are. So, yeah, runway gen four, I think, is probably in Sora territory, maybe a little bit better. I mean, we'll see. It is just dropped. So I'm sure the reviews are gonna be coming out. But, I I think, you know, Google VO might have some competition.

Jordan Wilson [00:03:37]:
And, hey, in terms of availability, Runway Gen four is available to everyone like OpenAI's Sora is, whereas Google's, v o two tool is not available to everyone at least inside of their platform. You can access it from third party, platforms, though. Alright. Our next piece of AI news, another record breaker. OpenAI has officially secured a record breaking 40,000,000,000, with a b, 40 billion dollar funding round valuing the company at $300,000,000,000. So OpenAI has closed that historic $40,000,000,000 funding round, making it the largest private tech investment ever. Alright. So it does, value chat the chat g p t creator at $300,000,000,000, and the round was led by Japan's SoftBank contributing 30,000,000,000 of that amount with additional investments from Microsoft, interesting there, Thrive Capital and others.

Jordan Wilson [00:04:32]:
The funding does come with a condition at least from SoftBank. Their investment could drop from 30,000,000,000 to only 20,000,000,000 if OpenAI does not fully transition into a for profit entity by the end of twenty twenty five. And that would require approval from both the California attorney general and Microsoft in a resolution of these ongoing legal challenges from Elon Musk, which I think are pretty much theater. Alright. And this also comes as OpenAI did just announce that their weekly active users has jumped up to 500,000,000, and, OpenAI CEO, Sam Altman, did just say on Twitter that they added literally a million people in an hour probably with all the Ghibli, AI studio, you you know, video or, photo generations. Also, this comes on the heels of OpenAI just announcing that they would release an open model. So pretty exciting news there. So make sure to follow along.

Jordan Wilson [00:05:30]:
We'll be following that news. Alright. Last but not least, some big news from Amazon. They've unveiled NovaACT, a new AI agent to compete in the agentic race. So, NovaAct from Amazon is an AI agent capable of independently navigating web browsers to perform basic tasks such as filling out forms, making reservations, or ordering food. So the Nova Act SDK is a toolkit for developers, and it's available now as a research preview on nova.amazon.com, allowing developers to, prototype agentic applications. So NovaACT is developed by Amazon's AGI Lab, co led by former OpenAI researchers. So according to Amazon, NovaACT outperformed OpenAI's and Anthropics agents in internal tests.

Jordan Wilson [00:06:20]:
But despite those claims, Amazon has not yet benchmarked NovaAct on some more widely recognized, evaluations like WebVoyager. Also, reportedly, NovaAct will play a critical role in Amazon's upcoming Alexa Plus upgrade, a generative AI enhanced version of Alexa, potentially giving Amazon a competitive edge through its massive user base. Alright. So we're gonna have a lot more on those stories and everything else you need to get ahead on our website at youreverydayai.com. So make sure you go there and sign up for the free daily newsletter if you haven't already. Alright. Enough chitchat. Let's get into Gemini 2.5 pro.

Jordan Wilson [00:07:01]:
No one's talking about it. It's wild. Like, the fact that we have a large language model that's available now, with the capabilities that Gemini 2.5 pro has and hardly no one's using it, no one's talking about it is pretty, telling, right, of a couple of things. I think this is a case of shiny AI syndrome as I like to call it. Right? I think that what Google has released can change fundamentally how we all do business, yet so few people are using it just because there's a new shiny AI object in the room, which is the new, four o ImageGen from OpenAI. Really a groundbreaking visual model. And, yes, we are gonna be doing a show on that sometime soon. That one's gonna require a lot of research.

Jordan Wilson [00:07:50]:
And and speaking of that, even today's show, we're actually gonna break it up into two chunks. So today, we're just gonna be talking about, kind of the bullet points, high level, what's new, and then we're gonna be doing a show maybe later this week or next week, kind of a part two. So let me know, what more you want to see, what you want us to test from Gemini 2.5. So, you know, our second part is gonna be more, hands on in use cases where today, we're really just going over, the bullet points of what's new. So, make sure, to let me know, livestream audience, what you wanna hear, from, part two use cases. You wanna see, all that good stuff. Speaking, hey. Good to see you.

Jordan Wilson [00:08:31]:
You know, our our YouTube, family here, our LinkedIn. Thanks for tuning in. Michelle, Samuel, Jose, Shires, Kyle, Sandra, Jean, big bogey, Christopher, Brian. I can't get to you all. Thanks for joining. But do let me know what questions do you have on Gemini two point five, but let's start here. What the heck is new in Gemini two point five? Well, there's a lot, and it's also a little confusing because, you know, you might be having deja vu. You know, you're you might be saying, okay.

Jordan Wilson [00:09:01]:
Wait. There's new Google Gemini updates, you know, out of nowhere. Didn't this just happen? Yes. It did. So we're gonna be doing a quick recap also of what was released, like, literally two weeks ago. But first, let's talk high level of what's new in Gemini 2.5, and then we're gonna be going over, all of these kind of piece by piece. So, some of the biggest things is now it is a technical hybrid model, although Google did not choose to call it a hot hybrid model. But what that means is it has built in thinking.

Jordan Wilson [00:09:33]:
So Gemini 2.5 pro is a thinking model. It uses chain of thoughts reasoning kind of under the hood. You know, we've been talking about this a lot over the last few months and we'll continue to talk about it a lot in 2025. Kind of this is the new direction that large language models are going. So Google following suit here, with Gemini 2.5. So, think of it this way, you kind of have your quote, unquote old school, you you know, transformer models, and then you have your reasoners that essentially use more compute to kind of do this chain of thought thinking or chain of thought reasoning under the hood. So Google Gemini 2.5 kind of combines both. So if you have simpler task, at least in my testing, it still kind of goes through that, that reasoning or those thinking steps although it's pretty quick.

Jordan Wilson [00:10:19]:
So it kind of depends or or or sorry, Google Gemini 2.5 kinda decides how much, compute or how much thinking it needs to use. But that's probably one of the biggest, you know, what's new. The other is the context window. Enormous. 1,000,000 token context window, which is, roughly like, 750,000 words, or, 1,500 pages. So we are talking literally multiple, books. I mean, we're talking, 30,000 lines of code as an example. So if you are brand new and you're like, what the heck is the context window? Right? That's essentially how much a large language model can remember at any one given time.

Jordan Wilson [00:11:00]:
This is different than memory. Right? But essentially, you know, think if you're having a chat with a large language model and you're giving it some information and you're going back and forth, right, with older models. Right? So, you know, let's even talk chat g p t. They're kind of a little behind in terms of context window. Right? They've had a, you know, roughly 32,000, context window on their front end chat products. So that means, hey. After, you know, 26,000 words, ChatGPT is gonna start forgetting. So, with this, at least in AI Studio, I did not see Google, clarify anything on the front end if you're using this inside of, Google Gemini on the front end chatbot.

Jordan Wilson [00:11:40]:
We will be testing that though. We'll probably share it in the newsletter. But, hey, essentially, a 1,000,000 context window, 1,000,000 tokens is wild. That means that the chat is pretty much not gonna forget, right, until you use it, like, incessantly. Right? Like like it like, until you are going wild and you're not leaving that chat, you're dumping thousands and thousands or sorry. I should say hundreds and hundreds of pages. It's still gonna remember, which is huge. Another thing, advanced coding.

Jordan Wilson [00:12:08]:
Some of the top benchmarks score for, you know, SWE bench as an example and complex code generation. So if you are big into software development, if you're big into coding, or even vibe coding. Right? This this whole concept of, hey. I'm just gonna open a large language model, have it code something for me, you know, have it code a Chrome extension for me, have it code a little desktop application, you know, have it code a simple CRM. Right? This was one of my bold, you know, twenty twenty five AI predictions is everyday people like you and me would just be using AI, to code our own little pieces of software. Gemini is great for that. Right? And the big like, the good thing is you don't have to know anything. You don't even have to tell it what coding language to use.

Jordan Wilson [00:12:51]:
Just be like, yo, Gemini, I want a a Chrome extension that does this, build it for me, and then give me simple step by step instructions on how I go ahead and install and deploy it. So fantastic for advanced coding. Already, I would say it's not the top coding model in the world. I still think Claude Sonnet three seven, inches it out a little bit. There's a lot of different coding benchmarks, but, you know, essentially, Claude from Anthropic was so far ahead. It's like it wasn't even close. Right? It was like they were one a, one b, one c. They were probably even number two.

Jordan Wilson [00:13:25]:
Right? And everyone else is so far in the distance. Now, Google has closed that gap and they're essentially, one b, with Gemini 2.5. Benchmarks, human preference is huge. So, you know, I've talked about this a little bit. I think a lot of the, kind of the AI labs, especially in 2024, were kind of overfitting models. So what that means is when they were building them, going through post training, all that, is they were doing it to get certain scores on benchmarks. Right? So Google Gemini 2.5 does that. Right? Like, not saying they overfit it to get certain benchmarks, but it cleans up on benchmarks and, you know, essentially has top scores, you know, either number one or number two on every important and telling benchmark that there is.

Jordan Wilson [00:14:10]:
However, the big one is the Elo, kind of the Elo score. So in the LM arena, this is essentially I talk about this a lot of the show. Think of it as a a blind taste test Pepsi versus Coke. You put in a prompt, you get two, outputs, and you choose which one is better. Those outputs are not names. Right? And that kinda gives you an Elo score. Generally, when a new model comes out. Right? So a GROC three or a GPT four o latest or, you you know, or quad three seven.

Jordan Wilson [00:14:40]:
Right? Usually, the new state of the art model, will go into first place generally, on the Elo scores, but maybe by only by, like, two points. Generally, it's usually like a two to four point jump anytime a new state of the art model comes, and it's like, oh, it's the most powerful model in terms of what humans prefer because that's extremely important. Right? In this case, Google Gemini came out by a 39 margin, which is literally unheard of, has not happened. So, yes, it checks the box in terms of benchmarks, but it definitely checks the box in terms of human preferences, which I think is usually more important. Right? And the LM arena has, I believe multiple millions of votes, right, not millions of vote yet, for Gemini two point five. But already with enough qualifying votes, it is the top in humans preferred by far. The other big thing, it's free. Right? Google snuck this in actually over the weekend, so they announced Gemini two point five last week.

Jordan Wilson [00:15:40]:
A couple days later, they're like, oh, guess what? We're gonna make it available for free. So if you do have a Google Gemini, account, so you can just go to gemini.google.com. You know, you can use your your Gmail or Google Works based credentials, and you'll find Gemini 2.5 in there, and you can start using it for free right now. Alright. So that is the high level. Alright. And, hey, livestream audience. Let me know what your thoughts are, you you know, of Gemini 2.5.

Jordan Wilson [00:16:10]:
Sandra's asking you can use it to code a widget for you. Yeah. You can use it to code anything, Sandra. But, yeah, you do have to, you know, as an example, a a Chrome extension or, you know, something that runs on your desktop, you have to still execute that, but it will write the code for you and tell you how to, you know, install or execute it. So let's go over because you might be thinking, didn't this just happen? I'm confused. Wasn't there just new Gemini two point something updates? Yes. There were. Okay.

Jordan Wilson [00:16:40]:
So, about two weeks ago, mid March, if you go back and listen, to episode four eighty two, if you want the full updates, we gave it to you there. I love Google's new strategy here. Right? I I I think they had their original, you know, kind of, December 2023 snafu where, you know, they put out this fancy marketing video about their AI, and it turns out a lot of it wasn't true and it didn't work, and they kinda got dropped through the mud. And they spend the better part of of 2023 and 2024 way behind. Ever since. I love what Google's doing. They're not coming out with flashy advertising, flashy marketing, big announcements, big hype. They just come ship.

Jordan Wilson [00:17:21]:
Right? They just ship updates that are pretty amazing. So, they did two weeks ago announce some pretty impressive updates that I still don't think people talked about. So, if you wanna know about that, you can go listen to that in episode four eighty two. Again, for free on our website. Yeah. If you didn't know on our website, you can go and listen to every single episode we've ever done, interviewing some of the world's top experts on AI. But here's essentially what was announced in the mid March version so we can get this out of the way. So, Gemini two point o multimodal, which was huge.

Jordan Wilson [00:17:52]:
I think that kinda set the stage for this whole, you know, g p t four o image gen, multimodal, by default. Amazing. I went over that. You can literally create a blog post with inline images. Yeah. Wild. Right. You can also edit images with natural language kinda like what you can do now with g p t four o's ImageGen.

Jordan Wilson [00:18:13]:
So mid March, Google announced the Gemini two point o multimodal. They announced deeper research was updated to the two point o model whereas previously, it was running on 1.5. They announced personalized Gemini, which I think some people like, some people don't like. Right? But it essentially takes into account your search history. So it's a mode that you can select. So that was new. They also announced Gemma three, which is wildly powerful for a super small, open source model so you can run it locally. They announced Gemini robotics running on Gemini two point o, as well as big updates to my favorite AI tool, NotebookLM.

Jordan Wilson [00:18:52]:
Also, they upgraded that as well under the hood to the Gemini two point o in a, model versus previously it was running on Gemini 1.5. Alright. So if you're scratching your head and being like, wait. Is Jordan, like, a month late on this? No. Gemini just did or sorry. Google did just have a ton of big updates a couple of months ago. Alright. Let's get into it now.

Jordan Wilson [00:19:14]:
Let's go over kind of, point by point here. Again, this one's not gonna be a super long one because we are gonna have a point two, or sorry, a part two. But here's kind of what's new in the Gemini 1.5, pro launch. Are you still running in circles trying to figure out how to actually grow your business with AI? Maybe your company has been tinkering with large language models for a year or more, but can't really get traction to find ROI on GenAI. Hey. This is Jordan Wilson, host of this very podcast. Companies like Adobe, Microsoft, and NVIDIA have partnered with us because they trust our expertise in educating the masses around generative AI to get ahead. And some of the most innovative companies in the country hire us to help with their AI strategy and to train hundreds of their employees on how to use Gen AI.

Jordan Wilson [00:20:10]:
So whether you're looking for chat g p t training for thousands or just need help building your front end AI strategy, you can partner with us too, just like some of the biggest companies in the world do. Go to youreverydayai.com/partner to get in contact with our team, or you can just click on the partner section of our website. We'll help you stop running in those AI circles and help get your team ahead and build a straight path to ROI on GenAI. So launched late March, by Google and Google DeepMind as their most intelligent AI. The biggest thing, like we talked about, it focuses on built in thinking or using that chain of thought. This is huge. This is huge for reasoning, coding, context handling. I mean, it's there's a lot of new capabilities and it does change what's possible for businesses.

Jordan Wilson [00:21:05]:
Alright. So how can you access Gemini 2.5 pro? Well, like I said, over the weekend, Google just with a tweet just said, oh, by the way, we're making this for free. So, it is available for free, to Gemini app users. So if you're using this inside the kind of Gemini chat, so gemini.google.com, Right? Especially if you're on a paid account, you do, have the option to turn off model training. So, you know, you don't have to worry about the the data that you share being used to train Google's model. So, on the front end, you can access Google Gemini that way. You can also access it for free in Google AI Studio, which is a kind of a more experimental version and more of a sandbox. And I'm glad that Google has shifted, their strategy after I, like, I don't know.

Jordan Wilson [00:21:53]:
I feel I, like, did so many rants in 2023 and 2024 because Google, for, like, a year kind of, quote, unquote, hid their most powerful and capable models inside Google AI Studio, which is more for developers. And then they didn't even label or tell you what was powering their Gemini chatbot. So you have no clue, but usually it was running a model that was up to six months old. So not anymore. I love Google's new strategy here. Put the newest, the latest, the greatest model inside the front end Google Gemini chatbot, but you can still use, Gemini, 2.5 inside Google's AI Studio, and that is where you're going to be able to get that full 1,000,000, context. Just keep in mind in Google Studio, Google's AI Studio, it is free, but there's, no data protection on that end. So, yeah, don't go and put, you know, confidential proprietary company data inside Google AI Studio.

Jordan Wilson [00:22:45]:
It is more of a sandbox. Also, the enterprise path, it will be coming soon, to Google Cloud Vertex AI in the coming weeks. Right? So there's technically I know it's a little confusing. Right? There's so many different ways that you can access Google and Google Gemini, you know, as well as inside their apps. Right? So they didn't say yet, you know, as an example, if their Gmail, Gemini, integration has been upgraded to 2.5. I'm not sure, but, at least for right now, you can go access it. Gemini.google.com even for free. If you have a paid account, you have higher limits, as well as you can, access it for free inside Google's AI studio, and it will be coming soon, kind of across Google's, family of products, via the, Google Cloud Vertex AI.

Jordan Wilson [00:23:34]:
Alright. Let's talk a little bit about the reasoning. So it does have that built in chain of thought like we talked about. So what the heck does that mean? Well, it kind of plans steps internally before it gives you an answer. And the cool thing is, right, you can click that show thinking. I don't know why. I feel a lot of people don't read that. At least people that I talk to, I highly encourage you.

Jordan Wilson [00:23:58]:
If you wanna get better outputs out of any large language model that shows it's kind of chain of thought, you should be reading that. Right? Because you'll see what happens a lot of times, especially with these, kind of hybrid models that can reason, they take a little bit longer, which is okay. Right? Because outputs in general are exponentially better, more accurate, more robust, more complex, much better. However, it does take a little longer. So what I always do while I'm quote, unquote waiting, right, it might be ten seconds. It might be two minutes depending on, how complex of a query you're giving the model. Read the chain of thought. Always read it.

Jordan Wilson [00:24:39]:
Right? If you want to, you know, be be future proof in your job. Right? If you wanna be the smartest person in AI in in in your department, read the chain of thought and go ahead and accordingly make updates to how you use that model, how you use the prompt. Right? All the thinking models work a little bit differently. Right? So you have Claude, three seven SONNET. It's a hybrid model with thinking. Although if I'm being honest, I think that was more of a marketing thing because you still have to click the extended thinking. Anyways, you know, you also, have goo sorry, OpenAI's, all their models there, o one zero one pro, o three mini, o three mini high, o three mini pro should be coming out via the API soon. Right? So, always, no matter what thinking model you're working with or reasoning model or hybrid model, look at the chain of thought, see what's going right, see what's going wrong.

Jordan Wilson [00:25:29]:
I always tell people, have a conversation, reprompt to get better results. Also, you gotta talk hey. One one, maybe we'll do a dedicated show on this. I don't know. Livestream audience. Let me know if you wanna know more about this humanities last exam benchmark. It is a newer benchmark put together by, I think they said, hundreds of subject matter experts. Essentially, it's a benchmark that in theory shouldn't be in any training data yet.

Jordan Wilson [00:25:56]:
So it did get the new Gemini 2.5, got an 18.8% on the score, which you might think like, oh, 18% of a hundred, AI is dumb. Alright. Humans, I doubt any human out there listening. Any single human could get a 1% on this humanity's last exam. Let's be honest. Right? But the previous high score, was, OpenAI's GPT 4.5, which got a 14%. Anthropic, Claude's three seven got an 8.9%. I believe DeepSeek was, shortly there behind in the mid eight percents.

Jordan Wilson [00:26:29]:
So, yeah, Gemini, two point five and eighteen point eight percent. So, hey, in terms of it being able to, solve and tackle very complex problems that the single smartest human in the world could never solve. Right? You'd have to get hundreds of people working together to be able to make a dent in this humanities last exam. You you know, Gemini did a great job. Also, it excels at complex logic and math without needing external tools. So we did talk a little bit about some of these benchmarks, but like I said, right away out of the gate in the LM arena, which is just human preference. You you know, being 39 points above the last, or the next best model, super impressive. Some other top math and science scores on the AIM 2025, got an 86%, and the GPQA diamond science, it got an 84%.

Jordan Wilson [00:27:25]:
Very impressive. And then the, an 81% on the MMMU, which is the multimodal equivalent of the old standard of AI testing, which is the MMLU. Alright. And y'all, even though this even though Gemini 2.5 has only been out for a couple of days, they've already made multiple updates to it since. Alright. So since launch, so a couple things. Number one, they made it open and available for free users. So that was not available at the time of launch.

Jordan Wilson [00:27:59]:
It was only available to paid users inside, Gemini.Google.com. So on the front end chatbot. So now it's available to all free users. Also, just hours ago, y'all, this is this is why sometimes I don't sleep and why I don't always, you know, sometimes I'll do prerecorded shows. But literally just hours ago, CEO, Sundar Pichai, just did re kind of announced that the new canvas mode, is available in 2.5 pro. So I did use it actually while planning this show, going through my notes and, you you know, having it put together, kind of some interactive elements to help me better learn and understand, what was new. So there are some things in the canvas mode that worked very, very well. There were some things that were buggy if I'm being honest.

Jordan Wilson [00:28:44]:
Right? It it is experimental. Keep that in mind. Also, they added support for third party tools like Cursor AI. That's huge. So, we we should see because, Claude and, you know, Anthropic's Claude has been making a living, right, essentially by being the coding, LLM of choice by the top software developers, by the top, engineers. So we'll see. I do see that potentially changing, especially when you look at API costs. The the, cloud models are rather expensive and the Gemini models aren't.

Jordan Wilson [00:29:22]:
So we'll see what happens. And if, anthropic still is kind of the the the defacto, model chosen for, software developers. Also, Sundar Pichai hinted at future MCP support for Google Gemini. Pretty big news. So that's model context protocol. I know we're gonna do a a dedicated MCP show soon, but, essentially, you you know, if you've been seeing this little acronym floating around, you're like, what the heck is it? Right? So, essentially, you have APIs. Right? So, in the SaaS world, in the software world, right, APIs is essentially a language that softwares can talk to each other. Right.

Jordan Wilson [00:30:00]:
It APIs can sometimes kind of work for AI tools and large language models, but not necessarily. So they're a little different. So this model context protocol, was actually developed, by Anthropic, but it's being used now and supported by just about everyone. OpenAI last week announced, support for it and, you know, so Google and Google Gemini may support, you know, MCP as well, which is essentially I like to think of it. It's a little more complex than this. Think of it as the API for large language models. It allows, different AI systems and different large language models to talk to each other and to talk to other APIs into other softwares. Alright.

Jordan Wilson [00:30:39]:
Next, the coding abilities. Alright. So we already talked a little bit about this is one of the, more, unique or, at least one of the angles that Google is taking with Gemini really pushing and promoting, its proficiency, in coding. So, very impressive. So, you know, they they put some demos out there, but also in agentic coding, you know, it scored a 63.8% on that sweet bench, which, you know, when it comes to agentic coding, I I do think that is the benchmark to look at. You you know, go go play with it. Right? And the good thing is is now you have that canvas mode inside Google Gemini 2.5. So you can literally go code anything you can think of with natural language.

Jordan Wilson [00:31:22]:
Code me this, build me this. Right? And you can render it or run it in the new canvas mode. So it's a little different, than OpenAI's canvas mode, which I think is more of like a Google Docs esque collaborative environment. You can run certain, coding languages inside OpenAI's version of Canvas or ChatGPT's version of Canvas. But I'd say in my limited testing of Canvas so far, which came out, a couple of weeks ago, I'd say it is more like, it is more like in tropics artifacts feature in terms of it can render and run a lot more languages. Right? And go go have fun. Right? This whole vibe coding thing, right, it's it's it's been kind of this this this trending topic. Go vibe code yourself something.

Jordan Wilson [00:32:07]:
See if you can. Right? And then if you can get it to run inside Canvas, then that means, okay. It's working and you could go, you you know, deploy it somewhere else whether you need to, you know, have it running on a a a full stack, kind of app, you know, running on some service online or whether you would run it on your desktop, whether you might, you know, as as a Chrome extension, etcetera. Right? I one shot it, which was pretty fun. And I shared it in the newsletter yesterday. I don't know if anyone saw it. I I I did a little, you know, simple, Chicago inspired, game. Right? A little, you know, side runner, two d, you know, very, you know, early Nintendo ask type game, but just one shot.

Jordan Wilson [00:32:49]:
I said, hey. Do it like this. You know? Working in all these Chicago elements, you know, hot dogs and pizza and potholes. Right? You know, make it kind of, you know, bring in elements that I like from Super Mario. Right? I need one shot, and it worked. Very amazing. Right? So, it like I said, instantly from a coding software development, you know, we're we're gonna put it through some more testing and maybe we'll do that in part two if that's something you wanna see, but very proficient in coding. Next, we have to talk about the multimodality and the context window.

Jordan Wilson [00:33:26]:
So Gemini in their two point o versions, everything is multimodal by default. So what that means is it understands not just text, but it understands images, audio, video, code inputs, or a a mixture of all of those things together, which is pretty amazing. So we we we are getting that as well and getting close to that, from big models like, Anthropic Claude and, OpenAI's chat gbt, but not quite there yet, specifically with video. Right? That's kind of a different modality, for at least OpenAI's chat gbt. Claude, I don't think is really gonna play too heavily in the multimodal by default space, although I feel they should. Right? I I I think Claude was really hoping they could carve out their niche, just with software, just with just with coding. But, you know, Google's like, hey. Hold hold my, hold my MCP.

Jordan Wilson [00:34:22]:
I mean, the 1,000,000, token context window, amazing. Like I said, we're gonna be putting that to the text test, inside of the, Google Gemini front end. I do know and have done some testing on the back end of AI studio. Context window, super impressive. Also, Google did announce that they're planning, soon for a 2,000,000 token context window. I mean, that's wild. Right? So one thing I'm gonna probably do is get together transcripts. Right? Like, I have almost 500 episodes, of the Everyday AI Show.

Jordan Wilson [00:34:57]:
That's thousands of pages of transcripts. So that's probably something I'll do for a test, upload everything. But, yeah, like, as we get to multiple million token context windows, I should have put this in my, you know, AI twenty twenty five, you know, road map series. So, you know, if you haven't listened to that, make sure you go on our website and go listen to free, for free to those is a five part series. I don't know. I think the, you know, rag is going to become a little less important, in 2025 and 2026. I'm not saying it's not gonna be needed. It's still gonna be needed.

Jordan Wilson [00:35:32]:
Right? But I think, so many, especially smaller companies and small use cases, you know, they heard this rag terminology, you know, really in late twenty twenty three and 2024, and everyone's like, oh, I need to build, you know, retrieve a log minute generation. Right? But, okay, what if you don't have a ton of data? Right? What if you don't actually have a ton of files and it's not a lot? Right? You might just be able to work in that 2,000,000 token context window. So, you you know, the context window is actually extremely, important to the future of AI development. Alright. Let's talk a little bit about some of the early feedback. So like I said, we've shared about this in our newsletter, but there's been some very impressive, you know, one shot generations, people building video games, precise image analysis, three d simulations, extremely impressive, and we'll be doing, some of those, in our part two of this series. Audio skills, being able to instantly get accurate transcriptions, very impressive, and just positive. Right? Just positive, you you know, people are always like vibes.

Jordan Wilson [00:36:33]:
Right? The vibes on 2.5, Gemini two point five, pretty pros positive, so far. Let's look at the market impact. So Google right now, they are trying to be the leader in thinking models. Right? They're kinda beating OpenAI to the hybrid punch. Like I said, technically, INTROPIC was first with Claude three point seven SONNET, but I don't know. I I I actually talked to a couple of people of this at the, a couple of people about this, at the NVIDIA conference at GTC in the few, you know, two minutes, free time I had between, like, the 15 interviews I did out there. We still do have, like, one or two more shows dropping from, GTC, by the way. A lot of people were like, yeah.

Jordan Wilson [00:37:17]:
And I'm like, hey. What do you think about this this new hybrid approach from Claude? And they're like, oh, is it really hybrid? Right? You technically have to click if you want the extended thinking or not. But with this, I think Google is in the driver's seat at least right now when it comes to this new hybrid model approach, which we also heard from OpenAI is going to be their approach moving forward as well. So they've said when we get GPT five, it's going to be more of a system. Right? And you're not necessarily going to be able to choose which model, that you use, which some people might like. Right? If you look at OpenAI's chat g b t and, you know, in my pro account, I think I have nine different models to choose from. Some people might be intimidated by that. So, you know, at least the GPT five, is going to be more of a in in architecture, that's going to kind of use this mixture of models or mixture of experts or using, kind of trend, traditional, you know, quote unquote old school transformer models and hybrid or, you you know, these these reasoning and thinking models.

Jordan Wilson [00:38:18]:
But Google with this, they're they're the leader in it right now. I don't think Anthropic did a good job, with it. If I'm being honest, I don't I think a lot of people were not super impressed, with SONNET three seven. I know a lot of people, you know, defaulted back to SONNET three five. They weren't very impressed, and they didn't feel they had kind of enough control, right, at least on the front end users. That's what we're talking about, not on the back end. But I mean, this this play right here. So aside from being a leader in terms of the thinking model, also, I mean, with the enterprise game.

Jordan Wilson [00:38:50]:
Right? So, this isn't released for Vertex AI yet, which is probably a good idea. Right? Because it is buggy. I should probably say this. Right? I will say when OpenAI releases a model, at least this is in my personal experience, I'm using, you know, the main models every single day, multiple hours a day. The the OpenAI's models when they're released, yes, they're they're throttled. They may go down. Right? Gemini has better availability. Right? You know, a lot of times, if especially if you're on a free plan or, you know, the basic $20 a month plan with ChatGVT and a new model comes out, you know, it might be very slow or availability might be impacted.

Jordan Wilson [00:39:28]:
But when it is there, it works fairly well, I will say. So I will say that Gemini 2.5, although it's not you know, there's no slowdowns, there's no real outages, the availability is there. It has been a little buggy. Right? So the canvas mode, although it's only been out out for a couple of hours, it's been a hit or hit or miss for me, but when it hits, it hits. The same thing with the just the general, you you know, Gemini 2.5. It's been a little buggy, but it's experimental. Right? You you know, I always run a series of tests and sometimes we were getting, not I wouldn't say hallucinations, but some misdirections. Right? One thing I always do to test its Internet, capabilities, I say, hey.

Jordan Wilson [00:40:07]:
What's the what's the latest episode of the Everyday AI podcast by Jordan Wilson? You know, so I see, okay, is it actually able to navigate to the web and and find the latest episode? And instead, it gave me the weather. Right? The weather was accurate, but that's not what I asked for. So, you know, hit or miss so far, but I think once Google, irons out some of those things, it's going to, you know, be be a very, impressive and reliable model. But I think that's honestly why they haven't really released it for Vertex AI yet. Right? So that's, when you can, you you know, when you'll start seeing it deployed at scale, you you you know, across, many large enterprise organizations. But I do think that Google is taking more of a a tiered approach and making sure individual users, people kind of using their sandbox in AI studio have a good experience. They're gonna wanna squash some of those bugs before they release it to the masses. Alright.

Jordan Wilson [00:41:02]:
And then kind of last but not least and hey. Live stream audience, thanks for sticking with me. We are gonna have a part two. If you have any questions, get them in now. I'm gonna scroll through, see if I can answer any. But last but not least, we have to look at the future outlook and updates. So, we are going to see, some pricing updates probably soon for the API, because I do believe there's gonna be some heavy usage. Google has said that they're working on enhancing the reasoning and coding even further.

Jordan Wilson [00:41:31]:
So there will be some under the kinda under the hood updates. That's another important thing to to to to think about. Right? So even though we we saw this jump from, you know, Gemini two point o to Gemini 2.5, that doesn't mean that Gemini 2.5 won't be updated until we get something like Gemini three. Right. Yeah. You have to kinda keep up with, you know, sources such as everyday AI, right, to see when some of these more under, under the hood model updates come out. But I do see it, you know, I think they're gonna squash some of these bugs, make some improvements. But the biggest thing is the ecosystem.

Jordan Wilson [00:42:04]:
Right? I'll I'll I'll be, interested to see when and if Google announces, if Gemini 2.5, Pro or Gemini 2.5 is going to be rolled out with deeper integration into its ecosystem. So that means right? And I expect it better and deeper integration across, you know, Google Sheets, Google Drive, Gmail, Docs, etcetera. I would love to see if, we're gonna get 2.5 in notebook l m, in Google Gems, which is kind of their version of, you know, GPTs. Right? Creating kind of this personalized version of Google Gemini. Also, you gotta get ready for the clap back now y'all. That I think I'm gonna end on aside from your questions because here's what happens. Anytime a model like this comes out and it is met with fanfare, and by fanfare, I mean, a combination of traditional benchmarks. Google Gemini's got it.

Jordan Wilson [00:43:02]:
Human preference, they got it in the Elo. And then just overall vibes, like I said, overall, people are loving the Gemini two point five, but no one's talking about it. No one's talking about it because everyone's on, you know, OpenAI's new, four o ImageGen creating Ghibli studio pictures of their family. Right? And don't get me wrong. That is actually I'm more impressed by, you you know, if I had to compare the two even though there are two unrelated things, I'm more impressed, with the updates, from, OpenAI actually because it is really driving the multimodal conversation. And, you know, and even though we did get this multimodal, kind of by default with Gemini two point o a couple of weeks ago, being able to create and work with, images in line, being able to edit with them, I think the execution was just much better with the four o image update from OpenAI. But, from a pure large language model, standpoint, Gemini 2.5 just getting completely overlooked, extremely powerful. So we're going to be, we're gonna be testing this in a part two.

Jordan Wilson [00:44:12]:
So make sure and if you aren't listening on the podcast, thank you. You can always reach out to me or just respond. When you sign up for the newsletter, tell me what you wanna see in our part two. How do you wanna see us us put, Gemini two point five to the test? What use cases, demos do you wanna see us run? Big bogey face here asking, let's test its coding skills. We can definitely do that. Kabari, asking on YouTube in terms of capabilities on a scale of one to 100, where is AI now? Oh, that's a good question. I don't know. If we're talking about Gemini two point five, I mean, you have to say it's in the nineties.

Jordan Wilson [00:44:49]:
Right? And this ahead of everyone else. If you're asking in terms of the AI, you you know, as a whole, I don't I don't know. Right? Because that 100 or the ceiling is constantly being raised. Right? Again, if you would have told people two years ago, that we would have models this capable, this powerful, available for free, I think you'd say no. Like, oh, that's not possible, but it is. Here we are. So the ceiling keeps getting raised. Denny asking, what about content creation for writing needs? Most of what was mentioned are video or tech kinds of needs.

Jordan Wilson [00:45:22]:
So I will I will say this, Denny. Great question. And maybe that's something, we can test as a use case, just kind of creative writing. But I do think that Gemini has always had a nice knack, for creative writing. Right? I think ultimately with proper prompt engineering, OpenAI has always been best. But if you're talking about zero shotting and trying to get some good, kind of creative writing, I think people have always preferred Claude. I think Gemini is is right there in terms of, you know, what you can get out of the box with just, you know, hey. Here's five examples.

Jordan Wilson [00:45:55]:
Go mimic, go mimic this. I think Google Gemini is actually great for that, and that's something I did do some, some testing on a little bit. Jose is asking when did the live stream start? Yeah. They start 07:30 Chicago time. So, yeah, if you're on the podcast, if you didn't know, this is unedited, unscripted. You can come in here, hang out, network, ask questions. We try to tackle everything as best, as we can. Alright, y'all.

Jordan Wilson [00:46:19]:
So I hope this was helpful. Got a couple questions, couple comments in here at the end. Again, we're gonna have a part two where we're gonna be breaking all of this down, go over use cases, do some things live, really push it to its limits. So make sure you join us for that, and let me know what you wanna see, what you wanna hear. So thank you so much for tuning in. If you haven't already, go sign up for that free daily newsletter at youreverydayai.com. We're gonna be recapping the highlights, and what you need to know from today's livestream podcast. If you didn't catch everything, don't worry.

Jordan Wilson [00:46:49]:
It's gonna be in there as well as everything you everything else you need to get ahead to grow your company and career with generative AI. So, if this was helpful, please subscribe to the podcast. Please leave us a rating. I'd appreciate that. I'd also appreciate it always makes me smile a little bit. You know, if if you, are listening on LinkedIn, click that repost button if this was helpful. We spend so many hours, cutting through the BS, bringing you on, you know, hopefully, unbiased and just real information, to help you make better decisions on your AI strategy and implementation. So if you could repost this, if it was helpful, I'd appreciate that.

Jordan Wilson [00:47:24]:
So thank you for tuning in. I hope to see you back tomorrow and everyday for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI