Ep 530: Google I/O AI Updates: 15 new features and how they can grow your business (Pt 1 of 2)

Resources:

Join the discussion: Got something to say? Let us know here


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Exploring Google's AI Innovations: What Business Leaders Need to Know

Google's recent announcements at their annual IO conference have positioned them as a frontrunner in the AI arena, showcasing a multitude of updates that business leaders should consider integrating into their strategic plans. The company's aggressive push into the AI space unveils specific tools and technologies that can potentially enhance business operations. Here's a detailed look at some of the most significant AI advancements from Google and their practical applications.

Revolutionizing Image Generation with Imagen 4

The latest iteration, Imagen 4, has set a new standard for AI-generated visuals, moving away from past shortcomings. This new platform offers enhanced photorealism and better text rendering within images, critical for businesses needing sharp, customized visual content. The advent of Imagen 4 allows organizations to bypass traditional stock photography, offering a fresh, authentic touch to their digital presence—be it on websites or marketing campaigns. Additionally, its integration into platforms like Google Docs and Slides means there's a seamless transition for those already embedded in Google's ecosystem.

Upgraded Browser: Chrome with Gemini Integration

Google has transformed its Chrome browsing experience by integrating Gemini, a step towards smarter web interactions. While not groundbreaking, this update streamlines the way users can summarize and interact with web content. By facilitating autonomous web navigation, it provides businesses with the tools to perform complex online tasks more quickly and efficiently. Companies can anticipate this tool augmenting research, customer service, and competitive analysis tasks.

Personalization in Email—A Critical Step Forward

Google's new email personalization feature is set to overhaul communication efficiencies. It leverages one's writing style, past emails, and relevant Google Drive files to generate context-rich responses. Particularly for businesses, this feature will streamline communication processes, ensuring that customer interactions are both personalized and swift. This advance is slated to launch through Google Labs, promising to enhance user interaction with timely, relevant email responses.

Enhancements in NotebookLM

NotebookLM, Google’s cutting-edge AI tool, is now powered by Gemini 2.5, offering a hybrid thinking model. With new multimedia capabilities, including the creation of simple video overviews, users can expect an enriched experience in working with complex datasets. This positions NotebookLM as an invaluable asset for data-driven industries, offering a visual and auditory lens through which to interpret data.

Experimenting with Speed: The Gemini Diffusion Model

One of Google’s more avant-garde innovations is the Gemini diffusion model, which diverges from the traditional transformer models by leveraging diffusion techniques. This approach is fine-tuned for coding and math, offering faster and more accurate responses in these domains. Such efficiency can greatly benefit tech-driven sectors, particularly those that rely heavily on rapid, precise computational processes.

Breaking Language Barriers in Real Time

Real-time translation within Google Meet is set to redefine global business communication. Initially supporting Spanish and English, this tool offers seamless natural voice synthesis as if a live interpreter were present. This capability eliminates language hurdles, enabling a broader spectrum of business dealings internationally.

Mobile AI: Gemini App Updates and Gemma 3n

Google's Gemini app updates enhance user interaction with improved research capabilities and app performance. Meanwhile, Gemma 3n stands out as a powerful small language model, designed for on-device AI applications. Given its ability to function without cloud reliance, this model prioritizes speed and security, critical for businesses focusing on local data processing.

Google’s foray into AI at the IO Conference signifies not just a technological leap but an array of opportunities for business leaders prepared to embrace and implement these tools. With potential applications spanning from enhanced communication to advanced analytics, the potential to drive business growth and innovation is substantial. Remember to explore and utilize these state-of-the-art tools to adapt and stay ahead in the rapidly evolving digital landscape.


Topics Covered in This Episode:

  1. Google’s Position in AI Race
  2. Top 15 AI Updates from Google I/O Conference
  3. Imagen 4 Text-to-Photo Platform
  4. Chrome with Gemini Integration
  5. Personalization in Email
  6. NotebookLM Updates
  7. Gemini Diffusion Model
  8. Real Time Translation in Google Meet
  9. Gemini App Updates
  10. Gemma 3N Model


Keywords:

Google IO, AI updates, generative AI, large language model, AI race, Microsoft, OpenAI, Anthropic, Google leader, AI conference, top 15 AI features, business leaders, free daily newsletter, Google AI, Google Chrome, Gemini integration, Imagine 4, text-to-photo platform, Google workspace apps, personalization in email, NotebookLM updates, Gemini diffusion model, real-time translation, Google Meet, Gemini app updates, multilingual support, hybrid thinking model, AI-powered tools, personalized email, Google Drive, Chrome features, business use cases, international communication, Edge AI, small language model, Gemma 3n, multimodal capabilities, on-device AI, open-source model, pro and Ultra subscriptions, data security.


Podcast Transcript


Jordan Wilson [00:00:16]:
Google has come a long way in a very short period, which seems weird saying that about one of the biggest companies in the world. But when it comes to the AI race, let's be honest about fifteen months ago, I don't even think Google was in the top three, right? When you look at Microsoft and open AI and Anthropic, I think about fifteen months ago, Google was actually in fourth place. But now without a doubt, Google is the absolute leader in the generative AI in large language model landscape. And what they just announced at their IO conference is nutty. And I think if nothing else, it really just cements Google's place, at least right now, as the leader of the pack. We'll see how and when everyone else responds. But at least for right now, Google is just cooking when it comes to AI, and they released dozens of notable AI updates. And on today's show well, on today and tomorrow's show, we're gonna be breaking down what I think are the top 15 most useful.

Jordan Wilson [00:01:31]:
So, yeah, we're gonna have a part one, which is today, and a part two, which is tomorrow. But we're gonna be going over the top 15 most useful AI updates out of the Google IO conference for everyday business leaders such as yourself. Alright. So I'm excited to dive in. I hope you are too. If you're new here, what's going on, y'all? My name is Jordan Wilson. I'm the host of Everyday AI, and this thing, it's for you. This is your daily livestream podcast and free daily newsletter helping us not just keep up with AI, which is very hard, but how we can actually use it to grow our careers, grow our company.

Jordan Wilson [00:02:07]:
So is that you? Did that hit home? If so, well, you're in the right place. This is your home. It starts here on the unedited, unscripted livestream of podcasts. This is where you learn, but where you're actually gonna leverage this and put this to use is on our website at youreverydayai.com. Because once you're there, you can sign up for our free daily newsletter. We're gonna be recapping today's show, but we also keep you up to date with everything else happening in the world of AI. And, yeah, even though Google is sweeping the headlines, there's still a lot more happening. And then also on our website, you can go and listen to for free sorted by category, more than 500 past episodes.

Jordan Wilson [00:02:45]:
Whatever you're trying to learn, we've already spoken to the experts. It's all there already. Alright. So, normally, we start out each livestream with the daily news, but let's be honest. Google is the AI news today. Alright. So, I'm excited for today's show. What's up, livestream fam? It's good to see you.

Jordan Wilson [00:03:05]:
Yeah. If you listen on the podcast, maybe sometime drop by at 07:30AM Central Standard Time. You know, when we have guests on, what other place can you go and ask questions live to the smartest people in the world on AI? Today, it's just me. Sorry. But, what's up live stream fam? So, Christian, join in on YouTube. Good to see you. Brian and Michelle, doctor Harvey Castro, big bogey. Everyone else.

Jordan Wilson [00:03:29]:
Denny, good to see everyone. Let's just not tease you anymore. Here's at least the first half of our top 15, AI updates from the Google IO conference for everyday business leaders such as yourself. Here we go with 15 through eight. Number 15, imagine four. 14, Chrome with Gemini integration. 13, personalization in email. 12, notebook l m updates.

Jordan Wilson [00:03:56]:
I can't believe that didn't make the top 10. 11, Gemini diffusion, a whole new type of larger language model. 10, real time translation in Google Meet. Nine, Gemini app updates, and eight, Gemma three. And, Yeah. That's a lot. And y'all I didn't miss anything. We still have our, you know, numbers seven through one, but here's some that didn't even make the list.

Jordan Wilson [00:04:22]:
Alright? And if you've been following, the the the AI news over the past, you know, I don't know, twelve to twenty hours. These are big advancements that didn't even make our top 15 list. Alright. Gemini code assist, Synth ID detector, Lyria two, the virtual try on in shopping, Google Beam, which is enormous news even of itself formerly called Project Starline, Jules, the new autonomous coding agent, the a to a, agent to agent enhancement. So, yeah, when I say that there were dozens, I literally had to scratch my head and look at my list of, like, 50 and say, what are the top 15? Right? So, very hard to do. Very hard to do. Alright. So, there's probably some big things you're like, wait.

Jordan Wilson [00:05:09]:
Where are some of these big, ones? Well, those are tomorrow. Right? You'll notice I didn't even say the word Gemini two point five there. A lot of updates there or Veo 3. Yes. V o three, which is shocking. Alright. So we're gonna be going over those, and a lot more tomorrow. Alright.

Jordan Wilson [00:05:27]:
But let's stick on our top 15 for today. Hopefully, a, concise show or more concise, show than normal for for you all instead of doing a an hour and a half, show or something like that. We'll try to keep this one short. Alright. First, Imagen 4. So this is Google's updated, text to photo platform in Imagen 4. It's really good. For our livestream, audience, you can probably see if you're listening to the podcast, nothing overly visual or overly instructive today, but, you know, maybe you wanna check out what's on the screen.

Jordan Wilson [00:06:06]:
You can always do that by checking out your show notes, and, you know, go on our website and watch the video. But, look at this image. This looks beyond real. Right? So this is a a a young girl here, a young woman, looks like in a in a dorm room maybe with pink hair and earrings and, you know, kind of a grungy t shirt with light filtering in, you know, through the window. It looks like an amazing photo that was captured with a high end DSLR. This does not look AI generated in the least bit. Let's just start there. It is as somewhat, and, you you know, I don't talk too much, about my my background here.

Jordan Wilson [00:06:49]:
I just realized y'all. I just realized I don't even have my mic plugged in. This is, this is how much, this is how much work I was doing and maybe how sleep deprived I am. So, live stream audience. Give me, give me a second. Let me know. Let me know if you can hear me now. Hopefully you can.

Jordan Wilson [00:07:10]:
Can I get a thumbs up, from, from the live stream audience? I didn't have my mic plugged in, but, must have been picking up somewhere else on my computer. Alright. So, thanks to my computer for still delivering some type of audio, even though my mic wasn't plugged in. Alright. Hopefully hopefully y'all can hear me. All right. Let's keep it going. So, this is good.

Jordan Wilson [00:07:37]:
So imagine for let's talk a little bit about what's new. Thank you. Thank you, Marie and Laura for letting me know you can hear me. Appreciate that. Okay. So here's what's new in imagine for and what it is if you haven't heard of it. So maybe you've heard of, Midjourney. You know, there's the new kind of viral GPG four o ImageGen inside OpenAI.

Jordan Wilson [00:08:01]:
You You know, there's a lot of these AI photo generators, you know, stable diffusion, Flox. There's, you know, a good five to 10 pretty good ones. I'm gonna be interested to see where imagine four lands on the benchmarks. So in the same way we talk about the LM arena, which is kind of blind taste test for large language models, they have that for image and video models as well. So, I'll be interested to see where Imagen 4 lands on the list. But from early, I test and as someone, I was a photographer, You know, kind of before in my earlier life, I've probably taken, more than a million. Yes. More than a million photos with a DSLR.

Jordan Wilson [00:08:41]:
So I would say my eye is a little more trained than the average eye when it comes to looking at things like photo realism or even being able to decipher what's real and, what's real and what's not. And I will tell you, imagine four images are otherworldly good. You know, in the same way, you know, mid journey v seven, very good, but, jeez, these Imagen 4 photos. So good. So good. Alright. A little bit about what Imagen 4 is, what's new, when it's rolling out, all that good stuff. So this is Google's latest and most capable image generation model with improved detail and text rendering within images.

Jordan Wilson [00:09:20]:
That's a big thing. Midjourney can't render text and they kind of said, yeah, we don't really care about that. This is good. The ability to render text. Yes. GPT four o ImageGen does great at rendering text for whatever reason you may want. Right? So maybe you want this person wearing a shirt to have a a t shirt that says, you know, the name of, you know, University of Illinois or something like that or Chicago. Right? Some AI image generators struggle with that.

Jordan Wilson [00:09:47]:
Imagen 4 so far does a really good job like GPT four o ImageGen does. But in terms of photorealism, quality, Imagen 4 is pretty good. And by pretty good, it might be the best out there. Time will tell. So right now, it's rolling out now in the Gemini app. Also, this is pretty interesting. It's coming to all of Google's, different products. So Google Google Docs, Slides, and other workspace apps.

Jordan Wilson [00:10:14]:
So, yeah, I don't really use Google Slides, but now I'm like, okay. There might be some use cases where, you know, I might want to or maybe might need to in some instances. Right? So this is going to be in the new, included in the Google AI Pro and Ultra subscriptions. We're gonna be talking a little bit more about that tomorrow. But for these things to make sense, you have to know, previously, you know, Google had a couple of tiers. Right? There is a free tier, and then there was a Gemini advanced. And in typical Google fashion, they're confusing us all. So now they're still obviously a free tier.

Jordan Wilson [00:10:47]:
The new $20 a month plan is called, Gemini a or sorry, Google AI Pro. I'm already getting confused. Google AI Pro is the base $20 a month plan. And now you have the Ultra, which is ultra expensive at $250 a month. Technically, $2.49 99. And I think for the first three months, it's like half off. But, you know, the base plan is gonna be $250. So, this is already rolling out to those people who have either of those two subscriptions.

Jordan Wilson [00:11:18]:
So like I said, some of the standout features here significantly better text rendering and images, enhanced photorealism, improved handling of complex prompts, so prompt adherence. There's in painting and out painting capabilities. So if you wanna change something inside the photo, you can do that very easily. If you want to extend a photo, right, so whether it's a photo you start with or a photo that you create inside a match in four, you can out paint or extend it, to bring in more of the scene that was actually never captured originally. And also this supports a range of aspect ratios up to a two k resolution. So there is a coming soon for this, a faster version. I don't know if they're gonna call it turbo, but, apparently, it's going to get 10 times faster fairly soon. So what the heck could you use this for to grow your business? Well, first, get rid of those ugly stock photos on your website.

Jordan Wilson [00:12:11]:
They look horrible. Right. Also, starting from here, if you're creating any videos, for social media, anything like that, start with an imagine for image. Right? Yes. Start with images. If you're doing AI video, it turns out better. But there's no shortage of ways that companies can just use visuals. Chances are everything you're using, whether it's for internal or external purposes is either extremely old, extremely boring, or a combination of both.

Jordan Wilson [00:12:41]:
Alright. Number 14, Chrome with Gemini integration. Alright. So, what this is while the Chrome browser is finally gonna get a little smarter. Alright. So I can't pretend that this is some groundbreaking new feature. It's more of like, oh, in about time, because let's call a spade a spade here, card players. Microsoft and their Edge browser, which is actually freaking fantastic.

Jordan Wilson [00:13:10]:
It's based on Chromium. Right? So all your Chrome extensions, everything like that will sync over. Microsoft Edge has had this for, like, a year. Not all the capabilities, but they've had a built in Copilot, for, like, more than a year, and that's why I use Edge a ton. But about time, we're gonna be getting Chrome with Gemini integration more than just being able to summarize web pages and things like that. But it helps you also with web browser tasks. So this is also you're gonna have to be on a paid plan, and you can summarize web pages that can help you explain complex information, answer questions about page context, content. And eventually, here's the eventually and why maybe it's for paid subscribers and not available for everyone for free.

Jordan Wilson [00:13:55]:
Eventually, it will be able to help you navigate websites autonomously, which is pretty big. That's been a big shift over the last even month or two. A lot of companies, the the the DIA browser from the browsing company, perplexity coming out with a Comet browser, even Microsoft Edge with their vision feature, you you know, built in, you can see web pages. So the ability for browsers by default to perform tasks is not some future, you know, sci fi. This is it's already available. It's, but it's been, like, wildly popular the past, like, three months. So, this Chrome with Gemini integration, will be, eventually be able to do that, at least Google says. So, what are some business use cases for this? Well, pretty straightforward.

Jordan Wilson [00:14:43]:
Number one, it's gonna help you summarize web content faster. Right? Which if you haven't already just been doing that, in Microsoft Edge. I told you about it, like, I don't know, a year and a half ago, and I'm like, start doing this. So nothing super new there, but, obviously, the ability for Chrome to perform actions on your behalf without having to launch a separate agent, pretty big in terms of time savings, winning back time, all that good stuff. McDonald said this is very impressive. Alright? I'm talking about Imagen 4. So he says art director for twenty years and things like Imagen 4, very impressive. Yeah.

Jordan Wilson [00:15:21]:
I I agree. Like I said, I've been taken more than a million photos with the DSLR, getting paid to do so, and it's really, really good. Alright. That's number 14. Let's go to number 13, personalization in email. So this is an actual, not what's on my screen, but this personalization in email was one of the things that actually, Google CEO Sundar Pichai actually talked about during his keynote, which I found interesting. Because when there's literally dozens of, of updates that are huge, personalization and email at first, I'm like, okay. This this is no big deal.

Jordan Wilson [00:16:01]:
But when you look at some of the the the marketing materials, again, there's obviously a huge gap between what's being marketed, what's being promised, and what actually happens. Right? And Google is getting much better. Although their original track record on this a year and a half ago, two years ago, no bueno. Now they're just shipping. Right? So I actually do have a high degree of com, confidence. A lot of these things are gonna be shipped on time. But the personalization in email, something against Sundar Pichai mentioned in his keynote address with his limited time on stage. So for our livestream audience, you kind of see an example here.

Jordan Wilson [00:16:37]:
You know, so there's kind of this blue area that's shaded, a green area that's shaded, and then a yellow area that's shaded. And it's showing you how, Google and Gemini are going to be able to use personalization, based on your context. Right? So it's not just those auto replies, right, which had been in Google Gemini for a long time, and I don't really use them because I don't think they're good. This, when and if it gets released, will be actually really good. So as an example, you know, the things in blue, it's basing part of an email reply based on, your own writing style. So it goes and it sees how you respond to emails. So the type of words that you use, the format, is it long, is it short, etcetera. Right? So it bases it, number one, on your writing style.

Jordan Wilson [00:17:23]:
Number two, pulling in context from your past emails, which is obviously important. Right? We want AI to be smarter. And then also based, in the the the yellow portion there for our livestream audience is based on files in Google Drive. That's the part that I'm like, holy freak. This is really good. So in this example, right, it's talking about, someone's asking about, a package or a service this company offers, and it says our pampering packages range from $90 to $230 depending on your dog's size and the specific services you were looking for. So, you know, it's pulling in that information based on a Google Drive file according to, you you know, what Google, released here. So that right there, extremely impressive.

Jordan Wilson [00:18:11]:
Personalizing emails based on your writing style, based on past emails, based on files in your Google drive. When, and if this happens, I'm gonna love it. You know, I won't have, I'm embarrassed to, to do this live, but I'm gonna tell you guys the truth. Alright? I get just bombarded with emails. You know, somehow people find my my my personal emails, the emails for the podcast. Mainly, it's just a bunch of people wanting to push, you know, they're sometimes, sometimes garbage, you know, AI, products and services to you all. And I say no to many of them, but there's some great people that land in the email. But, you know, already today, I have dozens of emails and most of them are unread because right now, the Google Gemini, you know, abilities are not good, you know, to reply to emails.

Jordan Wilson [00:19:08]:
So when this happens oh, yeah. I I was gonna look. So I have 2,000, three twenty eight unread emails. I hate email. I hate it. Right. I get too many emails takes too long to respond because number one, I have to do these three things. Right.

Jordan Wilson [00:19:23]:
I have to write it in my own style. Right. I don't want people to think I'm using AI even though I will end up using AI. Right? You know, I need to pull in context from past emails. And, you know, in many instances, people are asking, hey. I wanna sponsor the podcast. I wanna do this and this. Will you come speak at our event? Right? I have all that information in different Google Drive files, but sometimes I forget it.

Jordan Wilson [00:19:45]:
So it takes a lot of time to go and do those three things. So, this personalization piece will be huge. So, this is launching in Google Labs, so you have to sign up for Google Labs. It's a free program that's essentially where you get beta access to certain tools and features. So right now, it's saying it's launching, in Google via Google Labs in July of this year. Initially, it's going to be on the web only, you know, so you can't use this inside different apps, and it's going to be English only at first. So, I'm excited for that, and the business use cases for that are obviously, off the charts. Cecilia, I a % with what Cecilia says.

Jordan Wilson [00:20:30]:
Cecilia says email is the bane of every professional, so anything that helps is more welcome. Absolutely. Absolutely. And I do know, you you know, spending spending a little, time, on on Twitter, last night looking at all the the new releases and everything else. Logan Kilpatrick, who I've had on the show a couple of times, he's lead of product, for Google and AI Studio. And he did mention that email, prior like, the email priority is extremely high because someone's like, yo. Is this actually gonna happen? And he's like, yes. It is gonna happen.

Jordan Wilson [00:21:06]:
So, you you know, vote vote of confidence there from, Chicago's own, Logan, who's been on the show a couple of times. So, yeah, I'm really looking forward to this one. Hopefully, it does come out in July. Heck, Google, I'll even take 2025. Please give us a working version of this in 2025, and the business world will be crying tears of joy. Next, Hey, tears of joy. If you're a notebook LVM user, you're gonna like these updates. It's actually crazy that this didn't make our top, you know, seven for tomorrow's show.

Jordan Wilson [00:21:41]:
But here's what's new in notebook l m. And if you don't know notebook l m, it won our 2024 AI tool or mode of the year award, and it wasn't even close. NotebookLM is an amazing piece of technology. It is powered by Gemini 2.5 now, whereas previously, it wasn't. So that just rolled out at Google Cloud Next, about six weeks ago. So if you haven't used, NotebookLM recently, you should go use it now because it uses a hybrid thinking model. So it's even better than it was before, but it is grounded in your data. So as an example, let's say I load it up, which I literally did for this show.

Jordan Wilson [00:22:19]:
I load it up with a bunch of information about Google IO updates, and I ask it about deep dish pizza. It's gonna be like, can't respond, don't know. So it is grounded in your data. It only works with what you give it, which is huge for trust, transparency, and being able to use, something with accuracy knowing that there's likely not going to be any hallucinations. So some of the cool things is, video is gonna be coming out, which is gonna be wildly fun. Alright. So not a ton of, updates yet, but there's kind of these multimedia features. One is the audio overview, which is essentially a deep dive podcast.

Jordan Wilson [00:22:58]:
It makes a podcast, between two hosts that sound very real. Right. And many of you probably feel like you even know those two AI hosts. Right? Because you listen to them all the time if you're like me. So you are going to be able to, have the default time, to either five minutes, ten minutes, or twenty minutes. So the default is ten minutes. If you click shorter when you go to customize audio overview, that's about five minutes. If you click longer, it's about twenty minutes.

Jordan Wilson [00:23:26]:
So that's great. I was able to already kinda do this with some simple, you know, quote, unquote, prompt engineering, which is just, you know, instructing it over and over, for a time or giving it more complex, request when asking it to customize to get it longer anyways. So, yeah, there's gonna be some simple video generation based on your files, which I'm excited, to see what that looks like. And then like I said, the ability for five, ten, and fifteen or or sorry, five, ten, or twenty minutes, for the, the audio overview. So, also, you know, they did update this to, 50 languages a couple of weeks ago. So, I don't think the video overviews, FYI, they're not gonna be like Veo 3 quality. Right? Something that you would, you know, produce and, you you know, go say, okay. This is gonna be our new, explainer video for our business.

Jordan Wilson [00:24:25]:
I don't think that's what we're looking at here. What we are looking at is more of a fun and, kind of cutesy way, at least the, kind of the examples that they showed were more, kind of, I would say, animated. Right? Like, more retro esque graphics, which is fine, but great for explaining more complex topics, which is something I use for anyways. So, yes, this is when you're talking about business use cases, this probably isn't something that you're gonna export and go put on the front page of your website, but I don't know. Maybe it will be or at least something that you might put on social media. I could see that as well. So, yeah, couple, new updates there. Also, there's obviously higher much higher limits for Google AI Pro and Ultra subscribers.

Jordan Wilson [00:25:14]:
Although, I think even the free limits for most people on a free plan, is more than enough. All right. Our next one, this one's interesting, a Gemini diffusion model. Okay. This is pretty big. This is pretty big. So, this is not a transformer, large language model. So, diffusion.

Jordan Wilson [00:25:52]:
How do I explain this? It's almost like a live denoising process. Alright. So Gemini and most most large language models are quote unquote traditional transformer models. Right? A very advanced, next token predictor. Right? Kind of, you could say, in theory, working from left to right, where a diffusion model, it kind of just starts with noise, and then it updates the whole thing. Alright? This is a very, nontechnical, description. Right? But this is an experimental text model using diffusion techniques, which, like I said, diffusion models are inspired by image generation methods, and this is to refine answers with exceptional speed. So, I have the example up here and and, what Google's, kind of going to be releasing this for initially is for things that are more, finite.

Jordan Wilson [00:26:48]:
Right? Things like math and coding, because that's what I think, diffusion models might be better at. You might be saying, like, okay. Like, why why do we need a diffusion model? Well, how about for speed? So Google says, their early testing show four to five x faster. Four to five times faster on math and coding text compared to comparable, you know, non diffusion models. So this is a completely, new technology. But if you do use large language model for coding, STEM, you know, specifically math tasks, I think it's going to be great. So right now, it's in limited preview and there's a wait list. So like I said, this is a very novel approach to applying, diffusion based methods to language models, which have not been used before.

Jordan Wilson [00:27:43]:
And it's really just focused on solving complex reasoning problems. So this is less about, you know, creating long form blog posts and more about working in areas that usually have more of a right or wrong answer and less about using them in areas where there's a ton of gray, if that makes sense. So, like I said, some business use cases, if you're in anything with coding, math, and if you're already finding a ton of utility by using, you know, Google Gemini or other large language models, but, you know, maybe you need more speed, this could be it. Right? So it's a completely new technology, diffusion for text based large language models. Diffusion, technology has been out there and been wildly popular for, image models. Right? So it's essentially denoising. So if you ever watch an AI image be generated in real time, which many of us do because you go and, you know, whether you're using GPT four o ImageGen or, you know, you're using, Imagine or, Midjourney. Right? It starts.

Jordan Wilson [00:28:45]:
You can watch it go live. Right? So whether it's five seconds or a minute and you see it kinda transform. It starts with this blurry, noisy outline. It's like a bunch of blobs and then slowly it comes into focus. So that's kind of like what a diffusion model does versus kind of going left to right next token prediction on steroids. So pretty interesting here with the, Google diff, or Gemini diffusion model. All right. We have three, couple, couple more here, in our part one of our top 15 features.

Jordan Wilson [00:29:22]:
Okay. So real time translation in Google meet. So this is really cool. And like I said, this is technically nothing groundbreaking. Microsoft copilot has already had this, for certain users. Right? So, Microsoft Copilot has had a version of this, for their, teams meetings, but you did have to have a certain, Copilot plus PC. So you had to be able to do this, locally on your device. So Google is bringing this to the cloud.

Jordan Wilson [00:30:00]:
So what is this? Well, it's very limited right now, but very cool. So, it is live speech translation during video calls that work like having a human interpreter present. So, in terms of availability, initially, it's only going to be available for people on the $20 a month pro or $250 a month, Ultra plan. And at least right now, it's only going to be Spanish and English. But Google did say there's more languages coming soon. So, essentially, it translates this in near real time with natural voice synthesis. So, if I was talking to, you know, some of my wife's family, in, you know, Bolivia or Chile, we could talk to each other. Right? And I was speaking English, and it would, use a voice that kind of sounds like mine in real time, translate what I'm saying to Spanish, and then translate what they're saying, from Spanish to English.

Jordan Wilson [00:31:00]:
And at least from the demos they showed, there's not a huge delay. Right? It literally sounds like a world class human translator or interpreter. Right? It like, you can't really tell, any lad. Right? So it's not like you say a full sentence and then, you know, ten seconds later, you you know, you know, the the translated version comes. It is milliseconds. It is almost instantaneous. Right? So, again, that's the demo. We'll see what actually, happens when this rolls out and specifically, how it rolls out because, you know, one of the things I'm wondering, and I am gonna be following up with, my Google contacts to get a lot of answers, to questions.

Jordan Wilson [00:31:42]:
So if you do have questions on this, let me know in the comments, because I will track down the answers. But one thing I'm wondering is, like, okay. Do both users need to have a pro plan? Right? Or can just one person, you know, beyond that $20 a month, because if both people have to have a pro plan, I think that really limits the, you you know, kind of the talking that you can do and and having this be great. But think of what this does for business. This is absolutely nutty. Right? Once this does roll out, to more countries and more languages, and I do assume, that Google will be, trying to update this, and I'm guessing that this would be in the latter part of 2025, to the 50 languages that notebook l m supports. That would be my guess. I don't have that on authority, but Google did say they're working on more languages, and it probably makes sense, to work on the 50 languages that they've already incorporated into notebook l m, which are the most widely spoken languages in the world.

Jordan Wilson [00:32:42]:
So this is huge. Even if just for right now. Right? Think if you have business in, Latin America, South America, the language barrier is gone. Yeah. You might have to, you know, if both users need to have a $20 a month, like who cares? Right. Imagine being able to talk to your colleagues from another country without a language barrier. This is huge. This opens up so many new business possibilities, especially when you look beyond where we're at now.

Jordan Wilson [00:33:16]:
Right? And like I said, this has already been out with Microsoft for many more languages, but the downside is, you had to have that running on your local device. So you had to have a newer, Copilot plus PC that essentially was running a language model locally on your own device. So, if Google can pull this off and expand it to 50 languages, This completely changes how you can do business, right? Maybe you've only been a domestic business for now. And maybe the language barrier is one of the biggest reasons why, right? This is huge. This is huge. Alright. One or two more here as we wrap up. So number nine, Gemini app updates a ton here.

Jordan Wilson [00:34:05]:
So there's been a lot of enhancements to both the Gemini mobile app and, the Gemini app. And we'll probably, in the coming weeks, we'll probably have a lot of dedicated, episodes covering this, and we're gonna be covering this a little bit more tomorrow, when we talk about Gemini live. So a lot of new updates there. But some of the Gemini app updates are rolling out now to both iOS and Android users. You get a lot of the core features free, with some of the more, premium capabilities for people on those, subscription plans. So some of the ones that I think are worth noting specifically within Gemini, deep research. Right? Anyone out there using deep research, like, every single day, like I am? I'm excited for that, but you can start deep research by uploading, PDFs or images, which is huge in terms of personalizing your deep research. And like I said, I think, a month ago, OpenAI was in a league of their own with their deep research.

Jordan Wilson [00:35:18]:
But now, I think Google Gemini is probably slightly ahead, because they did, change how their deep research works because they upgraded it to their, Gemini two point five model. So it used more thinking and reasoning and planning. Right? But if you don't know anything about deep research, essentially, it it you give it a a query, and it'll go off and spend anywhere from, you know, two to twenty minutes researching anywhere from a dozen to hundreds of websites. But now what makes it better, inside Google Gemini with 2.5 is it uses this thinking model. It plans it step by step. And a lot of times, it will make a turn. Right? It'll start going down one path, and then in its research, it finds it finds out, like, oh, I was wrong about that. So I should probably not go look, you know, at another 100 web pages if I found out I was wrong about my original plan.

Jordan Wilson [00:36:08]:
So then it'll deviate and pivot, right, which is what OpenAI's version of deep research has always done. But now Google Gemini's version does that as well. But the new thing here, is at least with deep research is being able to start with uploading a PDF or an image, which is huge. A lot of, new updates for Canvas, which we're probably gonna have multiple, shows in the very near future just looking at Gemini two point five Canvas and all of these new updates. You know, you can create now infographics, interactive quizzes, and then everything with the Gemini live that we're gonna be going over a little bit tomorrow. So, I mean, just improved response quality through personal context, more natural voice interactions with emotion detection, in the, the voice features. And when you talk about business use cases, I mean, there's a ton. Right.

Jordan Wilson [00:37:00]:
This is really where I think a lot of knowledge workers should be starting their day. Right? Whether it's it's chat, GPT, Google, Gemini, Copilot. Right? You should be starting so many of your tasks in a large language model, not in the middle, not at the end, but start with idea, strategy, research, etcetera. So a lot of these app updates, they're more than quality of life. They're changing what's possible. And then speaking of changing what's possible, and this is last on today's list but not least, GEMMA three n. So this is Google's latest fast inefficient open open source multimodal model designed for on device AI applications. So what the heck does this mean? Gemma three n.

Jordan Wilson [00:37:47]:
Well, first of all, it's scary. Good. This is a small language model, 4,000,000,000 parameters. So what does that mean? Well, without getting too technical, a small language model, a 4,000,000,000 parameter model can fit on a phone can fit on today's smartphones. Right? Edge AI and small language models have been saying this for years. This is the future of large language models because what's one of the one reasons that most enterprise companies or even individuals don't, work with large language models? Well, they're like, okay. Well, data security, you know, all those things. Okay.

Jordan Wilson [00:38:30]:
Sure. Makes sense. I don't wanna send my stuff to the cloud even though you already have everything in the cloud, and it doesn't matter. It's the same thing regardless. For for for those that aren't smart enough to make that connection and, you know, figure out that, you know, one plus one equals two, I don't know in the new the new math, the core math, if one plus one still equals two. But one plus one still equals two here, because when you talk about edge AI, that takes out all those data security things, because you're not sending any information to the cloud. You can shut off your internet and use Gemma three n on a local device. Right? And the performance is absolutely nutty.

Jordan Wilson [00:39:12]:
Okay? Claude three point seven SONNET is one of the world's most powerful proprietary models. Obviously, you have to use it in the cloud. Right? Because it is enormous. You know, we don't know how big it is, but chances it's a couple trillion parameters, or at the very least, at least a hundreds of billions of parameters, which just means size. Right? Think of like a like a gigabyte, of storage or something like that. Gemma three n is a fraction. I would say it is less than 5% of the size of Claude three point seven Sonnet. Yet, for chatbot arena Elo scores, so side by side comparisons, it is essentially the same.

Jordan Wilson [00:39:53]:
Right? There's only a four point difference. So that means when humans don't know the difference and, you know, everyone I don't I I'm not a huge quad fan, FYI. But quad's latest model, although there's rumors that they might be releasing, you know, quad, for SONNET or quad for Opus any day now, but at least their most powerful proprietary model. This itty bitty model that you can download, you can fork it, you can do whatever you want is just as powerful, Just as powerful. So, the availability is the preview of this is available now via Google AI Studio. Also, Google AI Edge, it's free for developers. You can download it, you you know, fork it, fine tune it with your company's data, etcetera. So it's engineered to run smoothly on phones, laptops, and tablets with minimal resource requirements, and it can it is multimodal as well.

Jordan Wilson [00:40:50]:
It can handle audio, text, image, and video inputs. That is amazing. So it's optimized for resource constrained environments while maintaining strong capabilities and modalities. Also, speed fast. Right? Not having to send something to the cloud and wait for the inference, for it to go do its thing on the cloud. It's happening on device. So it's faster. It's more secure.

Jordan Wilson [00:41:13]:
And I've been saying this for a long time ever since we saw the first version of Gemma three, a couple of months ago. I said, don't sleep on Gemma three. It is wildly powerful. Y'all this, this completely changes how we're going to work in the future because what this signals, what this signals is this is going to force the other big companies, open AI, anthropic, etcetera. Those companies that don't have an open model yet, this is going to force them to go open Because if you have a Gemma model right? And also, you you know, there's good open esque models from Mistral, from, Meta, their llama models as well. But, I mean, this right now, Gemma three ed is benching off the charts for how small it is. Right? This is gonna force big companies that are only doing proprietary models to offer open models. Right? OpenAI CEO Sam Altman did say that they're going to be releasing something, but this is huge because this means that probably within, I don't know, a year or two, most new computers I mean, well, I won't speak for Apple since they're still operating in the nineteen nineties when it comes to artificial intelligence, but you would have to think even Apple is gonna have to catch up.

Jordan Wilson [00:42:30]:
Most computers are going to come with a state of the art level, large language model that can run everything locally. So you won't even have to worry about data security because nothing's leaving your hard drive. It's the same thing as saving, saving a file to your local device, working with a model like Gemma three n. So this is huge. So that's our quick recap of what's new at least on the first half here. If you want some related episodes y'all, I've had some recent ones. So I was at Google Cloud Next, a couple of weeks ago and covered what was new there, with, Logan Kilpatrick. Already mentioned that.

Jordan Wilson [00:43:15]:
So if you wanna go listen to that, that's in episode five zero one. Also, there's been a lot of new updates just with Gemini 2.5 pro, and we're gonna be talking about some of those even newer updates tomorrow. So if you wanna go get caught up, go listen to episodes four ninety four and four ninety five as we do a two part series on Gemini 2.5. You don't gotta wait for anything. It's live. It's there. It's free at our website. Go listen to it.

Jordan Wilson [00:43:37]:
So a very quick rundown as we wrap things up, our part one. Here we go. Number 15, imagine four. 14, Chrome with Gemini integration. 13, personalization and email. 12, notebook l m updates, which I'm extremely excited about, 11, Gemini diffusion, a brand new type of large language model, 10, real time translation in Google Meet, only English, Spanish now, but more coming soon, nine, Gemini app updates, And eight, Gemma three and the world's most powerful small language model. It is insanely good. I can't wait for tomorrow.

Jordan Wilson [00:44:12]:
Make sure you tune in for part two. I'm telling you some of these things that we saw, mind boggling. I don't even know how I'm going to verbalize it with words even though that's all I do. So thank you for tuning in. If you haven't already, please go to your everydayai.com. Sign up for the free daily newsletter. Please make sure you join us tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI