Episode Categories:
Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
Maximizing Value with Google’s Latest AI Tools: Gemini 3.5 Flash, Omni, and Antigravity 2.0
Google’s 2024 I/O event unveiled more than 100 AI-centric updates, yet only a handful are immediately accessible and relevant to business operations. Key among these are Gemini 3.5 Flash, the Omni Flash video model, and the Antigravity 2.0 app. Each brings both opportunities and potential drawbacks to organizations navigating AI integration. Below is a breakdown of their particular business value, pitfalls, and specific areas that demand attention for maximizing operational impact.
Google Gemini 3.5 Flash Model: Performance and Cost Trade-offs
Benchmarks and Speed
Gemini 3.5 Flash stands as Google’s top-performing AI model to date. Its main advantage is speed; it processes and returns output faster than any previous Gemini model. In direct side-by-side benchmarks, Gemini 3.5 Flash outpaces its predecessor, Gemini 3.1 Pro, across almost all categories, especially in tasks related to agentic coding and long-running agents. Key benchmarks reveal significant performance improvements in terminal bench and agentic tasks, which are directly tied to productivity in automation and coding 22:08, 22:36.
Cost Considerations for API and Enterprise Use
However, the shift from “Flash” being synonymous with speed and low cost to now being fast yet expensive is critical. Gemini 3.5 Flash is no longer cost-effective for API-heavy businesses, with usage costs up to 10-20x higher than previous Flash models for the same token consumption 22:48. Additionally, Gemini 3.5 Flash is one of the most token-hungry models available, leading to higher expenses for each task executed, a key factor for organizations calculating TCO (total cost of ownership) for AI deployments 27:13.
Impact of New Usage Limits
Usage restrictions have shifted: what used to be an unlimited, “buffet-style” model is now bound by revised limits, which for many professionals can be reached quickly depending on operational activity 48:05. This change impacts scaling plans and could disrupt workflows for organizations previously accustomed to higher usage ceilings.
Gemini Omni Flash: Beyond AI Video Generation
Not Just a Video Model—A Multimodal Asset
The Gemini Omni Flash is frequently branded as Google’s next video generation tool, replacing the popular Veo series. However, its core capability is true multimodality: input and output across image, video, text, and audio 28:38. This means that instead of being limited to text-to-video, Omni Flash enables businesses to orchestrate grounded video editing, dynamic object/background modifications, and cross-media tasks within a single interface, streamlining creative and technical workflows 29:32.
Character Consistency and Conversational Editing
A significant improvement over prior video models is Omni’s character consistency and editing features. Users can expect preservation of identity, voice, and visual narrative across scenes, elevating the reliability of brand-aligned video content for marketing, training, or onboarding outputs. Conversational editing also allows quick, multi-turn revisions, vastly reducing time-to-delivery on creative projects 29:06, 29:32.
Access and Limitations
Currently, Omni Flash is limited to 10-second video clips per generation and is available primarily to AI Pro and Ultra subscribers via the Gemini app and Google’s AI platform. This restriction matters for planning marketing or production timelines 29:06.
Antigravity 2.0 Desktop App: Early-Stage Agentic Coding and Productivity
Hands-On Functionality and Interoperability
Antigravity 2.0 is Google’s first mainstream push into agentic coding desktop applications, compatible with Mac, Linux, and Windows without the need for an IDE. The platform provides project-based conversation, scheduling of tasks, and enables read/write access to folders locally, supporting tasks such as automated file management 10:12, 12:27.
Unlike some competitors, Antigravity 2.0 allows choice among multiple AI models—including Claude Sonnet, Claude Opus, open source GPT OSS, and the latest Gemini variants—directly within the workflow. This could be beneficial for organizations seeking flexibility in model deployment without third-party tool dependency 11:12.
Missing Integrations and Workflow Gaps
Despite its potential, Antigravity 2.0 currently lacks direct connectors to critical Google Workspace tools like Gmail, Drive, Sheets, and Calendar—an unusual gap for an app in Google’s ecosystem 14:06. Competing agentic coding tools already deliver seamless app integrations, raising questions regarding Antigravity’s suitability for central command-and-control or cross-system business automation.
Performance and Output Quality
Initial field testing reveals that while Antigravity 2.0 can generate fully functional interactive web apps (e.g., auto-curated episode databases with filtering and navigation tools), some essential links do not work as intended without manual corrections, and the design output is generic 18:04.
Moreover, although the app claims to orchestrate multiple parallel agents, many users report hitting usage or task execution limits rapidly even on paid plans, complicating the path to true workflow parallelization 08:39, 16:00.
Comparing Market Fit and Business Value Among Google’s New AI Tools
Model Selection and Transparency Deficits
Consistent user experience is currently lacking across different Gemini interfaces and account types. The same user may not have uniform access to model selectors or reasoning controls—such as extended thinking modes—across personal and business accounts, causing friction in standardized AI-driven business processes 35:03.
Chain-of-thought transparency, a feature seen in competitor products, is still limited in Gemini 3.5 Flash. Users get only a partial view of how the model thinks and solves problems, which makes auditability and compliance reporting less comprehensive 37:08.
API Builders Must Watch Token Efficiency
For businesses leveraging Gemini 3.5 Flash on the API side, the model’s high token consumption and associated costs cannot be ignored. While the model delivers state-of-the-art output, a direct swap from Gemini 3.1 Pro or earlier Flash versions will result in significant cost spikes without commensurate efficiency for all use cases 27:19, 27:56.
Where Each Tool Delivers True Value
Gemini 3.5 Flash: Delivers top-tier speed and intelligence for critical use cases where time-to-answer is business-sensitive, but requires budgetary oversight due to higher costs and revised usage limits.
Gemini Omni Flash: Best deployed where multimodal creativity, rapid content editing, or high-fidelity brand/character continuity is non-negotiable. Current version is ideal for short-form output and organizations exploring advanced interactive storytelling.
Antigravity 2.0: While still behind industry-leading competitors in integrations and output testing, it offers model diversity and basic task automation for businesses willing to pilot agentic apps without complex infrastructure setup.
Strategic Recommendations for Business Leaders Evaluating Google’s AI Suite
Several points surfaced that shape practical next steps:
Conduct detailed ROI and TCO analysis before fully adopting Gemini 3.5 Flash, given its step-change in usage costs.
Pilot Gemini Omni Flash for creative workflows that need cross-modal content generation and dynamic video editing, but verify output suitability within existing brand standards.
Approach Antigravity 2.0 as an early-stage tool; current limitations mean it may serve best in experimental or limited automation roles until more mature integrations arrive.
Track upcoming releases, namely Gemini 3.5 Pro, which is slated to address current performance/cost imbalances.
Stay alert to shifting usage caps and model access provisioning across Google account types to avoid workflow interruption during expansion.
The latest Google AI offerings open specific operational opportunities—speed, multimodal creativity, agentic automation—while surfacing immediate cost, usability, and integration considerations that demand measured deployment strategies.
Topics Covered in This Episode:
- Gemini 3.5 Flash Model Hands-On Demo
- Gemini 3.5 Flash Pricing and Token Usage
- Benchmarks: Gemini 3.5 Flash vs. 3.1 Pro
- Intelligence vs. Cost in Gemini 3.5 Flash
- Gemini 3.5 Flash for API and Developers
- Google Gemini Omni Flash Video Model Review
- Omni Anything-to-Anything Multimodal Features
- Google Omni vs. Video Model Competitors
- Anti Gravity 2.0 Agent Desktop App Overview
- Anti Gravity 2.0 Pros, Cons, and Use Cases
- Usage Limits in Google Gemini and Anti Gravity
- Chain of Thought Transparency in Gemini Models
- Canvas Mode Interactive Web App Demonstrations
Episode Transcript
Jordan Wilson [00:00:16]:
Google literally had a blog post called 100 things we announced at IO twenty twenty six, and most of them were obviously AI based. And on Wednesdays, we go hands on with the latest and greatest AI updates in our AI at work on Wednesday series. Obviously, we're not gonna dive into all 100 updates that we got from Google's IO conference. These shows are already long enough. But of the dozen or so key AI updates that are available to most users right now, three seem to be getting the most attention out of the gate for better or worse. And that's the new Gemini 3.5 flash, the Gemini Omni flash video model, and the new anti gravity two point o app. It's although the benchmarks and the key notes and the initial PR coverage might have you believing that everything that came out of Google IO was all token friendly sparkling unicorns. That was not the case as there's some new wrinkles in the Google Gemini armor that weren't there previously.
Jordan Wilson [00:01:26]:
So we'll fill you in on all of that and give the hands on live demo treatment to Gemini three flash, Gemini Omni flash, and anti gravity two point o on today's show. What could go wrong? Well, let's get to it and put AI to work this Wednesday and highlight the good and the bad of these three Gemini standouts and show you hands on in life how to use them. Alright. So if you're new here, this is Everyday AI, but I wanna zoom out and talk about the big picture. Gemini 3.5 Flash is actually a really good model, but it might not mean what you think it means if you are, well, familiar with the Google Gemini family. So Gemini 3.5 Flash is a very capable and technically really good model. It's actually Google's best AI model out there. Right? But the name flash used to mean fast and cheap, but it doesn't really mean that anymore because now it's just really fast, but actually kind of expensive.
Jordan Wilson [00:02:31]:
And it is Google's best model. Gemini Omni is an impressive anything to anything world model, but I think people are mostly judging it as an AI video model, and they're only looking at its ability to output video, which I think is a big miss. And then anti gravity two point o is Google's first mainstream foray at a desktop app for agentic coding, yet it's extremely underwhelming. And I'm trying to find any reason I can to use it, but I can't. So on today's show, stick around with me for the next twenty five to thirty minutes. Here's what you're gonna learn. You're gonna know why Google's new best model actually got both better and a lot worse for both Gemini and API users. You're gonna know why Omni is better than the eye test and side by side comparisons actually suggest.
Jordan Wilson [00:03:26]:
And I'm gonna tell you why you should probably just skip anti gravity two point o for now. Alright. Let's get to it. Welcome to Everyday AI. If you're new here, my name is Jordan Wilson, and, well, we do this every day. This is your daily unedited, unscripted livestream podcast and free daily newsletter helping everyday business leaders like you and me not just keep up with what's happening in the world of AI, but how we can take all this information and use it to actually grow our companies and our careers. So that is what our, you know, Wednesday show really does. I say, let's go.
Jordan Wilson [00:03:57]:
I'm gonna be sharing my screen, my desktop live, and we're gonna learn. This is one of those, and most Wednesdays Wednesday shows might be a little better for our podcast audience to watch the video version. So you can always do that on our website at youreverydayai.com. However, I am gonna do my best to narrate exactly what's happening on the screen just in case you're like me and you like podcasts the old fashioned way just with the audio. Alright. And make sure you do check out today's news, newsletter. We're gonna be recapping the highlights of today's show as well as all of the other AI news you need to know to be the smartest person in AI at your company. Alright.
Jordan Wilson [00:04:32]:
So like I said, a ton new happening, at Google. So, they had their IO conference now, about, two weeks ago act or no. One week ago. Time flies. Time flies when you're, in codex all day. So their, conference was last week, and they announced a ton. Let me just be honest. This isn't the the these three, products are not the three that I'm most excited about.
Jordan Wilson [00:04:58]:
Probably the most that I'm excited about is Gemini Spark. It's new, Google's new kind of always on agent, but not everything they announced is available to everyone. So, again, I've kind of joked around about this, but this is the reality. There's a 100 new AI updates that they announced. And number one, they're not available to everyone. It's across different tiers, depends on what plan you have, what, you know, settings you have enabled in your workspace, if they're on workspace, if you're using a personal Gmail. Right? So there's so many, you know, hoops and fences you have to jump over and through just to understand what's going on. But I will tell you this, A lot of the, more impressive or more kind of headline grabbing updates that Google talked about aren't available to the masses, like the Gemini spark agent, which is essentially, for lack of better terms, Google's version at an online or cloud version of OpenClaw.
Jordan Wilson [00:05:54]:
Right? That's kind of like what everyone's been announcing. Right? Here's an always on autonomous agent that runs in the background and connects to all of your data, but is secure. Right? So that's what, GeminiSpark is, but it's not available to everyone. Right? There's another, feature that I actually only have access to in my personal account, so it's not really worth showcasing in my work account, which is called the daily brief. Right? So there's so many dozens of actually really good and really exciting and, use cases I would personally love to show, but it's kind of hard because not everything is announced now. You know, some things are rolling out in the coming weeks, the coming months, you know, this summer. Right? So I don't know. You almost have to have an agent, and I actually did this.
Jordan Wilson [00:06:37]:
I had codex go through and check the 100 blog posts and map everything out. And even there, it's still so confusing. Right? Because as an example, I keep talking about Gemini Sparks. It's, their spark agent because I wanna use it. Right? So, if you have a Gemini business plan, the release notes said that it was available in Gemini enterprise. But on the main home page, promoting everything Gemini Enterprise, it's not on there. Right? So a lot of confusion on what's available, on what price tier, when, and how, and who, and, you know, do I have to know someone at Google to get access to all of these? I do know people at Google, and I still don't even understand which account I should use. Anyways, alright.
Jordan Wilson [00:07:17]:
Let's get into it. Enough of my my little riff there. So, actually, we're gonna start live. Do this one a little different. What could go wrong? Livestream audience. If you could let me know. Alright. I'm gonna be sharing my, entire screen here.
Jordan Wilson [00:07:32]:
Let me just make sure I'm not sharing anything, too personally identifiable here. Right? Alright. So, podcast audience, I am sharing my entire screen here. And, I have a couple of things, popped up. We're gonna start some things live now, and then we're gonna gonna get back to, some of the details on what's new, bullet point, all that good stuff. So what we're gonna do is we are going to start in anti gravity. So what is anti gravity? Well, to me, it's a, piece of software looking for a problem to solve. Because I'll say this.
Jordan Wilson [00:08:14]:
Most people are not Gemini exclusive. Right? A lot of people, I think, are using Gemini and something else. Right? So whether that's Gemini and Copilot, Gemini and Chat GPT, Gemini and Claude. Right? If you are using gemini.google.com. Right? So the Gemini chatbot, if you have that $20 pro plan or a $100, you know, the new $100 ultra plan, Google also brought down the $200 ultra plan, down from $2.50. Right? So if you have any of those plans, I'm guessing you probably use something else. So if you have a $20 a month chat GPT plan, then you have access to codex. If you have the $20 quad AI plan, you have access to, Claude Code and Claude Cowork.
Jordan Wilson [00:08:57]:
So this is, anti gravity is kind of Google's version, for lack of better terms, of something like codex or like Claude Code. So it is, their kind of fully integrated agentic desktop system where you can run agents in parallel and kind of connect to your data, but I'm not even really sure how. Anyways, so on my screen, I'm showing this, the usage because this is one of the big downsides in some of my earlier testing. The usage is is actually fairly bad now. Right? It's not anthropic bad, but we're gonna be following along, a a long live. So I had to switch over to one of my other paid accounts. So I'm on a $20 a month account. You'll see on my screen here inside of anti gravity.
Jordan Wilson [00:09:44]:
I am fresh. I have a 100% remaining. And then if we look into my, inside, Google gemini, gemini.google.com, same thing. I have a 100% of my usage remaining. So I'm doing this live. We'll see. Because of my testing, it was really, really bad. Alright.
Jordan Wilson [00:10:01]:
So first, let me give everyone a quick, kind of tour around anti gravity. It's gonna be quick because at least right now, there's not a lot to see. Right? If you are a heavy codex user like me or Claude code, Claude co work, cursor, etcetera, if you're using any desktop agents tool, you're gonna look at anti gravity and be like, well, where is everything? Well, this is it. So it is very beta ish. There's not a lot to see, But I will tell you one big advantage, that you can well, with Google, anti, Google's anti gravity two point o desktop app versus something like codex or cloud code is while you actually have different models that you can use. Alright. So it is all based on your plan. So aside from the new Gemini 3.5 flash, which is the models we're gonna be using, obviously, you have the other Google Gemini model, Google Gemini 3.1 pro, but you also have Claude SONNET four six, Claude Opus four six.
Jordan Wilson [00:11:01]:
You have the open source version of g p t o s s, the one twenty b. So this is actually OpenAI's open source model. So I do like that with Google. You have, you can use different models from different providers. Right? Obviously, you can do that with a lot of third party, you know, or cursor, you you know, but within Google, it's nice that you're not just quote unquote forced to use Gemini models. Although, obviously, the Gemini models are really good. Alright. So just kind of starting with the unique, kind of call out there.
Jordan Wilson [00:11:35]:
Aside from that, I don't know. To me, there's nothing really unique, at least right now, inside, anti gravity. If you look at the interface, you know, people are joking around. I already joked around about this, and Google's already been kind of dragged through this. It looks very codex esque. Right? Even in their launch video, the two minutes, it was like a sixteen minute long video, of going over aids I gravity. Literally, they had a desktop icon with the word codex in it. So maybe they drew a little bit too much inspiration from codex because it literally looks just like codex.
Jordan Wilson [00:12:08]:
Anyways, on the left hand side, you can start new conversations. You can create new projects, new chats that are not part of projects. So again, literally just like codex. And, kind of the big difference here, is you can give Google anti gravity, read, write access to any folder or folders on your computer. Right? That's one big, kind of, advantage of these kind of desktop coding agents as well as you have scheduled tasks. So that's the other thing. So you can, you know, let's just say you wanted to clean your downloads folder. I know this is a lame example.
Jordan Wilson [00:12:41]:
Let's say you wanted to clean your downloads folder once a week on Sunday nights. You wake up, you know, Monday morning refreshed. Right? So you can schedule something like that, and it will run all of these things on a schedule. Alright. So if I go into a, new chat here. Alright. So you can create a new project. Alright.
Jordan Wilson [00:13:02]:
And this is gonna pull up a, all your folders, right there. So I'm not gonna do that right now. I'm just gonna go into a folder I currently have, set up, and I'm gonna start a new conversation. So when you start a new conversation, for this one, it has access to the downloads folder, and then I can click backslash for different actions. So it does have a goal mode, which is nice. You can run things on a schedule, browser. So and then there's also built in skills. So by using the backslash command, you get some prescheduled, kind of actions as well as skills that you can use.
Jordan Wilson [00:13:38]:
Alright? Or if you click the at button, you can, reference other conversations or different agent rules as an example. Alright. So, nothing there to really see. Very straightforward. The thing that I didn't fully understand, and I'm like, wait. I don't understand this. Right? I'm clicking around the account, the permissions, the appearance, the models, the customization, the browser, all of these side panel settings. And I'm like, okay.
Jordan Wilson [00:14:06]:
Where do I kind of connect my my Gmail, my Google Drive, my Google Sheets, my Google Calendar, all these Google products that you would just assume, and I'm like, they're not here? I don't know. I keep thinking like, oh, I have to be wrong on that, but, you know, I'm I'm looking. I'm searching. I don't see them, which makes to me absolutely no sense. Right? Obviously, you know, codex, Claude Code, Claude Cowork, they all have connectors, cursor, right, apps, connectors, whatever you wanna call it. Integrations. I don't know. At least I'm looking at Google anti gravity.
Jordan Wilson [00:14:40]:
Nothing. Anyways, let's put it to the test. Alright? So here's what we're gonna do. I have a prompt ready to go, and we're gonna check back on this in a little bit, and we're going to check at the usage. So all I said here, just so you can kinda see what's cooking and you kinda see anti gravity starting to work here. I said carefully look up the start here series by everyday AI, and carefully and meticulously denote the main characteristics, takeaway, etcetera of each episode. Then create a simple but executable web app that is interactive, filterable, and sortable for a novice user who is wanting an easier way to navigate through the series. Make sure to include a way that ultimately drives people to everyday AI's official resources for the start here series.
Jordan Wilson [00:15:28]:
Alright. And then I'm saying at the end of your output, please give me the exact way to run or launch this. Alright. And you'll see, anti gravity. It can't be done yet. Alright. It might already be done, but let's see if it actually properly. No.
Jordan Wilson [00:15:43]:
Okay. It's still going. It's just breaking this down, into separate tasks. So I would love to be able to run another test in parallel. But in my testing beforehand, when I did that, actually, none of it worked because it all timed out. It all it hit me over my usage limit. So I will, make sure to do just that one, and then we'll go back in later and see if we can't do another anti gravity test. Alright.
Jordan Wilson [00:16:08]:
So, stepping away from the live demo for just a second. Alright. Well, I was gonna say, let's go back and check on it, but it looks like it's looks like it's already done. That wasn't, that wasn't too bad. I thought it was going to work for a little bit longer. All right. So I'm just making sure I can open it up here and there we go. All right.
Jordan Wilson [00:16:34]:
So let's take a look here. So it created a nice, you know, kind of local host. Right? So this isn't a a hosted website or anything like that, but, it does have, everything here. Right? So it's a nice kind of interactive, web page. It went through. I'm kind of looking here, to make sure that it got all of them. Alright. So unfortunately, when I'm filtering, I okay.
Jordan Wilson [00:17:09]:
Here we go. All categories. So first I'm seeing we have our start here series, which is the podcast for beginners, and people who are trying to grow in their, knowledge of AI. And it looks so far pretty good. Everything in here, it looks like it got all 25 episodes. I'm looking at the title. Right? So, volume one, was generative AI, how it works, and why it matters in 2026 more than ever. That is correct.
Jordan Wilson [00:17:40]:
Right? And then let's see the latest build by partner or wait, the four layer stack decision framework for 2026. So this is good. So it not only got all of the volumes correct to the titles correct. Pretty good overall. The design leaves a lot to be desired. I didn't give it any, you know, certain prompting or anything like that. So it looks very AI. Right? Like, very AI generated.
Jordan Wilson [00:18:06]:
I wouldn't call this, you know, visual slop, but if you've used any, you know, anything, on the vibe coding side to create anything, it looks very vibe code sloppy. Right? So the dark mode with the, you know, the neon gradients, so not terrible, but, you know, nothing I would actually want to use or to put out there. And then when you click on something so let's go click on our readers, sorry, our latest episode, the build by partner or wait. So I click in this. There's a nice little modal. It has some additional, key takeaways and action points, some related tags. It has the episode number, the Git four full playlist, and then the list in or show notes. So, let's actually see.
Jordan Wilson [00:18:50]:
So the Git full playlist does not actually work. That is not gonna get us the full playlist. And then the listen show notes, let's see where that takes us. Okay. So it just took us, I didn't even know this page existed on our website. So those buttons in theory, it took you somewhere, but it did not actually take you to where you want it to go. So, you know, as an example, if I was using codecs, it would have, number one, taken a lot longer. And I know because I've done this test before.
Jordan Wilson [00:19:21]:
This is one thing I'm, you know, thinking about making part of my, you you know, testing series that I do a lot because it actually requires a lot of different tool calling. Right? It has to go to, the website a lot of times. It has to read at least 25 different pages. Right? But I know codex actually when it builds these things, it clicks through and it can test the links. I think Cloudico does a pretty good job in that as well. So, I wouldn't say the, the anti gravity version failed here by any means. It looks fine, but, some some things that don't really work. But the fact that I'm critiquing something that looks like this.
Jordan Wilson [00:19:57]:
Right? I have a little bit of a background in, you know, web design. But pre AI to build something like this, it would have taken, I don't know, would have taken at least thirty hours of nonstop work. Right? Even if I'm working with a template to be able to go through, grab all this information, put it up on there, you know, it is nice, smooth. There's animations. There's filters. So it is almost like a small database. Right? It is 25 episodes with correct ish links out, you know, additional information, you know, the right episode number related tags. So overall, pretty good.
Jordan Wilson [00:20:34]:
Alright. Let's get back though to going over some of the details of a Gemini three flat 3.5 flash, the Omni flash video model, and then also anti gravity. And then we'll go back hands on with Gemini, Flash 3.5 in the Omni, video or any to any model. So real quick, Gemini 3.5 flash, this is not going to be Google's final 3.5 model. So Google did announce that they, next month, will be coming out with a Gemini 3.5 pro, model or Gemini Pro 3.5. So we will see what happens with 3.5 flash. But right now, it is actually better. The flash version is better than Gemini 3.1 pro, which we haven't seen that yet.
Jordan Wilson [00:21:28]:
Right? So always Google Gemini three, you know, Gemini x Pro is always Google's best model. So this is a little weird because the Flash kind of moniker in the Flash series up until now has always been about being super fast and super cheap to use. It is even faster than before, but it is not cheap. Alright. So depending on, what you're looking at, you know, you might actually be paying 10 to 20 x more if you are using this on the API side than previous, you know, Flash versions, and I do have some examples of that. But it does outperform at Gemini 3.1 Pro right now on a lot of different benchmarks. It's available now whether you're using the Gemini app. It is also, available in search.
Jordan Wilson [00:22:14]:
So the new AI mode is powered by Gemini 3.5 flash, which is really cool, And also it's available on the API as well. So speaking of those benchmarks, yeah, it's pretty good here. So showing a, a screenshot here of all the different benchmarks from Google's, announcement blog post. And you'll see, for the most part, it beats Gemini 3.1 pro in every single benchmark category. Some of them are actually pretty notable. Right? Like, if you look at terminal bench, you you know, which is important for agentic coding, you know, pretty big jump up from Gemini 3.1 pro. The same thing with some of these, more agentic, you know, the agentic benchmarks, it does really, really well in. I mean, we'll see here.
Jordan Wilson [00:23:07]:
I think, really what Google is, I think, going to be using, Gemini 3.5 flash four in the future, obviously, after, Gemini 3.5 pro is well to run long running agents. What's interesting here, in this Google, benchmark sheet that they, released, against the Gemini. So Gemini 3.5 Flash, Gemini three Flash, Gemini three one Pro, Claude Sonnet four seven, Claude Opus 4 sorry. Claude Sonnet 46, Claude Opus47, and GPT 55. The interesting thing is Gemini three Flash, even though these are cherry picked, benchmarks, obviously, because there's dozens of benchmarks, out there. So they cherry picked the ones that, you know, they thought told a good story. But interestingly enough, GPT 5.5 and Gemini 3.5 Flash are winning the same amount of benchmarks on the ones they chose. So, you know, when people are like, like, Jordy, like, you're crazy.
Jordan Wilson [00:24:09]:
You're talking about codex and g b d five five too much. I'm like, no. It's right now, it's just the best, and it's not really close. And, I continue to say this. I don't think Gemini 3.5 flash overall is anywhere in the realm of GPD 5.5. And when we look at the artificial analysis, artificial analysis intelligence index, you'll see that holds true. Right? So artificial, analysis, a great, kind of benchmark, conglomerator. Right? It just throws all these different benchmarks together and gives all the models an intelligence score.
Jordan Wilson [00:24:40]:
So, Gemini 3.5 flash, you know, did fairly well considering it is a flash model. So it is still on this benchmark at least well behind Gemini 3.1 pro, and still very, very far behind the best overall model in the world, which is still, GBT 5.5 extra high or GBT 5.5 pro. So that is still very far ahead of, you know, Claude Opus four seven or Gemini three one pro. But the thing that you have to call out is Gemini three five flash is no longer cheap. It is actually very expensive compared to anything in the Flash family. I actually think it would have been it's a more, apt name or a more accurate name. It just if if they call this, like, 3.2 pro or 3.1 pro fast because the only difference, right, it is not anything like the flash models from before except that it's fast, but it's actually really expensive. So the artificial analysis, has the intelligence versus cost kind of index.
Jordan Wilson [00:25:50]:
Right? So, I have this, kind of, screenshot here that I'm showing to our audience. And for the most part, you want to have a model that is intelligent. So, on one axis, it has the, artificial, analysis intelligence index score. And then on the other one, it's the cost to run the index. So this is a set of tests, and then artificial analysis literally runs all of these tests via the API and adds up how much it costs to essentially run all these tests to get these scores. Right? So essentially, you wanna be in the upper left hand quadrant, which is a very smart model, but doesn't cost a lot of money to run. So, obviously, you wanna be in the upper left hand quadrant. So Gemini 3.1 pro, well, it was.
Jordan Wilson [00:26:31]:
Right? It's in the upper left hand quadrant along with models like GBT 5.5 medium. You have some open source models that are kind of on the edge there. But for the most part, it was Gemini 3.1 pro, GBT five five medium, Quen three seven max, and then there's some on the edge there. But Gemini 3.5 Flash, again, it's not up there in terms of intelligence. It's not a, you know, top five, you know, intelligence model out there right now overall according to artificial analysis, but it's actually fairly expensive. It was much more expensive to run these tests on Gemini 3.5 Flash than it was, on Gemini 3.1 Pro. That's because it is token hungry. Gemini 3.5 Flash eats way more tokens than it needs to.
Jordan Wilson [00:27:19]:
And why is that important? Well, it's important both on the API side. So if you are building on these models in your company, if you're using these models, which is gonna become easier and easier and probably more relevant and more prominent in the future because I think kind of this whole, you know, token subsidy thing is going away. Right? So Google has kinda getting gotten rid of it like I talked about. They've introduced these new usage limits, which they didn't really have before. Anthropic's limits aren't really usable if you're paying for a, you know, 20 or 100 or $200, a month subscription. So the only out of the big names, right, is just OpenAI that still has these, you know, usage limits that seem like they're, oh my gosh. It seems like they're unlimited. So anyways, intelligence versus cost to run Gemini 3.5 flash costs a lot of money if you're using it on the API side in terms of the intelligence that it can output per dollar.
Jordan Wilson [00:28:15]:
Alright. So let's talk about Google's new anything to anything model, which for the most part, even myself, right, we're calling it their new video model because it does replace their VO series. But the Google Omni flash is Google's new multimodal model that combines Gemini's reasoning reasoning with generative video creation. So, what makes this different than your run of the mill, you know, text in video out model. Well, is it accepts a lot of different formats on the front end. It accepts image, audio, video, text inputs, as well as, outputting grounded video that you can actually edit. So this does replace VO, and it's available in the Gemini app in Google flow for AI plus pro and ultra subscribers. You can only generate ten second clips at a time, but there is multi turn conversational editing at launch, which is really cool.
Jordan Wilson [00:29:13]:
Some things that really stood out with Omni versus the v o 3.1, the last video model from Google, is improved character consistency, preserving identity and voice access across every scene, and it excels at video editing, like replacing objects, backgrounds, foregrounds, etcetera. So, this is the thing. This is a world model first. Right? One thing that I think Google has always been on the cutting edge on, and I don't know if any of the major labs are going to be able to catch them anytime soon, is just being multimodal by default. Right? This is something I I told you guys, like, two years ago. I'm like, why aren't more people talking about, like, you know, using the I think it was Gemini 2.5 Gemini 2.5 Pro might have been the first model that could ingest video. Right? Ingesting video, in Google's AI studio. Not just being able to, you know, actually have that as an upload source, but it could look at it every single frame and tell you what's going on.
Jordan Wilson [00:30:13]:
Right? So Google has always had this huge benefit that I think is gonna slowly reap the rewards in the years to come when we talk about world models, robotics, embodied AI, etcetera. Right? But this is where the Omni flash version might eventually show, some of the the the the fruits of those labor. Right? As well as there will be other versions of Gemini Omni according to Google. So So we might get either a Gemini Omni, you know, versus just the Flash version or, you know, a Gemini Omni Pro. So and then let's quickly bullet point what's new in the, anti gravity two point o. So this is Google's standalone agent first desktop app for Mac, Linux, and Windows with no IDE required. It is the central home for orchestrating multiple AI agents and executing tasks in parallel. We'll see about that because my test did not allow me to on a paid plan.
Jordan Wilson [00:31:11]:
It can dynamically add sub agents. You can schedule tasks in native, voice commands. It's a suite that also includes the new anti gravity c l I in s SDK in managed agents. And it is powered by Gemini three five flash and co developed using anti gravity itself and apparently, maybe codex as inspiration. And, one of the bummers, one of the downsides, a lot of people are upset about this. I'm not really a big command line interface CLI user, but it does replace the Gemini CLI, which is gonna sunset for consumer users next month, so June 18. And a lot of people are upset about that because Gemini CLI was a, I think, a great product, if you were more of a technical command line user, and it was open source. So Google shutting down a popular open source project and instead making it a closed source.
Jordan Wilson [00:32:04]:
So, alright, that's enough for the bullet points. Let's go back live, shall we? Alright. So I'm gonna go ahead and bring back up anti gravity. Alright. Live stream audience. Let me know if we got it back. Do we got it back? Here we go. Alright.
Jordan Wilson [00:32:22]:
So we got our one. Alright. And I'm gonna get another anti gravity prompt ready. We're not gonna sit around and watch it, but I'm curious, what happened with our, what happened with our limits here? So let's go to our models. Okay. So this is very strange. Alright. All our models still say a 100% remaining.
Jordan Wilson [00:32:52]:
So maybe I was running into a bug because when I did this one prompt, it exhausted all of my limits on a paid plan before. But you'll see now okay. Good. So, I'm clicking refresh again. Alright. So maybe we're good. I'm actually gonna literally close this down and reopen anti gravity just to double check. Alright.
Jordan Wilson [00:33:19]:
Alright. So, apparently, we still have a 100%, usage remaining. So I'm actually gonna just go ahead and do one more, one more little, prompt there. We'll check-in on it at the end, but let's go ahead and do the same now for Google Gemini. So I always have kind of a funky little, rubric that I take some models through, just to really see, are they smart? Are they fast? How do they work? All that good stuff. Alright. So one other thing about Google Gemini, the new app, it actually looks really nice. I love the new user interface.
Jordan Wilson [00:33:58]:
They did redesign it from the ground up, which I really like. So, some things if you are new to Google Gemini, they do have this new daily brief. I'm not gonna be able to bring that up because this is my personal Gmail account, because that's the only one where it's available, and my other workspace account ran out of limits from doing like nothing. So on the left hand side, this is where you can start a new chat. You can have the new daily brief if you do have that, images with VO, or sorry, images with nano banana videos now with the Omni model, your library, your Google gems, which are similar to GBTs. Right? Just smaller versions of Google, Gemini that you can create. You have your notebook with your notebook l m integration and then your recent chats. Alright.
Jordan Wilson [00:34:43]:
So in the main interface, I'm just gonna go ahead. I have a series of, prompts that I like to run. One thing. Alright. I'm gonna refresh this, see if this works. Okay. Interesting. Alright.
Jordan Wilson [00:34:55]:
One thing that's confusing to me is even the model selector is not consistent across, Gemini interfaces. So with my personal Gmail, Gemini account, which I have up, Right? My options are three, right, three five flash and then with thinking levels. Okay. Cool. Right. And then I go to my paid version of Gemini on my workspace plan, and I don't have that. All I have now is three five flash and three five thinking. So that's one thing I think the Google team needs to tackle.
Jordan Wilson [00:35:31]:
Right? Especially for people who are using, right, like myself. And I know there's millions of other people, I'm sure, that are using Google Gemini, on their personal Gmail account, in a work account. If you want to be able to replicate things, test things, run things through use cases as a front end user, which is really important as teams start moving to more front end surfaces. You have to be able to, let people understand, number one, what model they're using. Is there a knob or a control for reasoning effort? So at least me, I've been thinking. I've been reading about it. It's not marked anywhere. It's confusing.
Jordan Wilson [00:36:07]:
Alright. Anyways, so for this, I'm gonna be using the 3.5 flash, with, you know what? I'm gonna alright. I'm just gonna do what am I gonna do? Let's just go ahead and do extended. Alright. We'll see if this works. And right now, my usage is 0%. So I have 0% used. Alright.
Jordan Wilson [00:36:25]:
So my fun little rubric, it's about eating bananas and stuff. Alright. So I wanna see, the correct answer for this is five apples and three bananas. So I just said, I just woke up today with six apples and three bananas. Yesterday, I ate a banana and two apples. This morning, I will eat one apple and no bananas. However, I don't really like apples, and one banana may turn brown tomorrow. Assuming nothing else changes, how many apples and bananas will I have tonight? Alright.
Jordan Wilson [00:36:51]:
So, the good thing is, well, Google Gemini got this correct, so that's good to know. Alright. I did wanna show you this because one new feature that Google has that they didn't really talk about, which I like ish, I wish they went a little bit further, is you can finally see a little bit of the Gemini chain of thought. Not a lot. What is chain of thought? Well, that shows you how a model tackles a problem and what tools it calls. Right? So is it uploading? Is it accessing certain emails? Right? If I was connecting to my emails here, that would make sense. Or what emails is it looking at? What documents is it looking at? What websites is it going to? Is it using Python to write this code, or is it doing the math in its head? Right? So, as an example now, I can click these three little dot at the end, of the prompt answer, and it says more. And then I can see see thinking steps.
Jordan Wilson [00:37:40]:
So, unfortunately, it's, it's a bad chain of thought. You don't really get to see anything. But before Google Gemini had absolutely nothing. So sometimes when I really push, and add a lot of files and give it a complex prompt, I can see a little bit at least about some of the websites it goes to and some of the thinking steps, but, it is nowhere near what you get, with chat g p t and with Claude. I actually think right now, Claude has a slight advantage, in be in having a more transparent chain of thought that's easy to read. Chat GPTs is ultimately better, but it's a little more difficult from a user experience to get to it. So, yeah. So, Google, they kind of went for it, but it kinda didn't work as well.
Jordan Wilson [00:38:29]:
Alright. A couple little we'll do these fun little things. Alright. This is one that most, large language models get confused at. A man and his dog are standing on one side of the river. There's a boat with enough room for one human and one animal. How can a man get across with his dog in the fewest number of trips? Alright. Google, Gemini, 3.5 flash, got it right.
Jordan Wilson [00:38:49]:
One trip. Right? I don't know why. Or or maybe these things are finally in the training data, because I don't know. I talk about them a lot. I'm sure other people give similar prompts. Alright. Next one. If it takes three hours to dry 10 t shirts in the sun, how long will it take to dry 30 t shirts in the sun? It takes three hours.
Jordan Wilson [00:39:08]:
Got it correct. Alright. So so far, some of these trick little prompts, Gemini three flash doing a good job. If you have a single match and you walk into a room with an oil lamp, candle, and a fireplace, which do you light first? The match. That is correct. Alright. Next one. What color is an airplane's black box? Alright.
Jordan Wilson [00:39:31]:
The answer here is orange. Let's see how it does. Alright. Got it right. Bright orange. I'm gonna jump over to usage. FYI click my refresh. I swear.
Jordan Wilson [00:39:46]:
I swear y'all. It only says 1%, which is great, which is great. Before when I was doing these tests, it was really not letting me do hardly anything on my other account. Alright. Anyways, let's keep going. And then I'm gonna say, please give me seven jokes that end in the word blue. Two should be about animals. Three should be about some other topic in the body of this chat, and you can make up the other two.
Jordan Wilson [00:40:16]:
Alright. Let's go ahead. We'll give it a second. I'm gonna get my alright. Let's see if we got this right. Animal joke. What do you get when a sad puppy falls into a bucket of paint, a golden retriever that is completely blue? So no model has ever been able to make these actually funny, but I will say that this is coherent. It's ending in blue.
Jordan Wilson [00:40:38]:
I'm looking at the rest. It's taking some from the fruit topic. Both end in blue. Correct. Okay. And then the made up jokes all end in blue as well. We'll see if the made up ones are actually funny. Why did the jazz musicians stare at the ocean all day? He wanted to write a song that felt blue.
Jordan Wilson [00:40:56]:
Alright. So no jokes are actually funny, but they're at least coherent, and they follow the rules. Last but not least, in my silly little rubric, a box is locked with a three digit numerical code. All we know is that all digits are different. The sum of all digits is nine, and the digit in the middle is the highest. What is the code? This is a trick question because there's actually multiple, possibilities. So let's see here. There we go.
Jordan Wilson [00:41:24]:
Google Gemini 3.5 flash says there is actually more than one correct answer. Based on the clues provided, there are 14 possible codes. Alright. Let's see. Codes that start with a non zero digit. Alright. They all add up to nine. Middle number the highest.
Jordan Wilson [00:41:39]:
Yeah. So 180, that works. 270, that works. 360450. Yeah. So it did a good job. It got through that. So now just for fun, I'm gonna go ahead and, open canvas mode, and I'm gonna say, build all of my prompts and results into an interactive website using canvas mode.
Jordan Wilson [00:42:02]:
Alright. So, obviously, all of the old and previous features, of Google Gemini, obviously, still work with the new model. I haven't used canvas mode as much because I've noticed on my work plan, which is the one that I primarily use that's connected to all of my work data. My limits were disappearing like peanut M and M's, like upstairs. Right? They were just, vanishing. Alright. Anyways, we're gonna give that a second. We'll check back in on that interactive, on that interactive app there.
Jordan Wilson [00:42:39]:
And then we'll just do our last kind of test here, which is looking at the Gemini Omni. Alright. So, so let's go ahead. I'm gonna click the create video. And let me make sure. All I'm saying is explain a large language model, but using basketball terms. All right. So I do believe right now it is an eight second limit.
Jordan Wilson [00:43:07]:
I was trying a video to video, a video to video, example for whatever reason. This has been going for like two hours. So I don't think this one is going to work. I don't think that are, okay. Interesting. So I said, explain a large language model, but using basketball terms. And it said, I can only generate videos. Try a another prompt.
Jordan Wilson [00:43:33]:
Alright. So I'm gonna redo this even though. Alright. I'm gonna say, create a video that explains. Alright. How a large language model work, but using basketball terms. We'll do that and see if that works where the first one didn't. Alright.
Jordan Wilson [00:43:54]:
So let's jump in. Alright. There we go. We have our interactive, website version here, Don. Alright. So this is kinda cool. So it did, it created a nice little interactive website for the prompts that I gave it with the correct answer, as well as a an explanation. Right? Which is pretty cool.
Jordan Wilson [00:44:23]:
Right. So the river crossing, you know, it says on the left bank, there's a man and his dog with two emojis of each. The boat, it clearly shows that there's two seats in the boat, and then it kind of shows that you can, this is wait. This is actually pretty cool. I can click on them. Alright. That's kinda cool. Load the emojis into the boat and then click the sail across.
Jordan Wilson [00:44:45]:
Alright. And then it goes the other side. So, you know, I I've I've never been, shy about how much I love Google Gemini canvas mode. It is one of my, favorite and most used, kind of little AI modes, cross. I love the Gemini mode. It's great with front end. It's fast, and it usually works way better than artifacts or even the canvas mode inside of chat g p t. Alright.
Jordan Wilson [00:45:11]:
So pretty cool there. Let's check on our video generation. Alright. Here we go. We have, let me see. I actually have to, so we can get some audio here. I have to stop this screen share, and then I have to share the screen only using tabs. So let's go ahead and do that.
Jordan Wilson [00:45:30]:
So maybe, we can all hear. Okay. I mean oh, okay. I can share system audio. No. I can't. Alright. This is I had to do this in my other plan.
Jordan Wilson [00:45:44]:
Alright. So I'm gonna have to download this video, and then I'm gonna have to open it in a different window. This is why I love doing this stuff live, y'all. I love doing this stuff live. Alright. Let's try this again. Alright. And then let's drop in the video I just downloaded explaining.
Jordan Wilson [00:46:07]:
Alright. Let's try this here. Let's go Chrome tab. Thanks for your patience, y'all. Here we go. Chrome tab. Alright. We'll see, how a large language model works in basketball terms.
Jordan Wilson [00:46:27]:
From a visual perspective, we have a basketball player clearly, you know, it's supposed to be the the Bulls, but no Bulls. So, you know, Chicago Bulls colors, number 23, someone in an arena. I don't know if it's actually, the United Center because funny enough, there's, kind of Canadian banners in the background. But the physics, everything looks good. Court looks good. So let's go ahead and listen to the ten seconds. Can they explain large language models in ten seconds with basketball references?
Speaker C [00:47:01]:
Imagine a massive scouting report. An LLM reads every play ever made. It practices billions of shots to predict the next bucket. It is just ultimate teamwork.
Jordan Wilson [00:47:12]:
Alright. So nothing terribly inaccurate there. Alright. Some of the, physics were were fine. Right? He kind of shot a basketball, but, you know, at least you can kinda see the ball kinda goes through the rim there and at least where he's shooting it. It's kind of a close-up. He in unless he's, you know, Victor Wambiana or whatever his name is. His shooting arc is not going to be that close, to the rim unless he's standing on his teammate's shoulder.
Jordan Wilson [00:47:43]:
But, I mean, overall, decent ten second explanation using basketball terms. Right? Alright. So this one went a little long. Apologies about that, but here's the big takeaway. I think there's pros and cons here, to Gemini 3.5 flash, to Gemini Omni flash, and to anti gravity two point o. One of the underlying downsides is, well, the new usage limits. And although in my demo today, they didn't really creep up, thankfully, that would have been bad if I, you know, started something in anti gravity, and then I couldn't have done any other demos even in Google Gemini. But a lot of people like myself have been complaining that even on the $20 a month plan, it's, again, not as bad as anthropic clause limits.
Jordan Wilson [00:48:25]:
But, I mean, I think Google Gemini used to be a a buffet. Right? It was an all you could eat, from the AI compute factory and not so much anymore. So that's a downside. The other big downside with Gemini 3.5 flash, if you are using it on the API side to build, it's much more expensive. Right? It is one of the most token inefficient models there is. So although you can look at it one way and say, oh, wow. It's it's very smart for a flash model, and it's very, very fast. If you're using it, if you're just trying to swap out, you know, maybe you're using Gemini flash, Gemini three flash or Gemini 3.1 flashlight, and you're saying, oh, we can just swap it out for Gemini 3.5 flash.
Jordan Wilson [00:49:06]:
It's gonna cost you a lot even though it is very fast in a very smart model. It doesn't really, you know, keep to the original name. So, obviously, Google's, you you know, what they're gonna market Flash as is gonna change, you know, as they see, user demand change. But, again, very capable model. It's good. It's fast, but more expensive, for the most part. And we'll see what happens with Gemini 3.5 Pro. I would expect a much better, much more capable model, especially since the fanfare and reaction to a lot of what we saw from Google IO has been mediocre at best.
Jordan Wilson [00:49:42]:
Google Omni, I think people are misjudging it. Yeah. You saw my one little example there. My video to video example didn't work for this, but I think, ultimately, people are only looking at and comparing it to things like seed dancing. They're saying, hey. Text input, video output. That's not what Omni is. I think this really unlocks the future of creativity.
Jordan Wilson [00:50:01]:
Once Google is able to serve this up to more people, maybe extend, you you know, the time. I think this really changes the future of storytelling, being able to, you know, shoot a quick video and then change anything about it instantly in a consumer friendly interface is a big step change in creativity and what's possible. Last but not least, at least on the three things we went over today, anti gravity two point o. Not that good. Not gonna lie. Right? Yes. It can create nice front ends. It is fast ish.
Jordan Wilson [00:50:33]:
But when you're comparing it to what else is out there, I'm not gonna use it. Right? Even I think even if I had unlimited usage of Gemini anti gravity two point o desktop app, I'm not gonna use it, because it lacks what you get in cloud code, cloud co work, and codex by far. Hopefully, the team is gonna continue to invest in it. It does look like anti gravity is a bigger priority than I would have thought. At least when you would look at the end product and you're like, okay. This actually got a lot of shine and a lot of play at the IO, event. And then you look at what was initially released. There's been a lot of updates even in the past one week since it was initially released.
Jordan Wilson [00:51:10]:
But when you look at what was released, you're like, they maybe should have, you know, slid this one out as a footnote or maybe, you know, waited, to release something like this. I don't think it's there yet. It's not in the, you know, first or second tier in terms of, you know, Claude code, codex, cursor. You know, there's a handful of other third party ones that I think are really good and very far ahead of what we have in anti gravity. So I hope this was helpful. A longer version of putting AI to work at Wednesdays, that's what we get for trying to tackle three things live and run into a couple hiccups along the way. I hope this was helpful. If so, make sure you do check out the video version at youreverydayai.com.
Jordan Wilson [00:51:50]:
While you're there, make sure to sign up for our free daily newsletter. And thank you for tuning in. Hope to see you back tomorrow and everyday for more everyday AI. Thanks, y'all.
Midroll [00:52:00]:
And that's a wrap for today's edition of everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit your everydayai.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers, and we'll see you next time.
