Ep 751: Hands on with Google’s Gemma 4: How to Use The Open Source Model Locally and Why It Matters

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


Google's Gemma 4: Local AI Deployment for Businesses

The conversation around OpenAI and proprietary AI services has recently shifted as Google’s Gemma 4 model emerges as a leading open model, offering 2025-level AI capability directly on local devices. Gemma 4’s release under the Apache 2.0 license makes it entirely free for commercial use—with no usage restrictions—a development that impacts AI project costs, deployment strategy, and enterprise data privacy.

Open Source AI Model Performance: How Gemma 4 Ranks

Gemma 4 is not simply another open language model. The 31-billion-parameter dense variant equals or exceeds models as much as twenty times larger in performance. It currently ranks third globally on Arena’s open model leaderboard, scoring an Elo of 1452—outperforming models with parameter counts in the hundreds of billions, a milestone not previously seen in open source AI. The model supports complex reasoning, multi-step logic, coding, and image understanding, with a 256k token context window for large document and code analysis.

Cost Reduction with Local AI Deployment

Previously, leveraging top AI models meant budget allocations for per-seat subscriptions or costly API usage—which could escalate to thousands or millions in enterprise settings. Gemma 4 removes these recurring charges by enabling high-performance inference directly on consumer-grade hardware, with no API calls or subscription fees. This opens feasibility for always-on AI agents, background automation, and internal tools at a fixed, upfront hardware cost.

For instance, the latest mid-tier MacBook Pro can efficiently run the quantized 26B variant of Gemma 4—offering general-purpose performance equivalent to leading models from just over a year ago, which were only available from the cloud at premium rates. Higher-tier hardware can accommodate the full 31B dense model for cutting-edge tasks.

Privacy and Compliance Advantages of Local AI

Running large language models locally ensures all data remains on-premises, eliminating concerns over sensitive data transfer, cloud exposure, or relying on vendor privacy assurances. For industries such as healthcare and legal, Gemma 4 means regulatory compliance can be maintained while gaining advanced AI capability. The full commercial license allows any legal business use, including creating and selling AI-augmented tools or workflows—without legal review cycles for restrictive terms.

Practical Deployment Paths for Gemma 4

Business users do not require deep technical expertise to operationalize Gemma 4. Usable desktop clients such as Ollama and LM Studio provide an interface nearly identical to ChatGPT or Gemini, but run entirely on the local device. Installation involves a simple desktop download, followed by one-click model retrieval for the variant that matches workstation specs—eliminating developer-level setup. For edge needs, the compact e2b and e4b versions can run on mobile devices or Raspberry Pi hardware, with Google’s AI Edge Gallery app supporting on-device inference without connectivity.

Gemma 4 can be integrated as the local engine within agent frameworks or AI-powered business process tools, simply by pointing existing agent software at the downloaded weights or connecting through supported local APIs.

Real-World Model Capabilities and Evaluation

Head-to-head tests underscore Gemma 4’s strengths in logic, instruction following, and contextually relevant creativity. For example, when challenged with a variety of classic logic riddles and custom instruction-following tasks, Gemma 4 not only outperformed prior open and proprietary models on accuracy but also demonstrated advanced structure and rationale generation. For contextualized content, Gemma 4 was able to process uploaded transcripts and output on-brand, audience-ready newsletter drafts within seconds—delivering results that closely matched defined tone, structure, and formatting examples.

Shifting Competitive Dynamics for AI in the Enterprise

With the introduction of Gemma 4, vendor dependency in generative AI is no longer a given. Enterprises can now deploy AI workflows with complete privacy, cost control, and code-level customization. The performance-to-size ratio of Gemma 4 indicates that new open models are closing the gap with (or even exceeding) cloud offerings—at a fraction of hardware and operational cost.

The accessibility of edge variants also empowers organizations to pilot and launch AI-powered mobile apps that function offline, with no licensing or usage tracking concerns.

Strategic Considerations

Business and technical leadership should actively assess the transformative impact of models like Gemma 4, even within organizations currently invested in enterprise SaaS AI. Running advanced AI locally can be not only a backup for uptime and cost mitigation, but also a primary deployment model for internal tools, confidential automation, and new AI product offerings. With the commercial-friendly license and robust community support, the window for proprietary advantage is narrowing: the cost, privacy, and innovation calculus for enterprise AI strategy is rapidly changing.

Emerging toolchains such as Gemma 4 are making high-caliber AI accessible, manageable, and free for the enterprise. Keeping pace will mean not just evaluating cloud AI options, but integrating local and edge AI as a cost-effective, private cornerstone of business AI deployments.


Topics Covered in This Episode:

  1. Google Gemma 4 Open Source Launch
  2. Gemma 4's Apache 2.0 Licensing Explained
  3. Gemma 4 Model Variants & Hardware Requirements
  4. Small Language Models vs. Large Model Performance
  5. Benchmarking Gemma 4 Against Top AI Models
  6. Local AI Model Deployment Benefits & Privacy
  7. Hands-on Guide: Running Gemma 4 Locally
  8. Live Performance Test: Coding, Reasoning & Logic
  9. Instruction Following and Creative Output Demo
  10. Future Impact: Open Source AI for Businesses




Episode Transcript 



Jordan Wilson [00:00:17]:
If I would have told you a year ago that you could use the world's most powerful models on your local machine without having to pay for it, you probably would have looked at me and said, you're absolutely crazy. Well, I'm not because that day is here. Thanks to Google's new impressive Gemma four model. You can get at least twenty twenty fives frontier AI performance on your local machine, running it privately offline with this new impressive open source model. And what I want you to think back to a year ish ago, I'm not just telling you to okay. Think about maybe replacing that $20 a month subscription. That's not where this is going. What about if your company was spending thousands of dollars or millions of dollars on AI deployments internally or externally? Or when you think about running AI agents around the clock.

Jordan Wilson [00:01:11]:
Right? Anthropic recently said, hey. You can't use your Claude subscription anymore for OpenClaw. Well, now you can run with Google Gemma four. You can run it around the clock and not pay a penny. No. This is not too good to be true. Yes. It kinda feels like we're living in the future, and we're gonna go over it all live.

Jordan Wilson [00:01:30]:
So welcome to everyday AI. So, here's the big picture. What's happening with Google's new Gemma four, and it's, I think, rivaling some of its, trillion parameter giant competitors. So Google DeepMind released Gemma four, its most capable open model family to date. So there's four different variations. We're gonna be breaking them all down on today's show. But the big boy, the 31,000,000,000 parameter model outranks models 20 times its size. It is third on the global ranking of all open models on arena, and that's not even the best thing.

Jordan Wilson [00:02:08]:
The best thing, maybe, is that Google changed its licensing. And now at Gemma four is released under the very permissive, Apache two point o license, granting full commercial freedom with essentially no restrictions. So not only is this free, you can run it on your computer around the clock if you have the right hardware. I'm gonna be breaking all of that down. But you can also create and sell things with this model. So on today's show, here's what we're gonna break down. I'm gonna break down for you how a 31,000,000,000 parameter model now competes with one's 20 times its size. I'm gonna show you and tell you why running AI locally on your laptop or even phone changes the cost and privacy equation, and I'm gonna break down live, right, the exact tools and steps to download and run Gemma for for free today as we go hands on.

Jordan Wilson [00:03:03]:
Alright. Let's get into it. This is Everyday AI. Welcome. What's going on? My name is Jordan Wilson. If you're new here, well, we do this every day. And this is your daily livestream podcast and free daily newsletter, helping business leaders like you and me keep up with the ever changing AI landscape. I help you understand what's important and show you practically how to use that information to grow your company and career.

Jordan Wilson [00:03:23]:
So it starts here, but make sure you go to our website at youreverydayai.com. Decide up for the free daily newsletter. We're gonna be recapping today's show as well as giving you all of the other AI information you need to be the smartest person in AI at your company. So it's Wednesday. Wednesday, right, we have kinda different shows. On Mondays, we give you the AI news. On Wednesdays, we go hands on a deep dive, with one new AI release, something like that. And then on Friday, we go over in our Friday features, you know, kind of five to seven new AI features.

Jordan Wilson [00:03:55]:
Tuesday, Thursday, we rotate it a little bit. So if you are new here, that's the plan. But on Wednesdays, we get our hands dirty doing live demos of AI. But first, before we do that, I wanna talk about how Gemma four, how I think is going to completely change the landscape. And the cool thing about having a daily AI podcast where you can go listen to all the episodes for free, you can go see. I've been ranting and raving about the power of small language models since 2023. And, technically, this is a small language model, but you get large language model performance. Right? So the exact, definitions of what's a small language model, what's a large language model is ever changing.

Jordan Wilson [00:04:32]:
Right? But for the most part, you look at the number of parameters in a model. So we're gonna be comparing, you know, what you can get for free out of Gemma four with what you could get out of the Frontier models about, you know, fourteen months ago, which at the time were GPT four o, and Claude at SONNET three seven. But those are models. Right? GPT four o was reportedly 2,000,000,000,000 parameters. Right? So here, you have a model that's about an eighth of the size, and it's open source, and it's free. And here's why I think it's gonna change the landscape. Well, number one, I already talked about kind of this anthropic saying, hey. You can't use your Claude subscription anymore to run OpenClaw.

Jordan Wilson [00:05:13]:
You have to pay via API, and people are like, that's gonna be crazy expensive. Well, now you can run a Gentic AI around the clock for free with GEMMA four. Is it gonna give you the same model as an Opus five four or a GPT five four? Absolutely not. But there's a good chance for, you know, maybe 50 to 80% of what you're trying to do. This model is going to be good enough. And Google also changed the Gemma license to Apache two point o, which provides users unrestricted commercial freedom, and preventing corporate vendor dependency. That's the big thing here. I think that this is going to lead to a, resurgence, and I've referenced this on the show once or twice, but I think we're gonna kind of have this future that's kind of retro.

Jordan Wilson [00:06:01]:
I think that desktop software is going to come back in the same way that, you know, in the nineties, we saw this wave of personal computing. I think we're gonna see personal software. Right? So not necessarily software that's for your whole company, software that's for you. Right? And I think that it's gonna be models like JEMA four, right, that are gonna allow this to happen. And also the performance versus size ratio absolutely just reset. Alright. And I'm gonna break down what this means, but essentially think of it like this. Right? If you follow, I don't know, boxing or UFC, I don't really follow those things, but there's something called, like, you know, pound for pound.

Jordan Wilson [00:06:39]:
This is the pound for pound best fighter in the world. Right? If if someone, I don't know, fighting at a 150 pounds can knock out someone at a 180, That's pretty impressive. Right? This GEMMA is punching well above its weight class. I'm talking about it is competing with models 20 times its size on the open source side. This is something we've literally, quite literally, have never seen in the history of AI, which is why I think Gemma four is a huge deal. So even if you don't necessarily think that you or your company need to use this, you're like, okay. Well, we pay for JET GPT Enterprise or Google Gemini Enterprise, Claude Enterprise, whatever. Right? And we have more, robust agentic solutions already going.

Jordan Wilson [00:07:22]:
Okay. You still need to be learning Gemma four and building with it, not just as a backup dependency, but because it can run-in in probably in the future. Right? If we fast forward one more year, I don't know if any open source competitors are going to be able to truly catch up to what we just saw from Google and their JEMMA four. So let's talk quickly about the capabilities. So it can solve complex reasoning, math, and multi step logic problems effectively. There's native support for function calling, yes, in a local model and structured JSON outputs for agentic workflows. It has a context window depending on, right, your hardware. We'll give you all those specs in the newsletter.

Jordan Wilson [00:08:04]:
Right? But, if you have capable hardware, you can work with a 256, k token, window for analyzing large document and code bases. It can analyze text images and videos natively, but excludes audio support for the bigger models, but the smaller models actually support audio. And it can generate incorrect code efficiently as a local offline coding assistant. Here are the four different flavors of Gemma four. And, oh, FYI, as I take a sip of my coffee, yeah, this is unedited, unscripted. I hope that this is gonna be interesting for you. But if not, and if you listen to the podcast all the time, make sure you sign up for our newsletter because I put a poll in our, newsletter on Monday. I said, hey.

Jordan Wilson [00:08:49]:
What do you guys wanna see hands on on Wednesday? And you all voted Gemma four. So if you wanna see other types of, demos on Wednesday, make sure you read our newsletter. I'll usually put out a poll maybe Monday or the Friday before depending on how busy things are. So I'm doing this for you. This is what you wanted, FYI. But, I mean, I'm doing it for myself because I'd be doing the same thing anyways. But right now, there are four different, variants of JEMMA four. Easy.

Jordan Wilson [00:09:15]:
Right? But you have the e two b and the e four b. Those are essentially phone models. This is as edge as edge gets because it can actually even, run on a Raspberry Pi, right, and your basic, phones. And this is big. Right? Especially if you've always, I don't know, wanted to build a certain app for something, right, and you're confused how or, or right. Using the JEMMA e two b and e four b models can get you there pretty quickly. They are extremely capable. Alright.

Jordan Wilson [00:09:46]:
Then you have the two bigger boy models that you're gonna need, well, consumer hardware. That's that's the reality here. Right? Because to get this level of performance previously before, you know, rewind more than a week ago, you couldn't on a $2,000 laptop, you couldn't run anything. Right? That was a a top 10 open source model. Now you can. And I think a good way to look at this is the MacBook Pro test. Right? So generally, Apple, right, they usually have about three to four different versions of their MacBook Pro. So obviously, the souped up ones are, you know, a little expensive.

Jordan Wilson [00:10:26]:
But I say if you take the middle, right, variety or the middle flavor of a MacBook Pro off the shelf. Right? Walk into Best Buy, Apple Store, whatever, look at the MacBook Pros and say, give me the one in the middle. Now that one in the middle can technically run the 26 b version of JEMMA four. And that's because it uses the mixture of experts framework, and it only activates 4,000,000,000 parameters, and it's really fast. Right? So that model, the 26 b is actually faster, less capable than the 31 b. But by default, you can run that. You you know, you could technically run the quantized version on a 16 gigabyte MacBook, the base baseline. But if you just go for the middle flavor, of a MacBook Pro, which I'm trying to see here, you know, what this cost is.

Jordan Wilson [00:11:14]:
I have it open, on my other tab here. Let's see. Okay. No no trade in. No thanks. I don't want all this extra stuff. Let's see how much this is. Alright.

Jordan Wilson [00:11:24]:
So, $202,200. Right? Which any MacBook Pro I've bought over the last ten years have has been that price or way more. Right? So I don't think people understand the middle, you know, middle MacBook Pro that you buy off the shelf can now run a model that's about the same capabilities, as the best models in the world fourteen months ago. Alright. And then you have the last flavor. Okay? So e two, e two b, e four b, for the, phones and edge devices. Then you have the 26 b, and then you have the 31 b dense model. Alright.

Jordan Wilson [00:12:04]:
So when you're running this, you're running the entire thing. That's why it's dense. It's not the 26 b mixture of experts. Alright. And that delivers the highest quality reasoning and coding output. And you will need a more powerful, computer. But, you know, luckily for me, you know, I got a Mac Studio, a fairly capable one, and it runs great on my machine. But, to run this one, you will either need the most souped out, you know, MacBook Pro or Windows equivalent of that, or you will need a, you know, fairly capable Mac Studio or an NVIDIA g t, DGX, something like that.

Jordan Wilson [00:12:39]:
But you quantized you could squeeze this one on, something with about 32 gigabytes of RAM. It will run better at about 48. So, you know, as an example, my Mac Studio has 64 gigabytes of RAM. Alright. So now you understand the technical side. Like I said, if you have a newer middle of the line MacBook Pro, just an easy way to benchmark it. You're gonna be able to run the quantized version of the 26 b. You're gonna have to have a little bit more of a powerful machine to run the 31 b dense model, but it's capable, and I'm gonna show you here live.

Jordan Wilson [00:13:09]:
So here's why this is important. If you look at the biggest and best open source models in the world, right, like Kimmy k two five thinking, well, JEMMA four thirty one b is now in the exact same category, but at a fraction of the cost and a fraction of the parameter count to run it. And community testers have confirmed strong results in coding, reasoning, and image understanding. And the 31 b scored a fourteen fifty two on arena. Right? So the, arena AI and formerly the LM arena, this is blind taste testing. You put a prompt in, It kicks out, you know, two different, outputs from different models. You score which one's better. That's how you get an Elo score.

Jordan Wilson [00:13:50]:
Right? And fifteen months ago, the best Elo in the world was not fourteen fifty. Right? And that's what's crazy. So now this is scoring better at least, on an Elo score and all the scientific benchmarks as the Frontier AI models from fifteen months ago. So here's, I I I have a little chart here on my screen for our livestream audience, podcast audience. You know, this one's not gonna be super visual, but you can always go to our website at youreverydayai.com. Click episodes, and you can go watch today's show if you wanna see the video version of this, but it should be fairly straightforward. I'm not gonna be doing anything too visual. But we do have this from Google's announcement blog post that shows the model performance versus size.

Jordan Wilson [00:14:32]:
And you'll see this is literally the new Gemma four b is uncharted territory, because this is charting Elo score, on one access and then the total model size on the other. So previously, to get anything like a fourteen fifty on an open source model. Right? And we don't usually know, the size of proprietary models. Right? So your, you know, your Gemini three one pro, your, you know, Claude, Opus four six, your GPT five four, you know, etcetera. But, presumably, they're, you know, multiple trillions, you know, maybe one and a half to two and a half trillion parameters. Right? So think of this as like a a hard drive size, right, if you wanna simplify it. So to get this same level of performance, a fourteen fifty ish, score, from an open source model, you're looking at, something that's about 300, 400,000,000,000 parameters in size. So, again, this is about, 10% of that size.

Jordan Wilson [00:15:38]:
Some of them, right, like, Kimmy k, k two five thinking, which is a lot of people's most favorite open source model. This is like a twentieth of the size. Right? With the same, roughly the same level of performance, at least when it comes to Elo and most scientific benchmarks. So now let's just quickly talk about, well, why would you even wanna run anything locally? Like, okay, Jordan. What's the point? I don't care. $20 a month isn't a big deal. Sure. Right? But think about doing this at scale.

Jordan Wilson [00:16:04]:
Think about agentic AI. Right? Things like OpenClaw. Right? And now that, Anthropic has shifted away, a lot of people are having to now use some of these open source models, but using them via open router so it's not completely free because you can't run. You literally can't run models like quen three five, g O M five, kimmy k two thinking on anything less than a, you know, $8,000 computer. Definitely on, you know, not the, you know, your average, you know, MacBook Pro. Right? So that's the difference. This can't really run on true consumer, prior open source models that might be able to run, you know, your agentic, tools like OpenClaw or other agents. You couldn't run it on consumer models.

Jordan Wilson [00:16:48]:
You have you had to literally have, like, an 8,000, $10,000 computer or more. Right? And that's where the local aspect comes in handy. So being able to run things locally, keeping costs down, no subscription fees, no API keys, no loop no usage limits after you download something. But also, talk about privacy. You never have to send anything to a to a cloud. Right? Because, yes, keep this in mind. As long as you are on a, paid team plan with any of the big four, you know, and as long as you turn off model training or turn on the basic privacy settings, you're not sending technically, any of that private information, to these companies, or they can't really do anything with it. Right? However, I do understand with highly sensitive documents, how you might not wanna send that even if you've turned off model training.

Jordan Wilson [00:17:40]:
Right? So any sensitive industries like health care and legal can gain very capable AI without any cloud exposure. And then like I talked about, right, the combination of well, it's free. It can run twenty four seven. It's sensitive. It runs it all on your machine. You can literally turn the Internet off and use this. But then also with the new licensing, you can have full commercial use. So you can literally use this for anything.

Jordan Wilson [00:18:03]:
And there's three ways to run JEMMA for today, and then I'm gonna show you how to do this live. Thanks for sticking with me. I wanted to, you know, first kind of tell you how important this is and kind of set the context here. But there's, kind of three different ways, all for free, that you can run at JEMMA for. One would be a tool like, Ollama. Right? That's the one I'm gonna show you. This is essentially it gives local models a graphical user interface like chat g p t. Very simple.

Jordan Wilson [00:18:32]:
Right? So you can then, run a terminal command and download the models in minutes, or you can even run that command in the Ollama interface. Also, LM studio is a great one. Same thing. Offers you a visual chat interface similar to chat gbt for non developers. I guess another way so, hey, we'll just do four free ways to run it. You know, if any agentic system that you're running locally, you can point that, point that system to Gemma four, on hugging face or if it does a lot of the, you know, local agents that run-in your computer, they can run via, you know, Ollama as well. So you can run-in Ollama command, hugging face command, you know, point your agents to the download because that's the thing. You essentially download this thing and you run it.

Jordan Wilson [00:19:16]:
Also, for the, the other versions. Right? So the two that we're gonna be looking at or the one are gonna be the bigger variants. Right? But to run the local ones, people don't know this. Google actually has a great, app called Google AI Edge Gallery. You can download that for iOS or for Android, but this is huge. If you haven't known this already, you probably should do this because what that allows you to do is the equivalent of running it offline on your computer. Right? This is an app where then you can download, the the smaller mobile edge versions and then, hey, if you're ever in trouble or if you're ever somewhere where you just don't have service, you at least have a highly capable large language model on your phone that you can use at any time. Alright.

Jordan Wilson [00:20:00]:
So let's get going live ish. First, I'm not gonna download this line because it might take a while. Right? And, trying to download a large, a larger file like this while also streaming live doesn't always work. So here's what I did. I'm gonna use Olama in this case. So here's what you're gonna do. You're gonna go to olama, alright, .com. So that's ollama.com.

Jordan Wilson [00:20:27]:
If you haven't already, you're going to download the program. Okay? This is a simple desktop client, like I said. It just allows you to use any open source, open weight model on your computer, but it gives it a graphical interface. So in the same way that you would chat with chatgbt.com, gemini.com, claw.ai, etcetera, this allows you to work with open models in that interface because by default, you'd be interacting with them via command line tools or the terminal, which is not always ideal for nontechnical users. So go to olama.com, download that, install olama on your local machine. Alright. That's step one. Step two, you're going to download the actual model.

Jordan Wilson [00:21:05]:
So you're gonna search you can search models on Ollama. Just type in Gemma four, and then you're gonna choose the variant that your local machine can run. So for most people, that's gonna be the 26 b version. For me, I'm gonna be showing you the 31 b version. Alright. So all you have to do is once you bring that model up on the Olama website, there's gonna be a it says CLI. Right, but there's a little command. You're just gonna copy that.

Jordan Wilson [00:21:30]:
Right? So this one says, Ollama run gemma four thirty one b. Right? All you're gonna do is copy that. Then you're going to open Ollama. Right? And again, very simple. All you're gonna do is then paste in that command. Alright? And it's going to download the model. So, for me, you know, this model was about nine gigabytes. It took, I don't know, five ish minutes, to download, and then that's it.

Jordan Wilson [00:21:55]:
And then you're ready to run with it. So, let's go live. So here's what we are going to do. Live stream audience. Do me a favor. Let me know if you can see my screen. Alright. So I'm gonna be jumping around a little bit here because I'm gonna be having some, some copy and paste prompts.

Jordan Wilson [00:22:11]:
So, about fifteen months ago, I did a show comparing the latest version of Sonnet, which I believe was three seven, to g p t four l. So again, going back to how I started this show, this was about, fourteen, fifteen months ago. These were the best, general use case models in the world. And I had a series of prompts. I kinda had, like, a very, you know, unofficial fun rubric, that I would do, comparing models. And I'm gonna go ahead, run the exact same prompts. Right? So we're going hands on here. So I'm gonna first put in a message to Gemma, and this is exactly what I did previously just to kind of, level the playing field.

Jordan Wilson [00:22:56]:
So all I'm saying is for this chat, please respond with proper formatting and structured bullet points. Do not waste words. Answer in the shortest way possible while still being detailed enough to fill in the user answer request. Right? Right? So this is what I did for all the other ones. This is what I'm doing it for now. Alright. So here is our rubric. So, test one, this is just a trick question.

Jordan Wilson [00:23:18]:
It's logic. Alright. And when I did this, both Claude and, Claude three seven SONNET in GPT four o got it wrong. We'll see if, Gemma four gets it correct. Alright. Hey. Love, love, love when we get little bugs. Alright.

Jordan Wilson [00:23:38]:
I'm gonna have to run that again. It essentially went through the thinking. Right? That's the thing. This model thinks and it reasons as well. And for those watching it live, you can just see that it did this. So we'll see if this got it correctly. Right? The correct answer, should well, I should probably read it. I said, I just woke up today with six apples and three three bananas.

Jordan Wilson [00:24:00]:
Yeah. Live stream audience or podcast audience. Try to do this live. See if you can get it. I just woke up today with six apples and three bananas. Yesterday, I ate a banana and two apples. This morning, I will eat one apple and no bananas. However, I don't really like apples, and one banana may turn brown tomorrow.

Jordan Wilson [00:24:17]:
Assuming nothing else changes, how many apples and bananas will I have tonight? So a little trick question. GPT4O and Claude 37Sonnet got this wrong. Alright. Let's see. Alright. So it looks like also they both got it wrong. Alright. So the correct answer is five apples and three bananas.

Jordan Wilson [00:24:39]:
Alright. So, Gemma four close got five apples and two bananas, which technically not that we need to gauge, you you know, the level of correct versus, other models. Right. Claude saw it said three apples, two bananas. GBD four o said three apples, two bananas. So they all got them wrong. Gemma got it a little closer. But, hey, did you get this right? Alright.

Jordan Wilson [00:25:01]:
Our next one, the old man and dog crossing a river. Right? So this also shows that the model is thinking. Right? So, if you're listening on the podcast, you're probably not seeing this, but it's also showing its thinking trace. Alright. So the next one, I'm saying a man and his dog are standing on one side of the river. There's a boat with enough room for one human and one animal. How can a man get across with his dog in the fewest number of trips? What's so funny is I did all of these, all of these beforehand just because I want wanted to make sure that they would work. The first time, alright, it got this right.

Jordan Wilson [00:25:36]:
And now the second time, it just got this wrong, which is funny. But you can always go back and look and look at how it thought. So same thing. Claude three 7 sought it, got it wrong. GBT4 o got it wrong. They both said three trips. The first first time I ran this, Gemma four got it right. This one, I'm doing it live here.

Jordan Wilson [00:25:56]:
It got it wrong. And then just for fun, I reran it again, and it still got it wrong. It said two trips. Right? Interesting. That's the thing with large language models, they're generative. That's why it makes these, live demos always always fun. Right? Because doing it before offline, I'm like, okay. Cool.

Jordan Wilson [00:26:10]:
It looks like Gemma four is gonna perform, much better, but it's, again, getting it wrong, as did the best models in the world fifteen years ago, but it's getting it a little less wrong at least for the first two times. Alright. Let's try the next one. Alright. So our next prompt here, we're saying, it takes three hours to dry 10 T shirts in the sun. How long will it take to dry 30 T shirts in the sun? The correct answer is three hours. Alright? And for, reference, a year ish ago, Claude and, GBT got that correct. Alright.

Jordan Wilson [00:26:47]:
It is three hours. The time doesn't change. Right? Assuming, and it did say drying principle. The time required to dry laundry is determined by external factors, sun intensity humidity, not by the total quantity of items provided adequate space exists. So Gemini Gemini four, sorry, Gemma four not only got it correct, but it did provide a, some nice rationale as it thought through the problem. Alright. The next one. And this already answered.

Jordan Wilson [00:27:13]:
That was so quick. Right? Again, this is running all locally, and that was probably faster than I would have even gotten from, proprietary models online. Right? So I said, if you have a single match I see your modest. Did you see how fast that was? My gosh. So I said, if you have a single match and you walk into a room with an oil, lamp, a candle, and a fireplace, which do you light first? Again, these are just fun trick questions. The correct answer is the match. Alright. So it got that right.

Jordan Wilson [00:27:42]:
Claude and g b t four o also got that right. Alright. Our next one, what color is an airplane's black box? Alright. It's taking a second to think. Bright orange. Got that correct. Good. As well as the others.

Jordan Wilson [00:27:54]:
Got that correct. Alright. Here's one. We'll see if, this is actually correct. Because last time, Claude saw it and Jeep d four o failed on this one. So I said, please give me seven jokes that end in the word blue. Two should be about animals. Three should be about some other topic in the body of this chat.

Jordan Wilson [00:28:15]:
That's important. Right? Although, in fairness, to the JEMMA models, that technically has a little bit more to do, with the harnessing of Ollama in this case. Right. So not exactly an apples to apples comparison, just FYI. Alright. So, I said three should be about two should be about animals. Three should be about some other topic in the body of this chat, and you should make up the other two. So first, I'm gonna see, did it get the correct number? Yes.

Jordan Wilson [00:28:44]:
It gave me two animals, three about chat topics, and two original made up. So so far so good. Next, do they all end in the word blue? Blue blue blue blue blue blue blue. Yes. Alright. So so far good. And then I'm gonna see, as long as they make sense. Right? These aren't always funny, but it at least has to be a joke to pass this rubric.

Jordan Wilson [00:29:07]:
Alright. So animal joke. Why did the monkey fall into the paint bucket? Because he wasn't used to something so vividly blue. Alright. Is that a joke? Sure. Is it funny? Absolutely not. Alright. Let's look at the chat topics, see if it got it right, pulled the context incorrectly.

Jordan Wilson [00:29:23]:
Why did the farmer throw away the apples? They were no longer crisp, just a sad brown blue. It's borderlining nonsensical. Right? Let's look at the last one. What or or the next one. Why could the larger mat predict the drying time because the sunlight was so strangely blue? So these are borderline nonsensical. I could say you could make the argument. They make sense. They're on they're on the edge here.

Jordan Wilson [00:29:48]:
Then let's look at the original made up jokes. Hopefully, these are a little bit better. Alright. Why did the geometry student bring a fishing pole? Because he was hoping to catch something entirely blue. Alright. So the jokes are trash, but they actually follow, the instructions. So when we're looking at instruction following, this technically passed even though the jokes were absolute garbage. But like I said, SONNET previously failed in GBT four zero field.

Jordan Wilson [00:30:14]:
Alright. Next one. Alright. This one is much trickier. So I wouldn't expect Gemma four to get this right. So I said a box is locked with a three digit numerical code. All we know is that all digits are different. The sum of all digits is nine, and the digit in the middle is the highest.

Jordan Wilson [00:30:32]:
What is the code? Alright. So this is a very, trick question because there are multiple valid answers. Alright. But, both Claude Sonnet got this wrong. GPT four o got this wrong. So what I'm looking for in a correct answer here, number one, that it even gives me at least one correct answer, but there's multiple correct answers. Right? Like, as an example, 180270351 would meet all those criteria. So, Claude in g p t four o got this wrong when we did the original testing.

Jordan Wilson [00:31:04]:
Claude's math didn't add up. G p t four o did not follow the rules. It had ones that added up to nine. Right? But as an example, it gave me one two six, but that didn't follow the rules because the middle digits two was not the highest. So let's see. It fought for twenty two seconds here. It kind of went through the the deduction process, and it did give me a solution here. So it technically is correct.

Jordan Wilson [00:31:32]:
Right? Whereas the other models did not even give me one correct code. Right? So quad seven, quad three seven saw it, it said 172. That does not add up to nine. That adds up to 10. Like I said, GPT4O, gave me 126, which does not follow the instructions because the middle digit was not the highest. So here, technically, Gemma got it right. It didn't get it fully right, but it was the only one that got it right. It said the code is 243.

Jordan Wilson [00:31:59]:
This was technically a triple trick question because I asked for a code, but technically, there are, multiple codes. So it technically answered, but I would have loved a a super correct answer where it said, you asked for one correct answer. Here's one. But there's actually more correct answers. But I will say that, at least now, Gemma got the last two right. Where the last did not get any of them right. Alright. This one, we're gonna go into some gray area here.

Jordan Wilson [00:32:26]:
Alright. I don't wanna make this, too long because it'll probably take another, five to ten minutes to go through the rubric. So I'm just gonna find some other questions that the others, maybe failed or just look into some gray area here talking about some creativity. So this one, I said, generate, generate unique and creative marketing advertising strategies to grow the everyday AI podcast. Do not suggest general run of the mill ideas. Only pitch clever advertising and marketing tactics to specifically grow the Everyday AI podcast. Alright. So, for reference, a year ago, Claude said, run AI teasers, virtual cohost challenge, listener q and a, augmented reality experience, GP four o said monthly puzzles, art contests, custom recommendations, guest AI cohosts.

Jordan Wilson [00:33:14]:
Alright. So let's see what Gemma four said. So, it said partnership and cross promotion strategies, which is good because that's, you know, basics of growing a podcast. Right? Which the others didn't come up with even though it's not super creative. Alright. So it says AI tool integration ads. It said partner with niche specific non major AI tools. It's a good idea.

Jordan Wilson [00:33:34]:
Industry vertical sponsorships. Then it said content hijacking, viral strategies, doing an AI Mythbusters challenge, interactive prop battles, then community engagement tactics. So the AI challenge hotline, I like that. It says dedicate a specific call in segment where listeners call with a real world mundane problem. Should we do that? Should we do that? Alright. If you think we should do that, also shout out because, someone from Microsoft did suggest this to me, like, two years ago. So shout out. I do remember Nisiani.

Jordan Wilson [00:34:05]:
You you said I should do that, and I was like, yeah. We should. Alright. So if you think we should do that, just say hotline. Right? Do a do a a a comment in the live stream or leave a comment on the Spotify. Just say hotline if you did that be fun. Maybe it will. Alright.

Jordan Wilson [00:34:18]:
Then it also said micro membership prompt bolt. Alright. So, this is good. I would say these are much more impressive. Yes. This one requires judgment on my part. It is great area. But, looking at what Claude three seven sonnet and GPT4O gave me, Gemma four, much, much better.

Jordan Wilson [00:34:37]:
Alright. Let me do one other that for sure, some failed on. Alright. So, uploading, uploading photos, might do that. Although I don't have the original, photo that I use. Let me see. All right. Let's just do one other one here.

Jordan Wilson [00:34:57]:
Okay. We're going to do a, uploading a transcript. I like that one. So let's go ahead and, let me find this file here. All right. So, I'm going to go ahead and put this prompt in it's a little bit longer and then I'm going to be uploading, two different files here. So I want to make sure that I get these, get these correct. Alright.

Jordan Wilson [00:35:27]:
There we go. I should go in my downloads folder. That that would help. Alright. So here's what we're gonna try, and this will probably be the last one. Alright. So I said for this chat, you will turn a podcast transcript of me, Jordan, the host of Everyday AI, talking about AI news. Turn it into a choppy and engaging newsletter copy.

Jordan Wilson [00:35:48]:
I've attached examples of previous newsletters and how they should be written as well as the most recent podcast transcript. So this is my podcast transcript from yesterday where we did a start here series about vibe coding. And then I said, please write a newsletter for the attached transcript mimicking, the style as closely as possible to the example given. So we'll see here, we're getting a little dot dot dot. So if I'm being honest, I don't know the last time that I uploaded, two different file formats. So I uploaded a PDF and an RTF, file here, inside Ollama. So, again, this one is not the fairest, comparison because, again, technically here, we're also relying on the, technically, the the harness, of Ollama and not just the model of Gemma four. Whereas before, you know, when we're testing this against g p d four o and quad three seven, we were using it.

Jordan Wilson [00:36:39]:
Okay. So I do know okay. It is working. Right. I'm like, okay. This should be able to Right? Old Nama is amazing. It should be able to handle, you know, multiple, kind of file types. So we are.

Jordan Wilson [00:36:50]:
This one is, taking the longest so far. This one is the first time that we're probably gonna have it be able to think a reason for more than a minute or two. And again, y'all think about this. Just the fact that you can have a local model that now reasons without paying a cent is crazy. Alright. So it also gave me a checklist of adherence, which is great because I didn't even ask it to do that. But that is something that if I was rewriting this prompt that I've been using this for, like, two years, I would have rewrote this. So it went through.

Jordan Wilson [00:37:21]:
It created a checklist based on what it found from the examples that I, that I uploaded. So as an example, right, I gave yesterday's transcript, and then I gave a 30 page document of older newsletters. So it went through. It examined those. It actually only took twenty seven seconds. It kind of picked out the tone style, the format, the context source, all these things, hook intro quality. Let's see. Alright.

Jordan Wilson [00:37:46]:
It actually did a pretty good job because I remember this was my intro of the podcast. Alright. So I'll I'll read it. Let's be real. You can tell an and if you read our newsletter, let me know if this sounds like it might be in our newsletter. Alright. Let's be real. You can tell an AI your wildest dream home and poof, a building appears in front of your eyes and minutes.

Jordan Wilson [00:38:03]:
It's exactly what you asked for. You move in. It's awesome. But then you you wanna hang a towel rack. You run-in you run into a wall, and you realize the entire thing is held together with duct tapes, hopes, and dreams. There wasn't a permit, and the foundation is shaky. You're in trouble. Alright.

Jordan Wilson [00:38:16]:
This actually did almost too closely to my actual intro from yesterday, but as I'm looking for this, as I'm looking at this, it actually did a pretty decent job of writing something in my tone, kind of this short, choppy style like I told it to, you know, an emoji in each headline, which is what we would normally do. It has an actionable try this section, which is something that we also do in the newsletter. So, although this is not, you know, the best ever, right, it actually did a pretty good job. From what I recall, it did a little bit better job at instruction following, than Claude three seven sonnet. I do think Claude three seven sonnet did a little bit with the tone of voice, but it did a better job tone of voice, or matching the tone of voice than it did than GBT four ope did. So, overall, when I look at the, you know, six or seven, kind of different, unofficial rubric test that we did here with a free local model, comparing it to the frontier, general use case models from fifteen years ago or or sorry, fifteen months ago, the best in the world. It actually did better because even though it failed right? And I wish it would have gotten it right. Like, the two that it failed previously, it actually got those right the first time I ran it, which didn't happen with, three seven SONNET or g b d four o.

Jordan Wilson [00:39:39]:
But still, head to head in the this very unofficial rubik, it did markedly better than the best models in the world from a year, in three months ago. Alright. So as we wrap this one up, here's what I want to leave you with. Open source AI is getting smaller, faster, and harder to ignore. Alright. Google built Gemma four specifically for agentic workflows with native function calling. Alright. So even though I didn't give an example of, you know, running this agentically, you now I cannot tell you how important that is.

Jordan Wilson [00:40:12]:
If you have a middle of the road new, you know, MacBook Pro as an example, you can now have an agent that works for you twenty four seven that costs $0.00. It's a 100% private. Also, what that's worth noting, this is based off of the Gemini three model family. So you're not getting quite the Gemini three level, but, again, you are getting a top three open source model in the world and the only one that you can run on consumer hardware. So now users can route routine AI tasks locally and cut significant cost on AI bills And the gap between the free local models in paid cloud services keeps shrinking fast, and you can no longer ignore it. Alright. If this was helpful, tell us a bit about it. You know, if you're listening, live here on LinkedIn, take a second to repost this.

Jordan Wilson [00:40:59]:
I'd really appreciate that. If you are listening on the podcast, do me a favor. Take thirty seconds. Make sure that you're following or subscribe to the show. But then if you could, if any episode of Everyday AI has been helpful, right, because we spend literally countless hours helping you all understand how this works. So if this has been helpful, please leave us a rating, all those platforms as well. So thank you for tuning in. Make sure to go to youreverydayai.com.

Jordan Wilson [00:41:24]:
Sign up for the free daily newsletter. We're gonna be recapping today's show and a whole lot more. So thank you for tuning in. Hope to see you back tomorrow and everyday for more everyday AI. Thanks y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI