Ep 321: Meta Llama 405B and Llama 3.1 – What’s new and what you need to know

Episode Categories:

Llama 3.1: The Latest AI Evolution and What it Means for Business Competitiveness

Recent advances in AI technology have seen the looming launch of OpenAI's latest model, Llama 3.1, with capabilities far outstripping its predecessor. Now available on platforms like Microsoft, Google, Databricks, and Snowflake, this disruptive technology not only reduces costs but opens new horizons in AI responses, further strengthening the competition against Google and Anthropic.

Unleashing the Power of AI: Llama 3.1’s Capabilities and Impact

Llama 3.1 extends a unique model distillation function that allows larger models to train smaller ones — a groundbreaking feature for data generation and customization. Arguably, this evolution could revolutionize business operations by enabling firms to craft models of any size that can boost performance. However, it's worth noting the model requires powerful computing resources, especially for the larger 405b version.


The Power Struggle Between Proprietary and Open Source Models

The battle between proprietary closed source models and open source or open-weight models continues to intensify. Recent comparisons, such as the MMLU benchmark, reveal the somewhat shrinking gap between these two. Notably, the Llama 3.1 boasts a highly impressive MMLU benchmark score of 88.6, spotlighting the potential vitality of open-source models in businesses.


Disruptive Transformations: Implications of AI On Business Operations

There's no denying the release of Llama 3.1 alongside other updated versions like OpenAI's GPT 4 o Mini and 405 b could vastly impact businesses. These new models are already reshaping the role of AI agents compared to human counterparts. User-friendly features, such as chat transcripts sharing and forking and easy chat saving and unsaving options, only exemplify the increasing ease of using AI in everyday business.


Risk vs Reward: The Privacy and Security of AI in Business

However, the adoption of AI doesn't come without its concerns. As businesses start to embrace Llama 3.1, questions about ad retargeting have cropped up. It may be worth taking some time to fully understand the potential implications of placing such a powerful tool in the hands of employees and customers alike.


Future Thinking: Anticipating Change and Seizing Opportunities

The climate is ripe for businesses to utilize AI technology to grow. The rise of Meta's potential to have the most globally used AI model and the shift from metaverse to investing in AI and generative AI indicates a future heavily dependent on AI agents. The challenge now rests on businesses to anticipate this change and seize opportunities as they arise.

As AI continues to evolve, so too does the competitiveness of the tech industry. The release of new language models such as the 8b, 70b, and the significant 405b exemplify this movement. Businesses should consider these updates as opportunities for growth and seek to take advantage of their wide availability from providers such as Meta.ai, AWS, IBM, and many others. By embracing such innovation, businesses can expect a promising future in the era of AI dominance.

Topics Covered in This Episode

1. Highlights of Meta's New Model - Llama 3.1
2. Spotlight on the 405b model
3. Functionalities and User Interface Updates
4.Meta's Focus and Its Implication on AI and Business


Podcast Transcript

Jordan Wilson [00:00:16]:
The new four zero five b model from Llama and their Llama 3.1 updates are extremely powerful. I mean, it's open source, free to download, and I think it'll change the business landscape. But there's a lot of things that I don't think people are paying attention to, and it's probably not what you think. So today, we're gonna be going over that. Now I'm gonna tell you exactly what this Lima 405 b and 3.1, what these things are. I'm gonna show you a little bit of the new model live and go over 5 things that you need to know. Alright. What's going on y'all? Let's get this thing started.

Jordan Wilson [00:00:57]:
My name is Jordan Wilson, and welcome to Everyday AI. This thing is for you. Well, it's for all of us. It's a place where, you know, nontechnical people, but still people who want to be technical, can learn about all things generative AI, to grow your company and to grow your career. So if that's you, thank you for tuning in on the podcast. Always make sure to check out your show notes. We just launched a cool campaign I'm gonna tell you about here in a second, but always more information. And make sure if you haven't already, go to your everyday ai dotcom and sign up for the free daily newsletter.

Jordan Wilson [00:01:28]:
Alright. So before we get into all of this, and I'm excited to talk about what this new, update from Meta, the Llama 3.1, the 405 b, all what I think is going to change the very near and long term future of how businesses use generative AI. But before we get into that, let's start as we do every day by going over the AI news. Well, at least the news aside from, at a 3.1 because we're gonna be covering the llama model, obviously, in great depth today. But a new study reveals AI increases workload and burnout among employees. So a new global study reveals a significant disconnect between the optimistic expectations of AI's impact on productivity and the actual experiences of employees. So according to a new study by the Upwork Research Institute, 90 6% of c suite leaders expect AI to boost worker productivity. However, 77% of employees, surveyed reported that AI has increased their workload leading to challenges in achieving the expected productivity gains.

Jordan Wilson [00:02:37]:
This, the study highlights that AI is contributing to employee burnout with 71% of full time employees feeling burned out and 65% struggling with their employers' productivity demands. And nearly half of employees using AI are unsure how to meet the productivity expectations set by their employers, and 40% feel their companies are often asking too much of them regarding AI. Yeah. Seems like no one's actually trading people. They're just like, hey. Here's a bunch of AI. We're investing in it. Go make it work and be 10 times more productive.

Jordan Wilson [00:03:08]:
Not how it works. You gotta invest in your employees. You have to learn, but that's what we do here at everyday AI. So next piece of AI news. Bloomberg has just released its annual list of top 10 AI startups to watch in 2024. So Bloomberg has announced its now second annual list of the top 10 AI startups to watch, highlighting significant innovations and emerging trends in the AI industry. Yeah. These aren't, small little startups that started yesterday in someone's garage.

Jordan Wilson [00:03:37]:
These are all probably startups you all have heard of. So here is the list, but we'll have more information in the newsletter. So has OpenAI maybe that small little startup everyone's heard of. Then we also have Anthropic, Suno. Shout out to our Suno episode. I was just thinking about that today. We gotta get the CEO back on. 11 Labs, x AI from Elon Musk, Perplexity AI, Safe Super Intelligence Inc, the new company founded by, former OpenAI cofounder Ilya Sutskever, Cognition, Udi Udiow, which is similar to, Suno for AI Music, and Sora, the yet to be released AI video platform from OpenAI.

Jordan Wilson [00:04:17]:
Alright. Last but not least in AI news, and this actually segues us into our topic for today. Well, OpenAI and Meta have really intensified their AI competition with some recent updates from OpenAI. So OpenAI just announced free fine tuning for its gbt4o Mini model just hours after Meta launched its open source llama 3.1 model. So, this move and I talked about this yesterday, we just squeezed it in our newsletter, and I talked about it a little bit on Twitter as well. But this move by OpenAI, I think, is a direct response to Meta's new free open source model, in Llama 3.1, and it's highlighting the fierce competition in the AI industry. So OpenAI said that this free fine tuning offer for gpt4 o Mini, its new model aimed at developers, is valid through September 23rd and allows developers to create customized, model experiences for specific applications at no cost. This is wild.

Jordan Wilson [00:05:16]:
Yeah. So, normally, if you are working on a proprietary closed model like OpenAI, their their GBT models, like Entropic, like Google, you have to pay for training, which at times can be expensive, but this new announcement from OpenAI gives up to 2,000,000 tokens of training a day for free. This is wild. Like, I know sometimes I say that, but that's actually wild. Right? This is a lot of money that presumably OpenAI would have been bringing in, and they're just like, nah. We just saw what Meta released. Let's open up the floodgates and get for free. Alright.

Jordan Wilson [00:05:51]:
Let's get this thing started y'all. Before we talk about meta, just make sure if you haven't already, check out our new thanks a million campaign in today's newsletter, your everydayai.com. We're having a lot of giveaways as we celebrate our upcoming 1,000,000 downloads for everyday AI. Thank you for your supports. We're gonna be giving away a ridiculous amount of great, you know, both from a year long subscription to chat GPT consulting and a lot of other things. We're gonna be dropping a lot of prizes here soon. Alright. So let's get into it and talk about MetaLama 405, b, the new model, and Llama 3.1.

Jordan Wilson [00:06:29]:
So we're gonna give you what

Jordan Wilson [00:06:30]:
you need to know, what people aren't talking about, and we're gonna show you

Jordan Wilson [00:06:34]:
a little bit live. Alright. And, hey, for our livestream audience, thank you as always for joining, Michael and Brian and Fred. Couple Chicago people in the house. Denny, Cecilia, thank you all for joining us. So let me know. Have you all used, this new Llama 3.1 yet? Are you going to? What questions do you have? Please get them in now. I'll try to tackle them at the end, but let's just get straight into an overview here y'all.

Jordan Wilson [00:07:02]:
So here's what's new. So you have technically 3 different tiers of models from Meta. Right? The, parent company of Facebook. So, Mark Zuckerberg did all his media rounds yesterday, wrote a very long, you know, blog post about the future of Llama and open source AI. But here's essentially what you need to know. In April, Meta announced Llama 3. Right? And they essentially had they said that there's 3 different sizes. Right? A small, medium, and large.

Jordan Wilson [00:07:31]:
And that seems to be the new trend I think started by anthropic. Right? And essentially, there's use cases for small, medium, and large, and we'll talk about that a little bit more here. But back in April when Meta first announced llama 3, which is a huge upgrade, they only came out with the kind of small and medium models. So when we talk about what is 405 b and what is 8 b and 70 b, well, those are sizes of models. So, the big one, the 4 0 5 b, had not been announced until yesterday or is not available, whereas the other smaller, the small and the medium sized models were available, but they've been upgraded. Alright. So I'm gonna go ahead and, share my screen a little bit here, and we're gonna walk through some of the benchmarks because they are impressive. Alright.

Jordan Wilson [00:08:21]:
So, let's go ahead and share my screen here. So for for our podcast audience, I'm gonna try to do my best. I don't think you're you're gonna be missing out on anything here because I think I should be able to kinda walk us through this just a little bit. So like I talked about here in, livestream audience, let me know if you can if you can see this. But so multiple new models from Meta. So like I talked about, we have a think of it as a small, medium, and large, just like anthropic. Right? So anthropic has their haiku, small, their sonnet, their medium, and their opus, large. So for meta, they're actually just naming it by the number of parameters it's trained on.

Jordan Wilson [00:09:04]:
So they have their 8 b, which is their small, their 70 b, which is their medium, And then they're just now released 405 b. Big jump up. Right? And without getting too technical, right, because our audience here is, for the most part, not super technical people like myself. Right? The easiest way to think about parameters is the amount of data that it's trained on. Alright? So let's go and just jump into the benchmarks because that is kind of our first point that we wanted to

Jordan Wilson [00:09:33]:
talk about is these benchmarks are pretty impressive. And, again, I have to really hit rewind even, Because we have to draw

Jordan Wilson [00:09:44]:
a line. Because even if we're looking here at these benchmarks, and this is what everyone always talks about and rightfully so. Right? So, benchmarks are when all the smart researchers and scientists, from both meta and everyone from all the big companies, third parties, they all do these benchmarks, and you essentially get a score. Right? So think of, like, you know, a new car. Right? When a new car comes out, you know, you get, oh, the EPA estimated gas mileage and, you know, goes through all these, you know, 3rd party, you you know, safety tests, all these things, and it gets scores. Right? So, large language models are the same way. And there's all of these different, dozens of different benchmarks, but the one that we talk about a lot here on everyday AI is the MMLU. Okay? So the MMLU benchmark is the massive multitask language understanding.

Jordan Wilson [00:10:33]:
So it's essentially 57 different subjects, and you get a score. Right? So it's like the ACT or the SAT for large language models. And this is by far the gold standard. So when we talk about MMLU and this new llama 3.1, the large variety. So that's what we're gonna be sharing here in the screenshots. It is very impressive, and I have a, I have a screen here earlier that I'm gonna be showing. But we have to keep in mind, this is an open source model, y'all. So what that means, this is free.

Jordan Wilson [00:11:11]:
This is free to use. You can download these models, although, you you really have to have, like, the world's strongest computer to download the 405 b. But for the small and medium models, most of us out there, if you have a new ish computer with a a decent, you know, graphics processing, chip, if it has decent specs, you're gonna be able to download, these these models from, from Meta. So that's like before we get into these benchmarks anymore, we really have to talk about the importance the importance of this is open source. Right? You can download it. You can build on top of it without having to pay. Right? Those inference costs, those training costs over and over, which is huge. Okay? And this is also it's available in a lot of dev environments, which I'm gonna show you here on

Jordan Wilson [00:12:04]:
the screen soon. So with that, let's look

Jordan Wilson [00:12:07]:
at these benchmarks. So the MMLU, llama came in at an 88.6,

Jordan Wilson [00:12:14]:
which is extremely, extremely impressive. Alright. I might do a full

Jordan Wilson [00:12:22]:
show soon. Live livestream audience, let me know. Do you wanna see that? Might do a full, show soon on MMLU. What are these benchmarks? What do they mean? But, essentially, if you are an expert in one specific thing so, again, these models are trained or sorry. The MMLU goes across 57 different subject areas. Think if you are a world expert in 1. A world expert is gonna get about, an 89.8. Alright? An 89.8.

Jordan Wilson [00:12:50]:
A world expert on that one domain specific field. Okay? But in the other 56, the average person gets about a mid thirties. Okay. So think. When we say that Llama has an 88.6, that means that a free model that you can download now is essentially a world class expert or almost as good. It's like having the 57 smartest people in the world available for free their entire knowledge, and you can build with it. You can download it. You can work with it.

Jordan Wilson [00:13:24]:
You can make it your own. Okay? So that's what we talk about when we're talking about both MMLU and the power of having these high up scores in an open source model that you can download, you can fork, you can build upon, and you're not having to, you know, pay a big company each and every time. Alright. So benchmarks are extremely important. So, the llama 31, it is above quad 3 5 sonnet, but it is just below, GPT 4 Omni. So GPT 4 Omni, 88.7. Llama, 88.6. Right? One fraction.

Jordan Wilson [00:13:59]:
Right? One one fraction of a point, away, which I was surprised about. I was surprised that, Mata didn't sit on this for another couple of weeks and try to over engineer and try to squeeze a little bit more juice out of it. So it's the, world leader on MMLU. But regardless, a free open source model that is essentially now the most powerful model. Extremely impressive. So, yes, the human eval scores also, very, very top notch, an 89 when, the the leader Claude is a 92. Pretty good there. A couple, though, that are worth noting about.

Jordan Wilson [00:14:38]:
I'm not gonna talk about benchmarks this entire time. But 2 other things. I mean, gsm8k, which is essentially, you know, math. Right? Basic math. A 96.8, the most capable model right now in the world. Also, the arc challenge, which is reasoning, got the highest scores of any model. So, when we talk about benchmarks,

Jordan Wilson [00:15:01]:
extremely impressive. Okay? We also have to talk about availability, where this is available at. Right?

Jordan Wilson [00:15:09]:
So we talked about you can download this now. You can also go to meta.ai. I'm not sure which countries have access. I didn't get a full list from Meta. I'm sure they'll be rolling out with that soon, but you can go to meta.ai. You do have to log in with either Facebook account and Instagram account, but you can use it for free literally right now. And I'm gonna be going over a little bit of that live. But we have to also talk about where this is available.

Jordan Wilson [00:15:35]:
Right? Because if you are a business leader and you are looking let's say you're working at a Fortune 500 and Inc 5 1,000 company, a big enterprise. Right? This is available now showing on the screen here. This is available now in so many places. So many places here. So AWS, Databricks, NVIDIA's Foundry, IBM, Google Cloud, Microsoft, Scale, Snowflake. Right? So so where so many of these big companies, big enterprises house their data, where they're trying to marry their data with the right large language model. It's available now. It's available to date, which I think, I love that from Meta.

Jordan Wilson [00:16:20]:
Right? You can have whatever thoughts you want about Mark Zuckerberg. You can have whatever thoughts you want about the social media side of of Meta, you know, Facebook and and Instagram and WhatsApp and data collection, all of these things. You can have whatever thoughts. But the fact that Meta announced this, dropped this, and it's available all instantly, hats off. Right? Because Google is notoriously bad for, you you know, having all these big conferences. Right? At their Google IO conference, they announced all these things. This is months ago, and we haven't seen a fraction of them at least when it comes to their new large language model developments and in generative AI. Meta, it's like drop a blog post, drop some interviews, and it's live.

Jordan Wilson [00:17:03]:
It's ready. So you can probably, today, go and work with this in your environment right now. And, again, it's open source, y'all. Alright. So like I said, the first the first thing that you need to know is the specs are extremely impressive. So like I talked about that, the benchmarks are great. A couple other things. The 8 b and the 70 b versions, the small and the medium, those are upgraded.

Jordan Wilson [00:17:32]:
So that is the big jump from, you know, llama 3 to llama 3.1. So even if you look at just the improvements in the small and the medium model here, they're they're they're huge. They're, I mean, they're very impressive. So the small and the medium models, if you look at them in their respective, quote, unquote, weight classes, right, Especially the 8 b. Because I think the 8 b, probably within, I don't know, 6 months to a year of hardware, I think you're gonna be able to see this as an edge device, as a model that could, in theory, run locally on a phone. So that's the other huge upside to being open source and to be able to download a model is you can run it locally without Internet. So then, privacy concerns are less. Speed is more.

Jordan Wilson [00:18:28]:
Environmental concerns are less of a concern at that point, when you can run a model locally and you don't have to, you know, run it off of essentially a server. And so it's faster. The latency's lower. It's it's more secure, and you don't even need the Internet. Right? But the 8 b model, especially, I'm looking out of this, and this is what I don't think people are talking about. So if you compare it to Google's new JEMMA 2 model, so these are, again, I would call these something between a small language model and a small large language model. There's no actual definition because the goalposts are always moving. Right? But these are models in theory, llama 318b8000000000 parameters and Gemma 29 b.

Jordan Wilson [00:19:10]:
These are models that in theory can be running locally on a smartphone probably within 6 months to a year. Right? So right now, Google also has Gemini Nano, which is one of the ones that's running on current smartphones. But you have to also think when you see all of these announcements, we can't just look at the benchmarks and what they

Jordan Wilson [00:19:29]:
mean today. We have to

Jordan Wilson [00:19:30]:
look at what do they mean in the future. So I already told you, This has huge future implications, which we're gonna get into about how businesses can can run now. But you also have to think of what this means for the future of on the go AI, which is the future. I've said hundreds of times, the future of large language models is small language models and working with many of them and working with them on edge AI, on device AI. But the 8 b model just is out punching its weight class in almost every single benchmark, aside from, like, one. It is the top model. It is the top small large language model, and it's not even particularly close. And, again, open source.

Jordan Wilson [00:20:11]:
So think of what this means for the future of even apps that you use. Right? Because now all of these developers, you know, it's not like even with GPT 4 o Mini, which we're gonna talk about here, that brought the cost way down. But, you know, a month or 2 ago, you know, it could be expensive to to to go launch a brand new app that was powered by AI for, you know, enterprise business or just something that's, you know, fun for people to use. This changes things. This changes what people with low budgets can go and build in a in a weekend or in a day and launch immediately. This brings scalability to even small companies that don't maybe have, you know, compute power. Well, you you might not even need it with an 8 1,000,000,000 parameter model, which is fairly small. You can download this and run it on a machine that's not even very powerful.

Jordan Wilson [00:21:00]:
Alright. And the benchmarks are extremely impressive there. Also, another thing to note, and I promise you this is it for our, for our more technical side. Alright. So a 128 k context window. Huge. Right? Because the context window was very small before. I believe it was, 8 k or less.

Jordan Wilson [00:21:19]:
Also, meta is using a slightly different, and this gets into a little bit of the technical side, but mixture of experts or MOE is what a lot of these large language model companies have been using. So, meta will put it in our in our newsletter today. They went a different route going this herd of experts, versus the mixture of experts. Alright. So enough on the technical side. Let's go ahead and talk about the second thing that you need to know. And I think that this is shots fired. I mean, shots fired against OpenAI, but this is a developer focused strategy.

Jordan Wilson [00:21:55]:
Right? Because I think people are overlooking the vast improvements that were made in the 8 b and the 70 b, the small and the medium models. And what this means is I think they're just going after OpenAI. Right? We saw Anthropic with, Claude 35 SONNET. We saw Gemini Google Gemini with 1.5 Flash. Right? So, and Claude presumably will be releasing 35 Haiku, which is their small model. Right? So Haiku Sonnet, Opus. So over, I'd say, in the 1st quarter or 2 of 2024, OpenAI had lost traction, I think, among developers. They were jumping ship.

Jordan Wilson [00:22:36]:
We had a whole episode about this on Friday, but, you you know, right after they released their new model, gpt4 o Mini, which is essentially a light version of their big boy model. Right? So now OpenAI has essentially a small and a large. They don't have a medium per se. Alright. But this really changed the developer landscape. So OpenAI, I think, was losing customers. They were going to, Google Gemini for 1.5 flash. It was cheaper and more powerful.

Jordan Wilson [00:23:03]:
They were going to Anthropic for, Claude 3 Haiku. It was cheaper and more powerful. You know, OpenAI was essentially relying on either 3.5, which wasn't very powerful, or GPT 4, which was very expensive. So I do think that OpenAI did make a huge splash with GPT 4 o Mini. Did that happen to coincide with the fact that Meta decided days after to release their 3.1 updates? Potentially. Right? Because I think they are now going directly after developers. And it can also be lost on us how OpenAI responded. I talked about that at the top of the news this morning.

Jordan Wilson [00:23:41]:
Literally hours after meta released a very impressive 3.1. 3 models, benchmarks go wild. Right? Literally hours after that, OpenAI said, oh, you know what? We're actually gonna make cost 0 for the next couple of months to come and fine tune your models. Right? So hey. Hey, developers. Hey, businesses. Right? Hey, businesses. You wanna, you know, bring in rag.

Jordan Wilson [00:24:09]:
You wanna have this retrieval augmented generation. You wanna kind of, marry and create your own large language model with your data and some fine tuning. Generally, you know, a year and a half ago, very expensive, very hard to do. Now it's easier and it's cheap. So OpenAI said, hey. For the next couple of months, you can come in and do it for free. 2,000,002,000,000 tokens a day, which I think is in direct response. They saw what Meta dropped, and they're like, woah.

Jordan Wilson [00:24:35]:
These benchmarks, pretty dang good. Pretty dang good. Right? We don't have comparisons yet with the, 8 b, 70 b, and GPT 4 o Mini, but I'm sure those will be rolling out here soon, and we'll talk about it. But this is a developer focused strategy from Meta, and it is a direct shot, to to OpenAI. Alright. So I do think at least right now, it's a 2 player battle. It's a 2 player battle. But what this does mean is I fully expect, Google and Anthropic to be releasing updates in the next 2 months.

Jordan Wilson [00:25:14]:
Right? This changes. It it it's it's getting so so cheap where it's it's getting to the point to train a model, you're not gonna have a lot of costs. Right? We always talk about inference and and the cost per token of training.

Jordan Wilson [00:25:31]:
When we see a matter like a model, like, Llama 3.1, and they're 8 b, they're 70 b, It's free 99 y'all. Yeah. You gotta have

Jordan Wilson [00:25:40]:
a computer powerful enough to to handle it. You still need the skill set, but the cost to train a model on your company's own data, the cost to build something on top of a large language model is going down to, like, 3.99. Right? Whereas a year and a half ago, pretty expensive. And now the availability is everywhere. You saw the chart. Right? It's NVIDIA it's it's available inside, NVIDIA. It's, platform. Their their their Foundry platform.

Jordan Wilson [00:26:08]:
It's available inside of, Microsoft, inside of Google, inside of Databricks, inside of Snowflake. Right? So all these, dev environments, it's available there. So the cost is going down. And this means that I think Google and, Anthropic are going to be responding here soon. Alright. Number 3, the, thing that you need to know is it is ready to be built upon. Right? So we did kind of talk about that.

Jordan Wilson [00:26:37]:
You don't have to wait. You don't have

Jordan Wilson [00:26:39]:
to wait. It's literally available now, and it's available in all of those platforms. I didn't even talk about AWS. But, yeah, AWS's platform, Dell, NVIDIA, Grok, Grok with a q, not the large language model from Twitter that no one uses. IBM, Google Cloud, Microsoft, Scale, Snowflake. It is ready now. Right? Which I again, I talked about this at the top of the show. I love meta's approach here.

Jordan Wilson [00:27:04]:
Don't don't make us wait. Right? Google is is like, hey. We're here's some marketing. Here's some promise. Let's see how our stock goes up or down, and maybe we'll release these in the next 6 months to 2 years. OpenAI kind of in the middle. OpenAI, you you know, when they have their spring event, they drop their model that day. But a lot of the more powerful features, we're still waiting on, you know, we should be getting them in theory any day.

Jordan Wilson [00:27:31]:
But Meta Meta just literally, you know, Mark Zuckerberg, new hairstyle, looking looking fresh with the gold chain, just drop, mic drop. Here's everything. Go out and play. Available everywhere now. Right? Changes how business is done. All right. So number 4, the number 4 thing you need to know in open source, Lama 31 is more than your average large language model. Okay.

Jordan Wilson [00:28:00]:
Again, we'll be linking to this in our newsletter so you can read a little bit more from both, Zuckerberg's long blog posts on the future of open source, and what that means, but it means a

Jordan Wilson [00:28:12]:
lot of things. So he talked a

Jordan Wilson [00:28:14]:
lot about model distillation, distillation. So that essentially means the big model, the 405 b, can act as a teacher to, quote, unquote, train student models. Right? You can also companies can make models of any size. That's the other thing. If you are working, with a closed source proprietary model, which is what all the other big guys are. Right? So Google Gemini, Anthropic Claude, Chat GPT, etcetera. Those are all closed source proprietary. Right? It's it's it's not a a bespoke option.

Jordan Wilson [00:28:47]:
You know, it's one size fits all or you don't use it. So you might be overpaying or you might be underutilizing. But with this model distillation with meta 3.1, you can make models of any size. Right? You can use the big model to train the smaller models. Also, a huge thing here is the synthetic data generation. So, Llama 3.1 can generate high quality synthetic data to enhance the performance of smaller models. Right? We always talk about where is this data coming from? Where is this future high quality data coming from? Which I personally think is a huge problem. Right? Because models, for the most part, are trained on the open Internet.

Jordan Wilson [00:29:27]:
And guess what? People don't know this, but this predates chat GPT's release. People, you know, SEOs, content writers, they've been using the GPT technology since 2020. So I think you have this almost this data regurgitation issue where so much of the data right now is being built, by AI. Right? A lot of studies say by 2026, more than 90% of new content on the Internet will created with AI. So you have this problem. Where do you get this new this new data, this original data? So, what Meta is kind of doing here with this synthetic data generation is you use the big model to create, quote, unquote, new model, or sorry, new data to train these smaller models or to train your company's models. Another thing is the customization. Alright.

Jordan Wilson [00:30:15]:
And this is why it is much more than your average large language model. It allows for model architecture customization based on specific task or hardware constraints. So what that means is if you do want to build, let's just say a customer's, customer service large language model for your company. Right? Put in all your data. If you're using a GPT 4 0 Mini, if you're using a quad 3, if you're using a Google Gemini, you, in theory, are going to still be using the full model. Right? So you can essentially, with, an open source downloadable model that you can fork, you can build upon, you can essentially do anything you want. Think of it like this. It's a template inside of a Word doc.

Jordan Wilson [00:30:57]:
It's a template inside of PowerPoint. You can move things around. You can, you know, if it's a 100 pages, you can delete it down to just the 10 pages you need. You can't do that with the other models. Right? Think of them as like a PDF. You can't go in and update them. So a lot of times, you're either overpaying for what your your business may need or you're sacrificing speed. Right? Because you still gotta use the full model even if you only need

Jordan Wilson [00:31:20]:
a little sliver. Okay? And and also,

Jordan Wilson [00:31:24]:
I mean, we have to talk about our last thing that we need to know. Well, actually, let's look at this, a graph here first.

Jordan Wilson [00:31:34]:
This is changing. Alright. The the gap between proprietary

Jordan Wilson [00:31:41]:
closed source models and open source or open weight models is diminishing. So shout out to the original creator of of this, kind of graph that I'm showing here, Max

Jordan Wilson [00:31:54]:
Melaboni, I I I believe on Twitter there. So about 2 years ago, night

Jordan Wilson [00:32:00]:
and day difference. Alright. So, we have essentially MMLU on the left side and then we have, our dates on the other axis. And we see that there was such a huge disconnect about 2 years ago, even about a year and a half ago between these open source models, right, that are free and available and anyone can fork them and do whatever they want, and the closed source models. As of yesterday, there really is no difference anymore. It used to be night and day. It used to be that open source wasn't really a good option for your business because you were sacrificing on output quality. You were sacrificing on MMLU.

Jordan Wilson [00:32:36]:
You were saying, hey. We could use this free open source model, but it's kinda dumb. Not anymore. The gap is essentially nonexistent. It is point one. It is point one difference on the MMLU. Right? Whereas before it was 10 points, 20 points, 30 points. Now it's essentially a wash on is the model smart enough? Is it capable as capable as the others? Right? When we talk about the difference between the world leading model, gpt4 Omni, 88.7, and the now free 405 b from Llama, 886.

Jordan Wilson [00:33:17]:
It's no longer night and day. Right? It's no longer night

Jordan Wilson [00:33:20]:
and day. We're essentially talking about and looking at the exact same thing. Alright.

Jordan Wilson [00:33:29]:
Let's go here. So number 5, the last thing you need to know is obviously this new model has tremendous business impacts. Alright. I'd say the last 72 hours between OpenAI's GPT 4 o Mini with meta's 3.1 release, the 4 0 5 b release. And also now with OpenAI saying, hey, come play for free until September. This completely changes what's what's possible with your business. Right? And I'm talking a little slow here because I'm I'm thinking, and I hope that you're thinking about this too. And I wanna point something out that Mark Zuckerberg said yesterday in an interview.

Jordan Wilson [00:34:20]:
And I said this y'all. I kid you not. There's always receipts. I said this back in December, and I think people laughed at me and thought I was crazy. But back in December, my prediction was by 2024. We'll see. But I said there will be more AI agents than humans. Right? And funny enough, Mark Zuckerberg said the exact same thing yesterday.

Jordan Wilson [00:34:46]:
I hadn't heard any, you know, big person in tech say anything like that. And yesterday, I'm like, yeah. Exactly. I said that in December. And, you know, now Mark Zuckerberg, who you could argue as one of the most, prominent people in the world in AI right now. Right? At least a handful, you know, of 5 of the most prominent people in the future of AI, the future of artificial, general intelligence, AGI, in theory, the the future of ASI. Right? Mark Zuckerberg is one of the most important and prominent people in the world right now because AI and generative AI and large language models impact business. Alright? And you'll see, Meta has essentially thrown not thrown away, but they've set aside their whole metaverse thing.

Jordan Wilson [00:35:29]:
Right? 3 years ago, it's metaverse metaverse metaverse, and now it's like spending, you know, 100 of 1,000,000 of dollars on compute. Right? They spent, I believe, $700,000,000 just on GPUs to train meta 3.1 and their future meta models. But, you know, the future business impacts, Mark Zuckerberg predicted there will be more AI agents than humans. He said every single business is going to have multiple agents. Right? And he's not just saying on their platform. He's talking about the bigger picture. And he does think and hope that, the llama model is going to be the most used AI model in the world. And you have to think, it kinda might be possible.

Jordan Wilson [00:36:13]:
Right? It might be feasible because

Jordan Wilson [00:36:18]:
guess what? Their model is everywhere. Yes.

Jordan Wilson [00:36:22]:
Can go to Meta AI, which we're gonna show you very briefly here in a second, but it's also available in Instagram. It's available in Facebook. It's available in WhatsApp. So Meta has the reach. And guess what? Yes. When you use these models, you're making them smarter. So, you know, people are always like, oh, it's free. Right? Oh, have another episode at some point, the downsides of using AI.

Jordan Wilson [00:36:45]:
And if you're using something for free, you potentially, yeah, you you are paying for it by giving them your data, giving them your feedback, etcetera. But Meta does have the chance with their model to be the most used AI model in the world and it's pretty feasible. And so you have to pay attention when Mark Zuckerberg talks about how AI AI agents large language model are going to completely reshape what not only what is possible, but how business works. And I've been saying this since day 1. The future of business, especially here in the US, we are all going to be working with agents, large language models, small language models. We're gonna be prompting. That's what our future work is going to be dependent on. You are not going to be in I don't know if it's gonna be 2 years, 5 years, 10 years.

Jordan Wilson [00:37:33]:
You're not gonna be rewarded and promoted. Your company isn't going to grow by the experts, by the subject matter experts anymore, by the knowledge, by what you know, because all of that is being transferred into large language models. And in the future, small language models. Right? When you have these very capable models or agents that are gonna be trained and fine tuned on one very specific task. Right? Think of that one specific thing that you do. Right? And then let's say you're lucky enough to be, you know, in the top 1% of people. Let's say, in that one very specific thing that you do, you're you and a 100 you know, you and 99 other people are the smartest in the world at that one very specific task. Guess guess what? Very soon, a small a small language model, because of what's happened in the last 72 hours, there's gonna be a fine tuned model that does that one specific task much better than you and the other 99 smartest people in the world combined.

Jordan Wilson [00:38:29]:
So the future of how we work is getting the most out of these agents, knowing how, when, and ultimately, why we should be using them and how we can squeeze the most business value out of agents, out of generative AI, out of large language models. That's the future of business. And that's why I think Meta llama has tremendous business impact because over the past 72 hours, the future of how we can work and the timeline the timeline has shortened exponentially. Because even a couple of months ago, you'd say, you know, this this this is gonna take a while. Not anymore. This is like, you you know, this is the equivalent of, you you know I don't know. 20 years ago, it used to be hard to get your company on the Internet. Right? You had to hire someone smart.

Jordan Wilson [00:39:15]:
You had to put a lot of work in. Right? So imagine 20 years ago, someone just drops. Hey. Here's the best website in the world. We're gonna do it all for free. Right? Imagine the gold rush to a.com. Right? It wouldn't have taken 10, 15, 20 years, for companies to really thrive online. This is a legit generative AI large language model gold rush.

Jordan Wilson [00:39:38]:
The big players are putting things down and saying it is now free or free 99. It is cheap, and the possibilities are endless. Alright. So those are the 5 things you need to know. So now let's just take a quick look live. Alright. So for our livestream audience and hey. I'm gonna be getting if you do have questions, like Monica said, did you do a demo for this? Got one right here.

Jordan Wilson [00:40:05]:
Got one right here. But if you do have any other questions, go ahead and get them in. I'll try to answer them here at the end. Of course would be good. Should I do a meta course? Gordon saying I should do a meta course. Maybe. Maybe. Let let me know if you think that would be helpful.

Jordan Wilson [00:40:19]:
Alright. So, alright. Let's go. We're we're just gonna do some very basic things. Alright? I did a whole video, yesterday, running through this, but let's just do some common things here. So I am going to meta.ai. Okay? So if you you do have to log in you do have to log in to your either Facebook or your Instagram account. Alright.

Jordan Wilson [00:40:46]:
So a couple things you

Jordan Wilson [00:40:47]:
need to know, and I'm glad they did this.

Jordan Wilson [00:40:49]:
There's even there's even been updates since yesterday. I really railed against Meta, for a couple of things, and strangely enough, they fixed them. Some of those things have been fixed. A couple there was some confusion in the model selection. So now, you want to go click your profile and go to settings. Again, there's other ways you can access this. You can access this also on Hugging Face. You can download the models, run them locally.

Jordan Wilson [00:41:13]:
I'm just showing you probably the easiest way to use this, which is just going to meta.ai. Alright. This is not the front end interface that maybe you're used to. It doesn't have the features and functionalities of a chat GPT, of a Claude, even of a Google Gemini. But if you just want to see if the model is right

Jordan Wilson [00:41:30]:
for you for your business, it's a great way

Jordan Wilson [00:41:32]:
to do it. Alright. So you can go click on settings, and then you can go to change your model. Okay. So I already changed it to the big one, which is 4 0 5 b. So you can either use the 70 b. So, again, these are the 3.1 updates. So the small and the medium were updated to 3.1.

Jordan Wilson [00:41:49]:
So yesterday, it still said 3. And I'm like, meta, what's going on here? But you can either use the 70 b or the 4 zero five b. So you can't use the small one. I'm wondering if they're gonna change that, pretty soon. But so right now, I'm going to select meta. I'm gonna select llama 3145 b. Alright. So a couple things on the interface.

Jordan Wilson [00:42:10]:
It's super easy to use. Super simple. Right? So you have your chat history on the left hand side. At any time, you can collect, click new conversation. But all you can really do there's there's 2 things. Right? There's text. So you can do text to text or you can do text to photo. I'm not gonna be going over that, but they have this imagine feature.

Jordan Wilson [00:42:27]:
It is pretty cool because you can type something in real time. So I can say, Chicago, and it's gonna start I don't know why it brought up a a random dude. So I can type in Chicago Skyline, and it actually should be, doing things live. I can type in Chicago skyline hotdog. Right? So then there's a hotdog. So, you know, it does these live. I'm not gonna go over, this too much here, the, imagine feature, the the text to image. But the rest of this, you can click new conversation and then just chat with Meta AI like you would any other large language model.

Jordan Wilson [00:43:00]:
Alright. So, I'm gonna run a prompt that I run some time. And, again, I did this, yesterday. So I'm saying this is a logic, a logic prompt, essentially. So I said, I just woke up today with 6 apples and 3 bananas. Yesterday, I ate a banana and 2 apples. This morning, I will eat 1 apple and no bananas. However, I don't really like apples and 1 banana may turn brown tomorrow.

Jordan Wilson [00:43:24]:
Assuming nothing else changes, how many apples and bananas will I have tonight? I'm going to create an actual, like, quote, unquote testing series of questions that I can do with all these models, but this is one I generally use. So if there's a new model, you know, SONNET 35, GPT 4 o Mini, etcetera, I usually have a set of 5 to 10 prompts. These aren't scientific, but these are, either logic, reasoning, math, creativity. I have a couple prompts that I I I run. A lot of models get this wrong. So, what I like by default. Right? So, again, I'm using the 405 b. By default, which I like here, Mara takes essentially a chain of thoughts prompting, response even though I didn't tell it to.

Jordan Wilson [00:44:08]:
You gotta love that. The first thing it says is let's break this down step by step, which is a great prompting technique. Right? It's it's it's kind of the, the essence of of what we do in our free prime prompt polish course. Right? So it's saying let's break it down step by step. So I'm gonna go down to the bottom, see the answer. It says tonight, you will have 5 ample 5 apples and 3 bananas, which is correct. A lot of models get confused. I put in a lot of, nonsense, to try to throw the model off.

Jordan Wilson [00:44:34]:
It does a good job here. Alright. Let's go ahead. I'm gonna do another prompt. This one, so many models get wrong. Yesterday, MATA got it kinda right, kinda wrong, so I'm gonna run it the same. Also, what's important to know? I say this a lot. Generative AI, it's kinda like rolling the dice.

Jordan Wilson [00:44:50]:
Right? Which is why prompt engineering and understanding how models work is so important. You can run the same prompt a 100 times, get 99 different results. You could get 3 different results. Right? It is generative. Alright. So let's go ahead and try this prompt here. I'm saying a man and his dog are standing on one side of the river. There's a boat with enough room for 1 human and one animal.

Jordan Wilson [00:45:12]:
How can the man get across with his dog in the fewest number of

Jordan Wilson [00:45:15]:
trips? Alright. So it says a classic puzzle. Right? Which I I do think there's all

Jordan Wilson [00:45:22]:
of these kind of, like, brain teasers or kind of large language model logic tests that a lot of people have been doing. And I do think by now a lot of large language models, because they're trained on the open Internet, are scraping all of these. Right? So people post these to to Quora, to to Reddit, blog posts, etcetera. And so models, I think, over time, understand the answer because they gobble up all this information on the open Internet, and then they learn from it. So, let's see if this gets it right. But most models get this wrong. You know, so let's see how it does. Alright.

Jordan Wilson [00:45:53]:
So this this get it got it wrong. So it said, it essentially said 3 trips. Right? The correct answer is one trip. If a man and his dog are on one side of the river and they have a boat with enough room for 1 human and one animal, it just takes them one trip to get across the other side of the river. This one says it takes 3. Yesterday, it said, technically, it said 2 or one round trip. So it got it kinda right yesterday. Today, it didn't get it right.

Jordan Wilson [00:46:19]:
It got it actually, fairly wrong. Alright. Let's try one more. Again, this isn't supposed to be a full live breakdown of, Meta, but I wanted to at least show you some of the capabilities of what it what it's, what it's possible. So this is something that I think most businesses could relate to. So I'm saying, oh, maybe not. You're maybe not creating a new company, but, using this to brainstorm, to ideate, to strategize. That's what large language models are great at.

Jordan Wilson [00:46:51]:
Essentially, being a companion to help you build a business, being a companion to help you market your, department's campaign, being a a a companion to help you essentially poke holes in things. Right? So I'm saying here, create a a new company, and brand for a future smart home device. This will solve a problem that does not currently exist. To start, come up with the company's name and its first flagship product. Give the product a name, branding campaign, go to market strategy, tagline, and rationale for why it will work. Respond in a succinct way, keeping responses to short bullet points, short bullet points, but with ultra specific facts. Alright. Hey.

Jordan Wilson [00:47:33]:
I might want this one. So here's what Meta said in its, 405 b 3.1 inside of meta.ai. The company name is Echoplex. The flagship product is Dreamweaver. It solves the problem of sleep data overload. I could use that. The you know, I'm not I'm not gonna read the whole thing, but, you know, the product it has the product description in there. Looks good.

Jordan Wilson [00:47:55]:
The branding campaign, the tagline is unlock the hidden narrative of your mind. Not bad. I've seen worse. It even has a color scheme with, hex color codes. It gives examples for a logo. This is a zero shot prompt, y'all. A very poorly written. There's no prompt engineering and the results are pretty good.

Jordan Wilson [00:48:11]:
That's one thing I've noticed with, this new model. You don't have to do as much prompt engineering to get something decent. Right? A lot of times, short prompts and other models don't get you great things. It does always almost seem like that over delivers when your prompt underdelivers, which which I like. So it does even have this kind of, almost seemingly built in chain of thought reasoning. And it always seems to give you a little bit of more without just being Verbios. Right? So a lot of large language models, what they do is if if there's ambiguity because your prompt kinda stinks, it's just gonna spit a bunch of general nonsense content that you don't it's not really good. Meta doesn't really do that.

Jordan Wilson [00:48:47]:
I think they give you specific content that is actually good. So we also have the go to market strategy. It's it gives us a target audience, launch channels. It gives pricing, how much this device should be priced at. It has the rationale, which is pretty pretty spot on. Right? It says there's a growing interest in brain computer interfaces and neural implants. That's true. A lot of, startup money is going into that.

Jordan Wilson [00:49:11]:
Increasing awareness of mental wellness. True. There's a unique value proposition. True. It's even getting key partnerships. So, alright. That's that's enough, for a a live model. But, again, actually, 2 other things with the interface.

Jordan Wilson [00:49:25]:
I I did mention that there's some new things that weren't there yesterday, which I think are important to talk about. And I actually wish that more companies would have this. So, at any point, you can tell the model if it's a good or bad result. So there's a little interface you can hover over. You can copy content to the clipboard. So here's the 2 new things which I like. You can share. Okay.

Jordan Wilson [00:49:47]:
So I can then share this chat. I can copy the link and, whoever I send this to, as long as they have an account, can go in. They can see the, kind of the transcript, and then they can pick it up. So they can essentially fork the chat. It's not gonna be shared. Right? But, I can essentially give them access to everything I have created up until that point. It's like creating a copy of a document, and then that person can have all of that knowledge, have all that work, have the context window, which is also important, and then continue to work with it. Right? So that's a brand new feature that wasn't there a couple hours ago.

Jordan Wilson [00:50:18]:
And then also the ability to save or unsave. Right? And that's especially important. I wish chat gbt and Claude had the speech, feature because now on the sidebar inside meta, I can just click save and then it's just gonna have my saved chats. Alright. So that's enough of a quick overview. Let's wrap this thing up and get to your questions. Alright. So a couple of questions here.

Jordan Wilson [00:50:45]:
Denny asking, should we expect to see ad retargeting based on what we put into Llama? Same as can happen based on interactions on Facebook. You know what? I'm not actually sure. There was a lot of information that came out yesterday. I did a video review. I spent hours planning for this show. I haven't read through the whole policy, so we'll we'll look at that, Denny. I'm not sure. The MMLU score, yes.

Jordan Wilson [00:51:10]:
It is ranking correct answers. Right? So think of it just like a standardized test. When we talk about MMLU, it's essentially a standardized test. Right? There's right answers and there's wrong answers, and there's different prompting techniques. Right? So there's some MMLU scores which are based on 0 shot, which is essentially copy and paste prompt with no input output pairings that tell a model what's good and bad. But then there's also, you know, like 5 shot, you know, MMLU scores, which means you can do a little bit of prompt engineering and, you know, kind of help the model get to the correct score, or or sorry, to the correct answer. But for the most part, in MMLU and a lot of the benchmarks, there's right and there's wrong answers, and there's different prompting techniques or methodologies that you use. And then it's like, hey.

Jordan Wilson [00:51:52]:
Did you get it right or wrong? Just like a standardized test. Yogesh. Yogesh in the house, former, former guest on Everyday AI. What's going on, Yogesh? So saying, will Llama 3.1 make custom GPTs obsolete on chat gpt? It's a great question. Right? Because when we think of I said that this is a huge play for the developer community. So there's something about chat GPT and even creating GPTs. The user experience is so nice. It's so easy.

Jordan Wilson [00:52:24]:
So in its current state, will this very new powerful three one make custom GPTs obsolete? I'd say no. Number 1, because OpenAI still has the most powerful model in the world even though it's not by a lot. But number 2, the experience. The user experience is much easier. Right? The fact that anyone can go into chat GPT right now, create a custom GPT with no coding skills, you know, it's not like it's low code. You can create a g a literal custom version of ChatGPT drag and drop. No code. Just say, hey hey, GPT builder.

Jordan Wilson [00:52:59]:
This is what I need. Let me upload my my database, and then it just does it. That's amazing. We don't have that quite yet, with, you know, Llama 31. We do have that a little bit with Claude. We don't yet have that with, Google Gemini, but we may once or if Google ever releases their quote unquote gems, which is their kind of counterpart to, to, GPTs. So right now, I will say no. It's not going to make custom GPTs obsolete because the user experience, right, it's not quite there.

Jordan Wilson [00:53:31]:
But will is that going to be an area where, meta plays in? Potentially. Right? That is where you start talking about, agentic capabilities or agent capabilities. Right? When they're trained on very specific tasks. That's essentially what a GPT is. It's something, custom GPT that you can train on one very specific task. And you say, hey. Big, big model. Just focus on this one one one very specific skill set.

Jordan Wilson [00:53:53]:
Here's my data. Let me train you and make sure that you can do this task correctly. So I don't think out of the box it's going to.

Jordan Wilson [00:54:00]:
Alright, y'all. I hope this was helpful.

Jordan Wilson [00:54:04]:
Went a little longer than I had hoped, but, hey, that's that's just everyday AI for you. Right? We do this live. It's unedited, unscripted. I hope today's show was helpful. If so, tag someone. Right? If you're here listening, in the LinkedIn comments on Twitter, YouTube, whatever, share this with someone. If you're on the podcast, thanks for listening. Thanks for tuning in.

Jordan Wilson [00:54:24]:
Please make sure to follow the show on Apple and Spotify. Leave us a rating if you think we deserve it, and, also, remember we have that. Nope. This is the wrong, the wrong slide there. But make sure you check out today's newsletter, for the campaign that we just launched. Alright? So it's a referral campaign. So if this is helpful, we help you refer your friends. Right? We spend so much time every single day cutting through all of the AI nonsense so you can stay ahead.

Jordan Wilson [00:54:52]:
We've heard from so many people. Hey. Everyday AI helped me get, you know, this promotion. Hey. Because of Everyday AI, now I'm leading AI initiatives at my company. If this is helpful, please tell your friends. Right? We keep this free because I think there's a lot of people out there that are putting out bad information and they're just trying to swindle you. They're trying to sell you something.

Jordan Wilson [00:55:10]:
Right? Here at Everyday AI, we are unscripted, unbought, unedited, the realest thing in artificial intelligence for you so you can know what's what. We do all the hard work for you. So if you appreciate that, which I hope you do, check out the show notes in your news, in in the podcast if you're listening there. Check out today's, today's newsletter and join the Thanks A Million campaign as we near our 1,000,000 downloads. Please join that, campaign. It's gonna be a lot of fun. And like I said, we're gonna be dropping, in the coming days weeks, we're gonna be dropping new prizes. So right now, top prize is a free year of ChatGPT or whatever large language model you want.

Jordan Wilson [00:55:46]:
We're gonna throw in some consulting sessions, etcetera, so make sure to check that out. And make sure to tune in tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI