Ep 319: OpenAI releases new model, Big Tech used thousands of YouTube for AI – AI News That Matters

Episode Categories:

Resources

Join the discussion: Ask Jordan questions on AI

Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Crunching AI Developments

AI advancements are sweeping across industries and reshaping traditional business practices. The surge in AI enthusiasm reveals considerable opportunities for businesses to remain competitive, diversify their offerings, and improve their operations. However, along with the advantages, significant challenges and controversies are cropping up.

AI in Education: A Disruptive Innovation

AI is ushering unprecedented changes in the education sector, with a vision to integrate AI assistants alongside human teachers to enhance learning experiences. Some initiatives are focusing on realising this vision by launching avant-garde courses. For example, course LLM 101n aspires to help students train their own AI, enabling them to build a functional web application using Python, C, and CUDA.

The course materials will be accessible online to cater to both digital and physical cohorts. Interestingly, the future of AI assistance could mirror real people, much like celebrity chatbots. Yet, instead of taking a hands-off approach, these models shall emulate a "storyteller AI large language model," paving the way for more interactive and engaging AI assistants.


The Challenges of Exploiting AI in Education

However, the lack of dedicated education in rapidly advancing AI technology could pose significant challenges. The current educational resources might stress sales rather than providing comprehensive education in AI, causing a crucial educational gap that might impede successful implementation.

Simultaneously, uncertainties persist around plan particulars, such as funding, costs, availability, and timelines, further adding to the complexity.


Legal Challenges Mounting Round the Corner

Amid the aspiring AI initiatives, businesses have witnessed several lawsuits and accusations, raising serious ethical and legal concerns. Tech giants have recently been accused of illegally using YouTube subtitles for training AI systems without acquiring necessary permissions.

The dataset, named YouTube subtitles, being part of a larger collection called The Pyle, has become a hotbed for controversies with potential copyright infringement issues. This case sure adds another layer of complexity to the already multifaceted issue of training large language models on copyrighted material and its impact on creators and news organizations.


Regulatory Challenges in the AI Landscape

Regulatory concerns have led to tech giants like Meta and Apple withholding some of their future AI models and capabilities from certain regions. For instance, Meta's newly introduced multimodal model, LAMA, which can reason across video, audio, images, and text, will not be available in the European Union due to the unpredictable regulatory environment.


Balancing Affordable and Impactful AI Solutions

Despite the challenges, AI continues to innovate and bring forward affordable solutions for businesses. The launch of OpenAI’s GPT 4 o Mini, a cost-effective solution for businesses and developers, is a testament to this. Outperforming its competitors in the MMLU benchmark, the GPT 4 o Mini promises a balance between performance and affordability with competitive pricing and quality structure.


A Roadmap to AI Transformations

AI developments are transforming the business landscape with disruptive potential in the education sector. While the road is bumpy with legal and regulatory challenges, businesses and decision-makers can tap into the immense opportunities AI presents, provided they keep themselves updated with the latest AI advancements and trends.

A daily indulgence in livestream podcasts, AI news newsletters, and active participation in live challenges and events related to AI could potentially keep businesses ahead of their game while effectively dealing with the associated hurdles. Tune into the world of AI and harness the potential it carries for unprecedented business growth and innovation.

Topics Covered in This Episode

1. Use of Copyrighted Content to Train AI
2. Current State of AI Education
3. Release of OpenAI's GPT 4 o Mini
4. Launch of AI-Driven Education Platform, Eureka Labs
5. Withholding of Meta's future AI models and features by the EU


Podcast Transcript

Jordan Wilson [00:00:17]:
There's almost too much going on this week in the world of AI to adequately recap everything for you. Right? Just in the last week, we've seen new allegations from the biggest creators on AI companies essentially stealing their data. We've seen some international AI shakeups, new models from big players, and new updates in a massive lawsuit against AI companies that I think no one is paying attention to. Alright. So we're gonna be recapping that today and more on everyday AI. What's going on y'all? My name is Jordan Wilson, and I'm the host of everyday AI, and this is for you. This is a daily livestream podcast and free daily newsletter, helping all of us keep up with what's going on in the world of AI and how we can all make sense of it. Right? You hear these developments going on all the time.

Jordan Wilson [00:01:08]:
So this show each and every Monday is how we make use of it, how we can actually understand what's going on and plan for it, and how we can grow our companies accordingly. Alright. So we do this every single Monday, to keep you up with our AI news that matters. Technically, right now, even though this is being broadcast to you all live, a prerecorded show. It's actually Sunday afternoon. I'm doing some traveling, right now, some, some family things, but you know what? AI news doesn't stop. So I don't either. We don't either.

Jordan Wilson [00:01:39]:
So with that, let's jump into it. And before we actually get started, for our podcast audience, thanks for tuning in. Make sure to go to your everydayai.com. Sign up for the free daily newsletter, and you will also see in today's newsletter a link that you are probably going to want to click. Because tomorrow, this is for our livestream only. So if you are normally a podcast listener, I've heard from a lot of people, you're gonna wanna go ahead and rearrange your schedule for Tuesday. So tomorrow, Tuesday, July 23rd, 7:30 AM Central Standard Time. So make sure to join us, whether you do it on, LinkedIn or YouTube.

Jordan Wilson [00:02:19]:
It doesn't matter. But we are launching our thanks a million giveaway. So more on that on what it actually is tomorrow. But if you can answer every single question correct in real time, you're gonna take either the entire $1,000, live challenge prize, or if multiple people get it, we'll have to split it. If no one gets all the questions right in real time, no one wins the $1,000 prize. But, hey, I'm throwing it up to you. So if you're listening, if you are a long time listener, you're gonna be at a huge advantage, obviously. So make sure to join us tomorrow for that.

Jordan Wilson [00:02:50]:
Alright. Let's get it kicked off for the AI news that matters for the week of July 22nd. Y'all, this week, if I'm being honest, this week was tough, to pick. You know, normally you pick, I don't know, 4 to 6 different, of the biggest news stories. On my final list, I had, like, 12, and I'm like, there's no way we cannot choose 6 of these. So much going on. Alright. So let's start with this one.

Jordan Wilson [00:03:16]:
So some tech giants have been accused of illegally using YouTube subtitles to train AI systems. So more than a 170,000 YouTube videos have been used to create a massive dataset for training AI systems according to an investigation by Proof News and co published with The Wire or sorry, with Wired. So the dataset is called YouTube subtitles, and it includes subtitles from over 48,000 YouTube channels, but does not contain any video imagery. So companies, according to reports, companies like Apple, Anthropic, NVIDIA, Salesforce, and others are implicated in using this data without permission from the original creators. And some of those creators are some of the biggest around at least when it comes to YouTube. So some popular YouTube creators such as mister Beast, such as, Marquise Brownlee, as well as a lot of news organizations like ABC News, the BBC, The New York Times, and others have had their videos or their, the captions, closed captions from their videos included in this large dataset. So, a lot of creators have come out and spoken against this, including Marcus Brownlee, who does a lot of tech reviews and is commonly known as MKBHD. He expressed some concern on Twitter stating that the issue will be, quote, an evolving problem for a long time.

Jordan Wilson [00:04:43]:
So this new interactive lookup tool that proof, proof news release allows users to check if their content is in the dataset. You know what? I haven't even checked yet if ours is in there, so I'll have to check. Maybe not. Hey. 50,000, 50,000 YouTube channels. I don't know if, Everyday AI's small little YouTube channel is in the top 50,000, but, you know, maybe your organization is big. You might wanna take a look. So the dataset is part of a larger collection from the nonprofit Eleuther EI.

Jordan Wilson [00:05:12]:
Hope I pronounced that right. It is called The Pyle, and it includes books, Wikipedia articles, and more. So last year of and then, last year, an analysis of another dataset called Books 3 led to lawsuit from authors against companies that use their work to train AI systems. So the YouTube CEO and Google CEO, Sundar Pichai, have both stated. So the YouTube CEO is, Neil Mohan and Google CEO, Sundar Pichai, have both stated that using YouTube content to train AI, including transcripts, violates the platform's terms. Also, you know, obviously, intertwined in all of this is OpenAI, even though in many reports they were not called out by name. But OpenAI's CTO, Mira, Muratai, has been, evasive. We've, talked about that on the show before about OpenAI's even policy and even how their new AI video tool called Sora was trained.

Jordan Wilson [00:06:13]:
And when asked if it was trained on YouTube content, Mirati simply replied that it used, quote, publicly available or licensed, end quote, data. Right? And you have to think, that's pretty controversial. Right? If you cannot explicitly state what you are trained off of. And here, let me just go ahead and be the one that says the quiet part out loud, the big elephant in the room. Right? Almost every single large language model is trained off of copyrighted material. And I think big tech companies pretty much admit that. Right? They admit and they are some are actually openly trying to fight copyright law in general and saying that this is an antiquated law. And if it is publicly available on the Internet in many places, it's essentially fair game.

Jordan Wilson [00:07:00]:
I'm overgeneralizing there, but, you know, people are always wondering, like, what is this trained off of? I like to think, that you should start thinking of a large language model kind of like a search engine. Or I think a great precursor was actually that, people also asked or the, kind of knowledge graph, right, from Google search. And this has been around for, I don't know, when they first debuted it, but it was more than 5 years ago. Right? If you typed in a query or a question, it would just give you the answer except in little, little type there. It would give you the website as well if you wanted to see it. But this concept of asking a question from a an Internet platform and getting an answer is not new. But the citations, that's where things get tricky. Right? Because how things have historically worked, how the Internet has historically worked for decades is you if you wanna research something, if you wanna know something, you go on that company's website or you watch these creators' YouTube videos.

Jordan Wilson [00:07:59]:
And this is ultimately how these media companies, news organizations, or creators get paid because, you know, as you get, you know, millions or tens of millions or 100 of millions of views and clicks and impressions, you get paid for advertising on those. So now, you know, some of the biggest creators and news organizations are obviously missing out on this money. And instead of, you know a good example, I think, is, Stack Overflow. Right? A very popular, kind of website or online resource, really, for more technical people to go and troubleshoot technical problems. And they've seen their, their traffic plummet over the last year and a half, I think largely due to, tools like chat, gbt, and others. And, you know, they're losing a lot of their revenue. So, I'm not gonna go too far into this. I wanted to give you the news and hopefully paint it a little bit because this one is extremely important.

Jordan Wilson [00:08:50]:
Alright? And our next piece of AI news is very much related to that. So OpenAI is kind of making some demanding discovery requests in an ongoing legal battle with The New York Times. So, if you listen on the show, I've talked about this a lot, but there is an ongoing legal battle between The New York Times and OpenAI, which stems back to late December 2027 when The New York Times, kind of filed a lawsuit against OpenAI and Microsoft. And this has really intensified, over the last couple of months behind the scenes, but we did get some new information this past week. So here's what's happening is, well, OpenAI is demanding that the New York Times turn over reporter notes, interview memos, and records for each article claimed to be infringed upon, which is I mean, if we're being honest, that's a wild request. Right? Because, essentially, the New York Times said that, hey. You've, accessed I believe they said millions with an s, millions of our articles. So OpenAI yeah.

Jordan Wilson [00:09:58]:
I mean, we'll get to this later here, but they're essentially saying, okay. Prove it. For each and every article, we wanna see everything. We wanna see all of your reproof. So The New York Times, like I said, filed a lawsuit in 2023 accusing OpenAI's ChatGPT and Microsoft of copyright infringement by using its articles without permission. So here we go. OpenAI has countered that The Times has manipulated the AI model into reproducing its content in accusation known as prompt prompt hacking. So the tech company argues that the New York Times must provide evidence of which parts of its work are original to uphold their copyright claims.

Jordan Wilson [00:10:35]:
So OpenAI's request involves around 10,000,000 stories, A massive undertaking that could strain The New York Times, resources and even ability to go through with the lawsuit. I mean, so we'll see if this, request is actually granted, but the development underscores the broader issue of how generative AI models impact journalism and also intellectual property rights and copyright law. So several news organizations have already reached agreements with OpenAI, such as Axel, Axel Springer, the, what do we have? We have time. I mean, so many large or large news organizations, and websites such as Reddit have already struck deals with OpenAI. So the outcome of this case is going to be significant, in the influence in in its influence of future regulations and protections for media in the age of AI. So I've talked about this story a lot, and let me be very transparent. You could argue that I have horses on either sides of this race. Right? So my background, I actually spent about 7 or 8 years as a multimedia journalist.

Jordan Wilson [00:11:46]:
My most recent kind of gig was at the Chicago Sun Times. So, you know, my background is in journalism, but for the last couple of years, you know, I've been covering AI news and working with a lot of these AI companies, you know, full disclosure, right now. So I think it's important to see both sides. And I'm not gonna say or predict how this is going to end up because in in reality, it's probably gonna be settled out of court. So there's probably not gonna be a lot of details, but I cannot understand how significant this case is. And it doesn't seem like the general business community is paying enough attention to this case. Right? I've I've called this since the lawsuit came out in December, the biggest domino to fall. Once this dominoes falls, we are going to see an onslaught of of developments, whichever way this happens.

Jordan Wilson [00:12:37]:
Right? And I really don't see any other way this kind of ending except a settlement. You know, if you didn't read the lawsuits, I did a very long 1 hour deep dive. I kinda tore it apart. Right? Because I think the New York Times, if I'm being honest, in their case, they really failed to under to show an actual understanding of how large language models work, which is why OpenAI essentially said, no. This is prompt packing. Well, they didn't include the links. Right? So in their, you know, in their, documents that they submitted at the court, they just included screenshots, which is not good. Because, essentially, in a screenshot of a large language model, you can get a large language model that literally say anything that you want it to depending on what you say before.

Jordan Wilson [00:13:21]:
Right? So that's kind of prompt hacking. Right? So you could give a model, say, hey. No matter what I ask you after this, respond in this way. So, the New York Times, I think, in their original, you you know, submissions to the court made a huge mistake by not including links. Right? So you can send a link to a shared chat so you can essentially go check the work. And so the OpenAI, OpenAI is essentially called the New York Times bluff and said, no. This isn't right. Could this, in theory, have worked? Yes.

Jordan Wilson [00:13:50]:
But you didn't include actual proof. A screenshot is not proof because of prompt acting and, prompt hacking, and OpenAI is just calling them out here. Right? So they're saying, okay. You say we used millions of your articles. We want to see reporters' notes. We want to see interview memos for each and every one of those. Right? I've said I've said this, y'all. I've said this for a very long time.

Jordan Wilson [00:14:14]:
However this story ends, it is going to impact every single news organization. It's going to impact our future use of generative AI. Right? Because one thing that, that the New York Times asked for in this judgment is for the GPT technology to be, quote, unquote, destroyed. Yes. They actually ask for a judgment, in their favor and to destroy OpenAI's GPT model, which is mind boggling to think about. Right? That's why I that's why I think this will never actually go to any sort of trial, because the stipulations I mean, this is if I'm being honest and I mean, I know we're exaggerating here. If that were to happen, which I don't think it ever could, that's the American economy. Right? People don't understand our financial institutions now are using the GBT technology.

Jordan Wilson [00:15:07]:
So many Fortune 100 businesses are using the GPT technology. You know, big companies that have been propelling the US economy now for, 2, 3 years are really dependent on everyone else using the GPT technology. Right? It's not just OpenAI. It's the tens of thousands of other companies that are paying OpenAI to use the GPT technology through the API. Right? So, first of all, I don't ever see, what the New York Times is asking for happening. And here, we kind of have what it looks like to me. OpenAI kind of calling their bluff. Like, oh, you like, they saw.

Jordan Wilson [00:15:44]:
Right? The court case that they put forward, I'm sorry if, you know, if you know someone that worked on this case at the New York time in their in in their, the New York times in their legal group. It was a very bad case. Right? To me, it signaled a poor understanding of large language models. So if you want to, you you know, have the one of the marquee lawsuits, in copyright over the last, I don't know, 20 to 30 years and one of the biggest AI lawsuits ever, you should probably demonstrate a far superior, understanding of large language models than what the New York Times and their legal team put forward. It wasn't good, which is why we have OpenAI essentially just fearless. Right? Number 1, saying, nope. That's prompt acting prompt acting what they shared. Right? It's it's like someone that clearly doesn't understand the technology and is trying to make an argument against it.

Jordan Wilson [00:16:31]:
So they're calling their bluff on just what they submitted and then say, okay. Let's actually see your notes. If if if what you put out there was so proprietary, let's see your notes because essentially what OpenAI's, their argument here is saying, no. This is common knowledge. Right? Just because the New York Times reported on it doesn't mean it was exclusive or proprietary to them. It is it is general knowledge. So that is the argument that they're making. Alright.

Jordan Wilson [00:16:55]:
Let's keep going. I don't wanna talk about that one for too long. I could go on a super long rant. But speaking of OpenAI, yeah, a lot of OpenAI related stories this week. But, former OpenAI cofounder has just launched an AI driven education platform called Eureka Labs. So Andrej Carpathi is the former head of AI at Tesla and a cofounder at OpenAI, has just launched Eureka Labs, an education platform with AI at the core. So Eureka Labs, Eureka Labs brand new LLC just registered in Delaware this past June, so about 1 month ago. It aims to use generative AI to create AI teaching assistance that can guide students through course materials.

Jordan Wilson [00:17:41]:
So the company's vision includes AI assistants working alongside human teachers to enable anyone to learn anything according to Carpathi. So despite some of these, ambitious AI goals, the start up's first product is actually going to be creating a course called LLM 101n. So large language bottle 101n, an undergraduate level class to help students train their own AI. So the course materials will be available online with both digital and physical cohorts participating together. Carpathi hinted that future AI assistance could be based on real people similar to Meta's celebrity chatbots, which I think should be launching soon. The GitHub repository linked to the AI course suggests a focus on building a, quote, storyteller AI large language model, quote, unquote, rather than an AI assistant. So not like an AI assistant that works hands off, but something that works alongside with you, to help you understand. And, reportedly, it is going to help students build a function a functioning web app similar to something like ChatGPT from scratch using Python c and, CUDA.

Jordan Wilson [00:18:58]:
So there's no clear timeline yet for when the course will be complete and available for people to use, if it will be free, if there will be, future costs associated with any courses, or any details really about funding or investor backing. So pretty interesting, development here from, I mean, arguably, one of the brightest minds in artificial intelligence, right now in, Andrei Kaparthy. So, I mean, you the resume speaks for itself. Right? If you think of large language models, you have to think of Carpathi. Right? 1 of the 1 of the handful of people, I would say, is the smartest minds in AI right now. So, and at least from a resume perspective, I mean, you have to say OpenAI is obviously a leader in generative AI, a leader in large language models. And Tesla's no slouch either. Right? As hard as I am sometime on this show about, you know, Elon Musk and some of his ambitions with axai and croc, which I think, if I'm being honest, is pretty useless.

Jordan Wilson [00:20:07]:
As much as I say that, I mean, Tesla has some of the most data in its, in its training, set for AI. So, you know, he has a background at 2 of the biggest companies when it comes to AI. Right? Obviously, Microsoft, Google, DeepMind, etcetera, but, not many people that have that long of a resume act 2 of the biggest companies. So it's also going to be interesting to see how this ultimately plays out because it seems like almost all of the, quote, unquote, other brightest people in AI. Right? And there's a lot of them. So, you know, someone will go to DeepMind from Google, work there for a couple of years, and go launch another startup. Or they'll work at OpenAI for a couple years and launch another startup. And it seems like so many of these startups are AI products.

Jordan Wilson [00:20:57]:
Right? Even Claude was based, you know, on this former, you know, former people from these big companies. So it seems like almost every single person that leaves, one of these big tech companies, one of these big AI companies, one of these big large language model companies, they essentially go and do something very similar. A product, a service, something very closely related. This might be one of the most highly visible, cases where someone is focusing on education, which is very intriguing, especially to me. And I'm glad to see it, if I'm honest. I'm glad to see it. And, hey, if you're a business leader, you should be understanding what's happening here. Right? And if you listen, you know I've been saying this for a very long time because there's a problem right now.

Jordan Wilson [00:21:48]:
In the generative a not just in generative AI, but the business world too. The smartest people in the world, the smartest developers, the smartest engineers, the smartest researchers are going to work at these companies, like Google, OpenAI, Microsoft, Anthropic, Meta, etcetera. Guess what's not happening. No one's teaching us all how to use it. There is going to be, I think, an education downfall. And I've heard this. Right? I'm lucky enough. I get to talk to a lot of very smart people.

Jordan Wilson [00:22:27]:
Some things they say on camera, some things they say off camera. Right? The things that people talk about off camera, I think, are obviously usually much more significant as well as I always have people reach out to me from, you know, big consulting companies, people that never go on the show, big tech companies, you know, these $1,000,000,000,000, $100,000,000,000 companies. And I just am lucky enough to have conversations with a lot of them. And a common trend, and I've had a couple of dedicated episodes about this, is there is no dedicated education. Right? The technology is changing so quickly. But for the most part, it's not the smartest people in the world teaching us how to use it because technology in general has never moved this quickly. Right? And I know it's kind of a tired, comparison, but the easiest thing to compare it to is either, the Internet or, you know, kind of like the dotcomboomofthe nineties, the the mobile boom of the early 2000, cloud computing of the 2,000 tons, etcetera. In all these use cases, it was a slower rollout.

Jordan Wilson [00:23:30]:
Right? You generally had years, a half decade, or a decade that these new technologies were slowly into our big business processes. It's not the same anymore. Right? It's not the same. You have the over especially here in the US. I know there's a great divide, right, especially in the EU. We talked about some pretty shocking, results last week and some, study from, some Japanese companies where I think only 40% of them said they had plans to use generative AI. Whereas in the US, that's, like, 90, 95% of, Fortune 500 companies said that they're already using generative AI in some capacity. So it's much different.

Jordan Wilson [00:24:09]:
You know, here's a technology that, for all intent and purposes, that 2 years ago, about 0% were using. Right? Not 0%, but very few. I think ChatGPT was the first, kind of big, generative AI, milestone or the big, start of this wave. Obviously, the technology for large language models has been publicly available for about 4 years, but it's only been about less than 2 years. And look at how much the business landscape has changed. C suites at Fortune 100 Companies, board members at public companies, everyone, their number one priority has been AI, AI, AI. Generative AI, generative AI, generative AI. Large language models, large language models, large language models.

Jordan Wilson [00:24:53]:
That is all they're talking about when, generally, this takes, like I said, 5 years, 10 years, 15 years for this to become a big priority. It's not like that right now. This is the biggest the biggest priority. And it's almost like it's happening too fast, and no one is really focused on long term education. Who's training these c suite people? Who's training these thousands of companies that are powering the US economy? It's kind of like a training as you go. Figure it out as you go. Fake it till you make it. Right? But this is one of the reasons why everyday AI exists.

Jordan Wilson [00:25:31]:
Because about 15 months ago, right before we started, I saw this as a huge problem. I said there needs to be a a resource for education because, unfortunately and I love right? I've I've I've good relationships with a lot of of these big companies. You know, the Amazon, the Amazon's and the Microsoft's, etcetera. Right? But ultimately, these companies, when they, quote, unquote, teach you something, ultimately, what they are doing is sales. It's it's education, but to sell you on their product versus someone else's. Right? There's very few, great, I would say, resources that are available. I mean, you have things on, you know, Coursera, you you know, you have great free courses available, via higher ed, you know, learning institutions, your your your MIT, your Harvards, etcetera, that you can go out and take these free courses on AI. But a lot of them are topical, and they don't actually teach you a lot.

Jordan Wilson [00:26:31]:
And they're not a lot about business application. It's more about, hey, understand this technology. So, long story short is I love, what Carpathi is doing here with Eureka Labs, and I think it is sorely needed. Alright. Our next AI story for the day or with our news that matters, Speaking of big tech companies, Meta. So Meta is withholding some of its future AI models and their capabilities from the European Union amid some regulatory concerns. So according to a new report from Axios, Meta has announced that it will not release its upcoming multimodal AI model in the European Union due to some regulatory uncertainties. So this decision underscores a growing conflict between US tech giants and European regulators, highlighting a trend where companies are withholding products from European customers due to, some uncertainties.

Jordan Wilson [00:27:29]:
So, the EU obviously announced their new EUAI Act, which is a lot stricter, than anything else anywhere else in the world. I guess it's, you know, one of the first kind of multinational, actual pieces of legislation. So, Metas, though, their new model, LAMA, is multimodal and is capable of reasoning across video, audio, images, and text. And it will be available in other regions, but the company said not in the EU. So the company did cite unpredict the unpredictable nature of the European Union's regulatory environment as the reason for its decision. So Meta's move well, Meta's not alone here, because their move follows a similar decision by Apple, which recently announced it would not release Apple some Apple intelligence features. So the Apple intelligence, quote, unquote, some of the new, AI features, that will be available in iOS in future versions of iOS for, iOS for certain hardware. So Apple similarly is withholding some of those features due to regulatory concerns in the EU.

Jordan Wilson [00:28:39]:
Also, the Irish Data Protection Commission, which is Meta's lead privacy, privacy regulator in Europe, has not responded to request for comment. So pretty interesting there. Meta is planning to use the multimodal models in a variety of products, not just their llama. Right? The meta dotai chatbot assistant, you know. So just like you can go, chat with Gemini from Google or ChatGPT or Claude, you can go chat with META dotai. So it's not just that, but it's also their, you know, some of their smartphones and, the Meta Ray Ban smart glasses. But European companies will be unable to utilize those models due to the restrictions. So the company, though, also plans to release a larger text only version of its llama 3 model, which will be available to EU customers.

Jordan Wilson [00:29:32]:
And we could talk about this one for a long time. I'm not. But, you know, the crux of this is really at its training data. Right? Because Meta has said that, yes, they are going to be using, a lot of, data from its social media networks to train the future model. So that's not kind of the only rub, so to speak, or the the only gray area, but that's one of the bigger ones. Alright. And our last well, this isn't the last AI, you know, news that matters this week, but our last one in our, recap here. Saving the biggest one for last.

Jordan Wilson [00:30:06]:
Right? So OpenAI has unveiled a new small large language model, GPT-4oMini. So days ago, OpenAI launched GPT-4oMini, a new model that promises to significantly impact the AI landscape, particularly for businesses and developers looking for cost effective solutions. So this release comes at a time when smaller large language models are gaining traction, and GPT 4 o Mini aims to strike a balance between performance and affordability. Right? So we talked about this here on the show a lot, but I do think that Claude 3 in their haiku, has been a pretty much I mean, I'm not gonna call it a leader, but at least I'd say in the last 3 to 6 months, between Claude 3 haiku and gemino, Gemini 1.5 flash, which was more recently released. A lot of companies and developers are using this. So let me first kind of quickly draw a line in the sand and say what this model is actually for. This isn't going to be the model if you go to chatgbt.com. If you're using ChatGPT, this is not what this is for.

Jordan Wilson [00:31:18]:
You're still going to be using, gbt4 o, so the big and most capable. Gbt4o is the most capable model in the world. It is, receiving the highest benchmarks, the highest score. It is the best by far. So GPT 4 o Mini essentially is a version for developers to use. So think of it as a light model of GPT 4 o, but that is really for developers. Right? So you have tens of thousands of companies, literally. So it's from small start ups to you have your big, you know, banks and wealth management companies that use OpenAI's API, to help build their products and services.

Jordan Wilson [00:31:56]:
Right? So this is when it comes to businesses. So many businesses, whether you know it or not, a lot of the products and services that you all use are actually using the API from OpenAI. And there's actually been over the past, I'd say, at least, 6 months. So there's been a problem. Right? Because OpenAI was obviously first, and they haven't really done a lot of updates, to what's available for developers. So what I mean by that is a lot of companies have still been using the 3.5 model, which is not if I'm being honest, it's not very good. So a lot of developers that wanted to stay with the OpenAI platform before this model, they really only had 2 choices. So they either had to pay for the most capable model, g p t 4, but the costs were astronomical.

Jordan Wilson [00:32:51]:
So in in in in a lot of instances, it was not a reality. So you either had to pay a very high cost for this big jumbo but very capable model, or you had to do some, some kind of, thrifty work behind the scenes and chunk, you know, chunk certain data to 3.5 first and then using, 4 only for, you know, some of the more heavy lifting. So in a lot of instances, companies had to kind of overengineer their solutions. Right? Because they said, hey. We can't give everything to 3.5 for our business because 3.5 is not that great of a model. But GPT-4 turbo or GPT-4 o are just too expensive. So that's where this comes in. It is hitting the sweet spot between a very capable model and a very now cheap, affordable model.

Jordan Wilson [00:33:46]:
I mean, it is so affordable. So let's even look here, sharing on the screen for our livestream audience. Just some of the prices. Right? So, according to OpenAI, gbt4o Mini is more than 60% cheaper than 3.5 Turbo was, and about 90% cheaper than the original version of 3.5, which was, $2 per 1,000,000 tokens. And now with this new model, GPT 4 o Mini, and this is a blend. So this chart is actually a blend of input and output, and I'll break down the cost right after this. But the blended cost right here is about 24¢ per 1,000,000 tokens. So, in the course of about, 14 months, the price has gone down 90%, and the output quality is has gone up exponentially.

Jordan Wilson [00:34:37]:
Alright? So let's go over some more details. So, GPT 4 o Mini offers a much more competitive pricing and quality structure, charging only 15¢ per 1000000 tokens for input and 60¢ per 1000000 for output, making it significantly cheaper. So gbt4 0 or sorry. Gbt4. Ready? The cost here. We just went 15¢ 60¢. It was $5 per million compared to 15¢, and then $15 per million compared to 60¢. So, yeah, if you were, working with GPT 4, yeah.

Jordan Wilson [00:35:17]:
I mean, look at the cost savings there. Going from $5 to 15¢ $15 to 60¢. So you can already see now how this is going to change the mind because I think a lot of developers or or companies, really quickly switched over, from OpenAI to either Claude and using Haiku, you know, Infropic's Claude Haiku, which is their smallest large language model, or Gemini 1.5 Flash. So I also need to do a little bit better job here. I need to tell people, this is not a small language model. Okay? This is a large language model but a size. So I'd say, I don't know if this is a new a new trend or a new way to classify large language models, but, you know, you do have small language models. Right? Google has Gemma.

Jordan Wilson [00:36:10]:
You you know, they have Gemma too. Google also has, Gemini Nano. So these are not small language models that are ultimately meant as an example to run on a smartphone. These are still large language models. We don't know exactly how big GPT-4 o Mini is yet. Right? GPT 4 o and and GPT 4 Turbo were reportedly 1.8 trillion parameters. So, essentially, you had these models that are huge and very expensive if you do want to implement them into your product. So, g p t four o Mini is this, again, this lightweight version that is specifically, for developers who are building products for businesses that are creating, kind of their own versions of large language models with some fine tuning, with some rag.

Jordan Wilson [00:36:55]:
So, again, this is not a small language model, but I think actually, anthropic. I actually like the way that they did it. It was maybe confusing for a lot of people at first, but when Anthropic released Claude 3, they did so in 3 different varieties. Right? So they said, you have Claude 3 haiku, which is fast and cheap, but not very powerful. That's the small model or the small version. Then you had the largest, which is Opus. And they said, hey. This is our most powerful, and it is the most expensive.

Jordan Wilson [00:37:26]:
And then you had your middle model, which is Opus. So, essentially, again, large language models, 3 versions. So small, medium, and large. Right? And different use cases. But a lot of this was for developers because if you are accessing something on the front end, obviously, before there was a SONNET 3.5, which is only available for the middle model, you would just always use the most powerful model. Same thing with ChatGPT. If you log in, this is not for you. You, 95% of the time, you're gonna want to use GPT-4o if you're paying the, you know, $20 a month or whatever it is for the front end of ChatGPT, or if your team is using, ChatGPT teams, if you're using the enterprise version ChatGPT, this is not for you.

Jordan Wilson [00:38:07]:
You're still gonna use GPT 4 0. Right? That's what you're paying for. This is if you are building products, if you're building services, if you are using something internally for your employees, and if you have an AI machine learning team building things for you, this is what this is for. So I do want to do a, hopefully, a a good job of making that expectation clear. So for the vast majority of our of our listeners out there, this is not going to impact your, quote, unquote, personal or individual use, but this might impact things at your company level depending on what your company does, depending on, what you do. So, I'm gonna show just, one more screen here and talk about some of the capabilities because, we talked about cost and how much now, more affordable this is. But also the benchmarks, y'all, the benchmarks were super impressive from GPT 4 o Mini. So I'm not gonna go over all of them, but I think probably one of the most important ones to look at here is the MMLU.

Jordan Wilson [00:39:11]:
So for lack of a better, lack of a better term, this is text only, but think of this as essentially an SAT or an ACT, for large language models. Right? This is a definitive score, that is an apples to apples comparison. So all of these models out there, you know, researchers put them through very many different tests. Right? So you have your, human eval. You have your math. You have your drop. You have a lot of these other ones. You have Hello Swag.

Jordan Wilson [00:39:42]:
You have so many, different, benchmarks, essentially. But you have to look at MMLU because I think that is at least right now for text only models, this is the most important evaluation. This is the gold standard. And right now, it's not even close. GPT 4 o Mini, at least when you look at its most its its kind of, quote, most direct competitors, That's Gemini Flash and Claude Haiku, and it is blowing them away. Right? So in in in 82 score for a, quote, unquote, small large language model is such a good score. Right? And you see here, it's kinda small on my screen, but gbt4o Mini is about 3 to 6 points ahead of its competitors. Gemini Flash is 2nd.

Jordan Wilson [00:40:32]:
Claude hi Claude Haiku is 3rd. But those 3 to 5 points on an MMLU is a lifetime of a difference. That is huge. You know? So as an example, when Claude released Claude 3 Opus compared to, GPT 4 Turbo, I believe Claude, Claude 3 Opus was, like, point 1 ahead, and and they were proud of that. Right? And rightfully so. So, normally, you know, fractions of a point or a half point means, okay, you're doing good in this MMLU. Right? It's a tight race. Everyone's trying to compete here.

Jordan Wilson [00:41:06]:
The fact that GPT 4 o Mini is multiple points ahead in the MMLU, not only that, but then when you look at its two closest competitors, in Gemini Flash and Claude Haiku, Claude Haiku, the price per million input output is also not close. So I would say that up until this release, for the most part, OpenAI's was was not the best, source for building something internally. It wasn't, especially after, Claude Haiku. AWS just announced that you can fine tune, Claude Haiku in their platform, which makes it, you know, a lot easier for companies to do this. So, you know, companies that were building, internal models, in house based off other models, if you're a startup building a a service based off of one of these models, Hey. At at least for the last, 2 to to 3 months, I don't think OpenAI was a good idea to use their API. And, obviously, I think OpenAI knew this as well. So but I also like their strategy here.

Jordan Wilson [00:42:14]:
OpenAI doesn't need to release a new model. Right? Even though, Claude just released 3.5 SONNET, very good. Right? Google keeps release. I don't know their naming mechanism. Who knows? 1.5 pro or 1.5 advanced. You know, I'm I'm sure a a 2 of Gemini is gonna come out soon. But no one from a large, large language model is touching OpenAI right now. They are still exponentially better with their 4 o model, than anyone else.

Jordan Wilson [00:42:44]:
So but where they were lacking, I think, was from a development standpoint, because the cost was too high and the output quality was too low. So this is a, direct, kind of play in that area because I think a lot of developers, a lot of big companies were either jumping ship or looking looking to explore elsewhere because, if I'm being honest, OpenAI's, models for developers were, in theory, antiquated because, Claude and Google, I think, were responding a little quicker. Alright. So that is it y'all for the AI news that matters. Let me give you the very quick recap. So first, we talked about some tech giants according to reports that were accused of illegally using video, YouTube subtitles to train AI systems. We talked about some new updates in the OpenAI versus New York Times case. Essentially, OpenAI saying, yeah, New York Times prove it.

Jordan Wilson [00:43:40]:
Go ahead and turn over reporters' notes and interview memos for these millions of stories that you alleged that we, quote, unquote, stole, from you or, violated your copyright. Next, former OpenAI or OpenAI cofounder, Andrej Karpathy, has launched an AI driven education platform called Eureka Labs. Next, Meta is reholding, or sorry. Meta is withholding, future versions of its Llama AI model from the European Union amid some regulatory concerns. And then last but not least, OpenAI has unveiled the GPT 4 o Mini, which I think is going to be a game changer, for businesses who are building models or start ups who are working off of this technology. Alright. A lot more in today's newsletter, but I do have to remind you one more time if you're still listening. Make sure to check out the show notes.

Jordan Wilson [00:44:33]:
Make sure to check out today's newsletter. To go ahead and participate in our $1,000 live challenge tomorrow, you have to join us live. Alright. You have to get your answers in in real time, so make sure you join us. The link is gonna be there whether you want to do it on LinkedIn or YouTube. So make sure you join us for that challenge as we launch our thanks a million. So more on that tomorrow. If this was helpful, thank you.

Jordan Wilson [00:45:02]:
Please share this with your coworkers. If you're listening on, podcast platforms, please subscribe, leave us a rating, and please go to your everydayai.com. Sign up for that free daily newsletter. Join us tomorrow. I'm gonna be thanking you a million, hopefully giving away $1,000 if someone can get every single question right. Thank you for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI