Ep 379: OpenAI not profitable until 2029? Microsoft goes all in on AI/health & more – AI News That Matters

Resources:

Join the discussion: Ask Jordan questions on AI


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Google and the Future of AI Dominance

Despite Google's impressive control of over 90% of the search market share, experts expect that to fall below 50% by 2025. This comes amid concerns about the tech giant's data and AI infrastructure and the perceived under-utilization of its access to training data for AI models, such as Gemini. However, Google's endeavors face regulatory challenges as the U.S. Department of Justice eyes interventions aimed at enhancing search transparency and limiting data gathering tactics.

Tesla's Trials in AI Vehicle Development

Tesla's AI advancements extend to ambitious attempts like the CyberCab and RoboVan prototypes, envisaged to be on our streets by 2027. Despite their AI-driven ambitions, Tesla's projects miss investor expectations leading to a tailspin in their shares, spurred by competition from companies like Waymo.

Microsoft Leads AI Innovation in Healthcare

While some companies dip, others ascend with Microsoft unveiling AI tools designed to alleviate administrative burdens for healthcare workers. The company's efforts demonstrate AI's potential to address pressing societal issues, including healthcare worker burnout. The innovations span medical imaging models, healthcare agent services, and automated documentation for nurses.

Debates on AI Reasoning Abilities

A recent study steers conversations towards AI and reasoning capabilities. With a paradigm shift from traditional pattern matching, AI researchers explore the models' potentials for logic and reasoning. As these AI models approach AGI stages, debates on such capabilities heighten. However, questions persist as leading AI models exhibit fluctuations in performance and minimal progress as task complexity increases.

OpenAI’s Financial Future

OpenAI finds itself treading a fine balance between financial sustainability and advancing AI research. Despite generating billions annually, the company's substantial spending projections raise concerns about long-term viability. However, major investors like Microsoft, with a $13 billion stake, seem undeterred.

AI in the Fight Against Election Misinformation

Misuse of AI to generate fake content threatens the integrity of democratic processes. The emergence of AI-generated misinformation targeting US elections underlines the need for vigilance and proactive countermeasures in protecting information integrity.

Human Reasoning vs. AI

AI’s capacity to mimic human cognition continues to polarize researchers. A study by Apple challenges the idea that AI, despite its vast data processing abilities, can replicate human cognitive processes.

Exploring AI Education

With AI’s pervasiveness increasing, education about its use becomes imperative. Free AI courses like ChatGPT offer high value and interactivity, ensuring broader understanding and useful application of AI technologies.

Recognizing AI Pioneers

As AI continues to reshape industries, pioneers in the field are earning global recognition. The awarding of the Nobel Prize in Physics to AI researchers underscores the significance of AI evolution and its impact on human lives.

The AI landscape continues to evolve rapidly, offering both challenges and opportunities. As we move into 2025, decision-makers must monitor unfolding developments to leverage AI's transformative power strategically.

Topics Covered in This Episode

1. Google's Role in Advertising and AI
2. Tesla's Struggles with AI-powered Vehicles
3. Microsoft's Advances in Healthcare AI
4. OpenAI's Financial Outlook and its Role in the AI Industry
5. Impact of AI Models on Logic and Reasoning

Podcast Transcript


Jordan Wilson [00:00:17]:
You may think we had a quiet week in AI because we didn't get a new large language model update, but that is not the case because, reportedly, OpenAI may not be profitable until 2029. We got a lot of new AI powered innovation from Tesla, and Microsoft is going all in on AI and health. So, yeah, even though we didn't, I think for the first time in seemingly months, get huge updates from one of the big large language bottle makers, AI news did not stop, and neither do we. What's going on y'all? My name is Jordan Wilson, and I'm the host of Everyday AI. Welcome. Every single Monday, we bring you the AI news that matters so you don't have to spend hours every single day trying to keep up with what's happening in the world of AI and how it might impact your company or career. You can just do that with us every single Monday on the AI news that matters. So if you're new here, welcome to everyday AI.

Jordan Wilson [00:01:20]:
This is a daily livestream podcast and free daily newsletter helping us all do just that. Keep up, not just keep up actually, but how we can use all of this information to get ahead. So if you are brand new here, thank you for joining us. Make sure, if you haven't already, go to your everydayai.com. That website should be your best friend y'all. Because on our website, it is literally the largest free unbiased source of everything generative AI on the web. You can go back. We have nearly now 400 episodes.

Jordan Wilson [00:01:54]:
Whatever, you know, if you care about marketing, sales, non profits, whatever it is, we have categories on our website. You can go learn from the leading experts in the world. So if you haven't already, please make sure to go check that out. Alright. Enough chitchat y'all. Good morning to our livestream audience joining us. So everyone, Jay and Marie and Kurt, Michael, Kathleen, Brian, everyone else, sabbatical. Thank you for joining us.

Jordan Wilson [00:02:17]:
Let's just get straight into it y'all. A lot going on in the world of AI. Let's get caught up. So OpenAI is reportedly facing a $44,000,000,000 in losses before profitability in 2029. So that is according to some recent reports from the information, that says OpenAI is projected to incur $44,000,000,000 in losses, before reaching profitability in 20 29. So, according to reports, OpenAI is attributing that to substantial, expenditures on training AI models, employee salary salaries, and data acquisition. So the report estimates that OpenAI could generate around a $100,000,000,000 in revenue, but its current spending is approximately $7,000,000,000 on model training and $1,500,000,000 on staffing annually. Also, Microsoft, a key investor with a $13,000,000,000 stake in OpenAI, is expected to take a 20% cut of the company's revenues, which may obviously impact OpenAI's profitability.

Jordan Wilson [00:03:35]:
So according to the report, OpenAI is currently generating about $2,000,000,000 from ChatGPT and $1,000,000,000 from access fees for large language models totaling an annual revenue somewhere between 3a half to 4 and a, 4 and a half $1,000,000,000. But market analysts suggest that OpenAI will need additional funding to subs to sustain its operations following that recent $6,600,000,000 fundraising round we talked about 2 weeks ago that pegged OpenAI at a, valuation of a $157,000,000,000. So there's been a lot of concerns, related to the long term viability of AI startups as investor interest may be kind of peaking, potentially leading to takeovers or increased scrutiny of financial practices. Seems like, if I'm being honest, a lot like, I don't think a lot of people aside from those that work at, you know, private equity, venture capital firms have really worried, about the longs long term sustainability of companies like OpenAI and Anthropic because let's be honest, when any of those companies want to raise money, they can very easily raise 1,000,000,000 of dollars on command. So, you know, people are like, oh, you know, this means OpenAI is gonna shut down. They're facing financial, strain. No. That's not what it means.

Jordan Wilson [00:05:02]:
So let's just call that, exactly what it is. This is, you know, typical. This is typical going from a a start up to profitability. So, you know, don't freak out over these recent reports that says OpenAI may not be profitable until, you know, 2029. That doesn't mean you shouldn't be using their technology. You absolutely should be using their technology or a profit Claude or Google Gemini or whatever is best for your business regardless of if these large language models are making money themselves. Because right now, we've talked about this all along. The biggest race for these companies is for you.

Jordan Wilson [00:05:36]:
It is for us. Right? That is what they ultimately care about. They care about attracting customers, retaining customers, and, you know, profit. They'll figure that part out. But right now, it's a it's a great time to be a consumer, a business owner, you you know, to take advantage of all that these large language models have to offer. Alright. Let's keep it going y'all. Here's a big one.

Jordan Wilson [00:06:02]:
So the Nobel Prize in Physics has been awarded to 2 AI researchers for their groundbreaking work in artificial intelligence. So the recent awarding of the Nobel Prize in Physics has gone to 2 prominent research, marking a significant milestone in the field of artificial intelligence. So John j Hopfield, for who's 91, and Geoffrey Hinton, who's 76, received the Nobel Prize for their pioneering research that laid the groundwork for modern artificial intelligence. So the majority of their work goes back to the eighties and before where the 2 have been instrumental in developing machine learning, a process where computers learn from vast amounts of data to perform tasks such as, I mean, anything AI can do, diagnosing disease, personalizing entertainment recommendations, doing your homework, whatever AI does. So Hotfield is a professor at Princeton University and invented the Hop Field Network in 1982, which was a neural network capable of mimicking certain brain functions and recalling information from partial data. Hinton is often called the godfather of AI based in Toronto, utilized Hotfields' inventions to create networks that classify and recognize patterns in large datasets, which can be applied in various fields, including image recognition. So, most of y'all have probably heard of Geoffrey Hinton. It was very, famous when he departed from Google, last May in May 2023 and was driven by his desire to openly discuss the potential risks associated with AI technology, emphasizing concerns over its misuse by malicious actors.

Jordan Wilson [00:07:53]:
And that has definitely been going on, and we have another, news story, today on that exact same topic. I don't know about y'all. Today today is one of those Mondays, I need my coffee a little stronger. Little sleepy today. I don't know about you guys. Yeah. And FYI, we are an unedited, unscripted podcast. We're just bringing it to you real.

Jordan Wilson [00:08:15]:
Alright. Our next piece of AI news, could Google be in big trouble with the Department of Justice? Maybe. So the US Department of Justice has some recent proposals to address Google's dominance in online search, and that could significantly significantly impact the tech giant's profitability and also its advancements in artificial intelligence. So the US Department of Justice, the DOJ, is considering remedies that may force Google to divest key parts of its business, such as the Chrome browser and the Android operating system, which are perceived by feds to support its monopoly in online search. So other potential actions include restricting Google from collecting sensitive user data, requiring transparency in search results, and allowing websites to opt out of their content being used for AI training. That is the big one y'all. So following the Department of Justice announcement, Alphabet stock, Alphabet, the parent company of Google, of course, dropped, 1.5%, indicating at least a little investor concern over implications of these remedies and on Google's long term AI plans. So, analysts are also warning about this that these changes could obviously weaken, Google's revenue and also provide more opportunities for competitors like DuckDuckGo and Microsoft Bing and, I guess, Yahoo.

Jordan Wilson [00:09:52]:
Does anyone use Yahoo search anymore? So especially as Google's US search ad market share is projected to dip below 50% by 2025. So, yeah, that's not their search share, which is, still in the nineties, but their search ad market share is expected to dip below 50%, which is kinda wild. So, yeah, I think there's a lot of reasons why this is a this is actually pretty big news. I would say the biggest is even though Google has been, I think, personally, unable to capitalize on the amount of data that it has. And let's be honest, just the head start. Right? The I'm I'm I mean, the transformer infrastructure was developed or created by Google researchers, yet Google has been fairly far behind kind of in the, large language model race, at least for front end products. I think Google and their Google Gemini has a great back end for developers. But if you're using the front end of, you know, Google Gemini, I think it's very far behind even, Anthropic and, you know, the much smaller start ups, Anthropic and OpenAI, even though Google has better, more direct, more up to date access to all training data via its Google bot.

Jordan Wilson [00:11:16]:
Right? So, if your company has a website, obviously, it wants Google and has wanted Google, to scrape it and to access that information for many years. And in now doing so, you also, by default, you know, when you're saying, hey, Google. Come to my website. You're also now by default saying, hey, Google. Go ahead. You can also scrape all that content and use it for your Gemini models and all your other AI. So, I think that part is particularly interesting. I'm not even so much focused on the DOJ trying to have, or or force Google to divest its Google Chrome browser and its Android operating system as much as I am on that AI training data piece.

Jordan Wilson [00:12:00]:
Right? Because that is what I think has been Google's biggest, advantage that I think they haven't really taken advantage of because for whatever reason, it seems like Google Gemini has this almost huge issue, to bring up to date relevant information to its Gemini product via that kind of Google integration, which I know is kind of wild. Right? Google knows everything up to the second, but Gemini, its front end chatbot kind of struggles bringing in recent and relevant information. Or I don't know. I could I could be alone in that live stream audience. What do you think? Is anyone out there using and loving Gemini on the front end? I haven't met too many people. Love love the new notebook LM. Don't get me wrong. But not not a huge fan of, what they have going on on the front end for Gemini chat.

Jordan Wilson [00:12:52]:
Alright. Our next piece of AI news, speaking of things that I'm not personally the biggest fan of, Tesla's cyber cab unveiling has failed to impress investors as their shares dropped immediately 9% and are now down more than 11% since the announcement. So Tesla's recent reveal of its AI powered cyber cab robo taxi concept has left investors very underwhelmed, resulting, like I said, in an almost immediate 9% drop in, shareholder value and 11% since the announcement. So if you didn't miss this, this was late last week, Tesla had an event branded as We Robot. Yeah. You don't wanna get Will Smith mad. You know? Don't don't don't get too close to Will Smith's territory there. So the new cyber cab robo taxi, concept.

Jordan Wilson [00:13:50]:
It's a concept. It showcased a 2 seater self driving vehicle that lacks conditional control, lacks traditional controls such as a steering wheel and pedals. And, instead, essentially, the AI self driving does you know, takes care of that for you. So, CEO Elon Musk announced plans to introduce the CyberCab by 2027 with an anticipated price tag under $30,000, but provided no specifics on manufacturing locations or really specifics on anything. So before everyone out there collectively loses their marbles on that price tag and this great innovation, it's also worth noting Tesla's recent history with new vehicle releases. So as an example, Elon famously or maybe infamously announced the Cybertruck in 2019 with a $39,000 price tag and a 2021 delivery date. We all know that didn't happen. Right? So the Tesla truck actually shipped out at the very end of 2023.

Jordan Wilson [00:15:02]:
And now today, the price tag is 95,000. So right after it came out, it actually oh, it's actually about 6 figures, not $39,000. So everyone looking at this, oh, cyber cab 2027. You know, AI cars with no steering wheels, and pedals are gonna be driving us around in 3 years. I would not hold my breath. I would honestly say it's either going to cost about 3 to 5 times that or and or it's not going to be coming this decade. So, yeah, I wouldn't exactly start saving your pennies, for a cyber for a $30,000 cyber cab by 2027, FYI. So a little bit more details about it.

Jordan Wilson [00:15:43]:
The details were impressive. If it ever if it ever comes to fruition, it's a completely other story. So Musk stated that Tesla aims to launch, which, I mean, we've been hearing about this one for years, its unsupervised full self driving FSD capabilities in Texas and California next year for existing model 3 and model y vehicles. Although the current FSD, that's full self driving system, still requires, yeah, human oversight. So kind of many different, announcements. So following the event like I talked about, Tesla's shares, marked a 12% decline year to date and 17% decline over the past year. So Tesla's not doing very good in terms of investor sentiment. Also, Elon Musk announced another new kind of AI driving, vehicle with no steering wheel, no pedals, called the Robovan.

Jordan Wilson [00:16:43]:
Yeah. That's Robovan. But the Robovan, I guess it's cooler if you pronounce it Robovan. So the Robovan is designed for transporting larger groups of goods or people, but that also seemingly failed to generate any real excitement at least when it comes to analysts. And Tesla has been getting crushed recently in the past couple of days since this announcement. Also, the event underscores the challenges Tesla is facing in bringing self driving cars to market, especially as competitors like Waymo already have this technology, and they've already launched successful robo taxi services in certain states. You know what? I was actually in Austin, Austin, Texas last week at a keynote speech I was doing, and I was super bummed that, I couldn't catch a a Waymo, kind of taxi from the airport, to my hotel. But, yeah, maybe maybe next time.

Jordan Wilson [00:17:40]:
Anyone out there, let me know. Have you guys, you you know, taken these Waymos or any, you know, self driving AI cars? I'm pretty interested in this, and let me know if you wanna hear more about this on everyday AI. We don't really go into self driving cars and, you know, AI powered cars too much because I think it kinda starts to blur the line. Right? And is that really, you know, impactful for your business and your career? Maybe not. Aside from the fact that, okay, if we do get fully automated, fully self driving cars, well, you can go do business in there, and you don't have to drive. So maybe you normally have a, you know, 45 minute commute, but if your car doesn't have a steering wheel or gas pedals and the law in your state allows it, okay. Yeah. That turns into 45 more minutes, you know, of working, I guess.

Jordan Wilson [00:18:30]:
But, you know, we don't really cover this too much, but, you know, I think when Tesla comes out with a pretty big announcement like this, it's worth paying attention to. So sabbatical here says, love to see the Waymo in Atlanta. Yeah. Yeah. And, Cecilia is saying there are already operational autonomous cabs in the Bay Area. Yeah. Absolutely. All of the innovation always strikes there first.

Jordan Wilson [00:18:54]:
Alright. Our next piece of AI news, y'all. Let's keep it going. Taking a shift here and going into health. So Microsoft has unveiled a new set of AI tools at reducing burnout in health care workers. I think this one's pretty important y'all. So Microsoft has announced a suite of new health care data and artificial intelligence tools designed to alleviate and administrate, sorry, alleviate the administrative burden on clinicians, a significant contributor to burnout in the industry. So right now, reports say that nurses spend up to 41% of their time on documentation.

Jordan Wilson [00:19:36]:
That's wild. So these new Microsoft AI tools could obviously be a game changer for the health care industry. Alright. Let's talk a little bit more about what these new AI tools are.

Jordan Wilson [00:20:32]:
So they include a collection of medical imaging models, a health care agent service, and an automated documentation solution specifically for nurses. So this comes in a set of multimodal AI models that can analyze various data types, such as medical images and clinical records, allowing health care organizations to develop more tailored applications for their needs. So right now, at least, Microsoft has partnered with Providence Health and Services to create a whole slide model for pathology, enhancing cancer mutation prediction and subtyping. So this model is expected, like I said, to be a game changer for health systems looking to improve diagnostic cape capabilities.

Jordan Wilson [00:21:39]:
Now let's talk the health care agent. So the health care agent service enables organizations to build AI agents that assist with tasks like answering clinical questions and identifying relevant clinical trials, potentially saving doctors a significant amount of time. Microsoft's automated documentation tool for nurses aims to streamline their workflows, which differs obviously from that of clinicians. So this tool right now is in development with input from organizations like Stanford Health Care and my hometown, place here, Northwestern Medicine in Chicago. So Microsoft's previous tool, DAX Copilot, has gained some traction among physicians, and the introduction of a similar tool for nurses highlights the company's commitment to supporting health care staff. So while many of these solutions are still in early development or previous stages, their potential to transform health care workflows is significant, particularly in light of the growing demand for efficient technology driven, solutions in the sector. Yeah. So, hey.

Jordan Wilson [00:22:45]:
Shout out to Jay. Had a great conversation with, Jay a couple of weeks ago about this very thing, about AI and health care and burnout. So he's saying since the advent of EHRs or electronic health record, time spent documenting has increased substantially causing burnout and for physicians and others a lot of time spent completing notes at home after hours. That is a huge a huge factor here, y'all. And I think the, I think if I'm being honest, the AI, kind of the intersection of AI and medical and health care is ripe for disruption. Right? Yeah. Especially that. Just note taking.

Jordan Wilson [00:23:23]:
Right? You know, so many times in the doctor, even over the last couple of years, I'm just sitting there waiting for the doctor to type notes or the nurse to type notes. Right? Luckily, in their professions, right, they're okay at typing, but it's like, why are we still living? Why are we still living in that, in that world? Also, even within the same health care system. Right? I remember my first appointment, like, 3 years ago where the doctor's like, oh, okay. Or, you know, maybe it was, like, two and a half years ago. The doctor said, hey. We'd like to record this, and our system will automatically transcribe. And I'm like, yeah. Like, why aren't we doing that at all times? Right? I'd rather much I'd much rather be talking to a physician face to face than talking to the physician's shoulder or the nurse's shoulder as they, you know, peck on the keyboard.

Jordan Wilson [00:24:11]:
That's not their fault. Obviously, the health care system, they are really overburdened. You know, they're overwhelmed. You you know, I think especially, since the pandemic, it's it's just resulted in this long standing burnout for our medical system here in the US. Our health care system's overwhelmed. So, yeah, I'm personally looking forward, to this Microsoft partnership, you you know, really rolling out to these other institutions, as well as the big medical players starting to integrate this technology into their existing systems. So, please, let's get this fixed. You know, the fact that let me repeat this.

Jordan Wilson [00:24:50]:
Nurses right now are spending 41% of their time on documentation, and it should be a fraction of that. That is what AI does best, right, to accurately transcribe words to be able to put that in a computer. So, hey. Health care system, medical innovators out there, let's get on board, whether it's with, you know, Microsoft. Google has, some great options for AI in the health care industry, but it's about time that we, unburden the health care industry by leveraging artificial intelligence. Yeah. We have to figure out privacy concerns. I yes.

Jordan Wilson [00:25:31]:
I get that. Private health information, PII, all that. But, yeah, let's tackle that because this is a no brainer to help, not just relieve the broken health care system for doctors, for nurses, all of these people who are just overworked, but also for us. Right? Us humans that wanna get into a doctor, and we have to wait 3 months, 6 months, longer, because the health care system is so drowned out. Alright. Let's keep this thing going, y'all. Couple more AI stories here. I'd say 2 2 big ones.

Jordan Wilson [00:26:04]:
We always save the we always save the best for last here on the AI news that matters y'all. Alright. So OpenAI has reported a rise in AI generated fake content targeting US elections. So OpenAI has revealed that its AI models have been increasingly exploited to create fake content fake content aimed at influencing elections, raising concerns about the misuse of technology in the political landscape. So in 2023, OpenAI announced that they neutralized over 20 sophisticated attempts where its AI tools were used to generate misleading articles and social media comments related to elections. Notably, a set of ChatGPT accounts was identified in August that produced content on US elections, highlighting the potential for AI to spread misinformation. The company also banned several accounts from Rwanda in July that were involved in generating election related comments for social media platforms. And y'all, this isn't like 1 or 2.

Jordan Wilson [00:27:11]:
This is in the 1,000. This is just bulk election misinformation. OpenAI found that many entities and state actors were using their tools for reasons that the company doesn't want to. So despite these attempts, OpenAI noted that none of the content generated achieved significant viral engagement or maintained a sustainable audience. The US Department of Homeland Security has expressed growing concerns about foreign interference in the upcoming November 5th elections with countries like Russia, Iran, and China potentially using AI to spread divisive information. Y'all, we are in that homestretch, right, for, election season. The election coming up here in the US in, about 4 weeks. And now is the time, especially if you are listening, to this podcast, if you're watching this livestream.

Jordan Wilson [00:28:10]:
Stay vigilant. Right? You you have to always double check anything that you read on social media. You have to double check what you read online because right now, and especially kind of in that last hour, right, this is when all of the kind of AI powered information is going to hit, its all time high. This you know, because here's why. Right? And there's you know, both sides are using this. Right? It's it's it's always kind of annoying to me when I, you know, read the AI news, and I'm like, oh, this political party, you know, used AI to do this. And everyone's like, oh, you're biased. And I'm like, no.

Jordan Wilson [00:28:50]:
This is just AI news. Right? There's certain candidates and certain politicians right now that are using AI to spread a little bit more misinformation, and disinformation. So you really need to stay vigilant about what you are reading, if is it true or not, what you are sharing. Right? All of these AI images now that look actually fairly real, and then these images are being used to create videos that are very real. So as the US election, comes up in a couple of weeks, a reminder to everyone, please always stay vigilant. Get your news from trusted sources. You know? Don't pay too much attention to what is spreading on Facebook and Twitter and YouTube because there is a good chance in the coming weeks when it comes to politics, there's a good chance a lot of it could be fake. Right? We've kind of known that this last month right before the election, it's going to be an onslaught because by the time that the general public is probably going to realize something is AI generated, it could be too late.

Jordan Wilson [00:29:53]:
So keep that in mind. Alright. Our last piece of AI news here, an interesting one. Apple researchers are challenging the logic of leading AI models, including OpenAI's latest o one. So a new study from Apple researchers raises the significant questions about the, logical capabilities of today's large language models, including OpenAI's latest reasoning model o one. So the research team at Apple developed a new evaluation tool called GSM symbolic, which builds on the GSM 8 k mathematical reasoning data set to test AI models more rigorously. So, yeah, if you're a dork like me and you follow all of the benchmarks for large language model, you'll probably recognize that GSM. Right? So the GSM 8 k is a fairly popular, kind of benchmark, to see how smart these models are.

Jordan Wilson [00:30:56]:
So that is essentially grade school math 8 k, with that 8 k standing for 8,000 math word problems. So, yeah, these are different, essentially tests for large language models, to see, okay. Are they actually smart? Can they actually reason? So this new one from Apple researchers is similar to the very popular gsm 8 k, but this one is called GSM Symbolic. So the study from Apple researchers tested a range of models, both open source and proprietary, including Google's, or sorry, Meta's Llama, Microsoft's PHY, Google's JEMMA, Mistral, and OpenAI's GPT 4 o and o one. So revealing that even top models struggle with real logic. So current accuracy scores on the GSM 8 k dataset are deemed unreliable with performance fluctuations observed across different models. For instance, the llama 8 b model scored between 70 80%, while PHY 3 varied from 75 to 90%. So, yeah, huge gaps.

Jordan Wilson [00:32:08]:
Right? 10 to 15% gaps in these, tests are kind of saying they're unreliable. So researchers noted that as task difficulty increases, the variance in model performance also grows suggest that handling this variation may require exponentially more data. And despite achieving high scores on benchmarks, OpenAI's o one model, so that is, you know, the one that we originally, you know, was codenamed QStar, and then it was called, strawberry, and now it's rolled out as the o one model. So, despite achieving high scores on those benchmarks, the o one model still exhibits performance fluctuations and makes basic errors, reinforcing the idea that they are more advanced pattern matchers than true reasoners. Alright. So pretty, this didn't grab a lot of headlines because it's a little dorky. Right? Essentially, Apple researchers went. They did a bunch of very advanced benchmarking across all these models and said, hey.

Jordan Wilson [00:33:16]:
These models can't actually reason. Right? And then they suggested a new benchmark going forward that could better test, this reasoning ability and in doing so, hopefully, make that wide range. Right? So, essentially, what they're saying is, number 1, these models can't actually reason. Right? They can't actually use logic. They are just NEXX token predictors. And one of the rationales behind that was, well, look at look at this wide variance in these scores. Right? Scoring anywhere from a 75 to a 90. Right? If you are able to reason and repeat that reason over and over, you wouldn't have a a 15 percentage, kind of variation in these scores.

Jordan Wilson [00:33:59]:
Alright. So this study highlights a growing debate between leading AI research institutions as OpenAI maintains that o one represents a significant step toward logical agents, while Apple researchers argue that the evidence points to limitations in reasoning. Let me just cut it to you real. I think there's a lot of different ways that this study can be interpreted. Right? Number 1, I think we all have to understand this kind of shift that's going on in the large language model industry. Even though this is kind of OpenAI's, language, but we are really taking a step. Right? Kind of OpenAI, you know, released slash leaked their 5 stages, toward AGI a couple of months ago. You know, stage 1 being chatbot, stage 2, being reasoners, stage 3 being, agents.

Jordan Wilson [00:34:58]:
So OpenAI CEO Sam Altman said, a couple of weeks ago that, hey, essentially, that OpenAI has achieved stage 2, which is reasoners, which are models that can reason, models that can, you know, in the o one model, essentially something that is more than a next token prediction. Right? Because right now, the best way to get the most out of large language models is a human has to continually guide it. A human has to right? I like having visuals. Right? And you say, oh, you know, you say something to ChatGPT, and then ChatGPT gives you a response. And then you go back and forth, back and forth. Right? And you you might be chatting, and you might have a, you know, 15, 20, 30 interactions with ChatGPT in the same kind of chat window, trying to guide it, to better outcomes. Right? And if you're doing this correctly, right, you might say, oh, that's, you know, some some good prompt engineering. That's, you know, using, you know, chain of thought prompting.

Jordan Wilson [00:35:58]:
You know, there's all these different prompting techniques. And, essentially, with these new models that are reasoning models, essentially, the big companies are saying, hey. Now the models themselves are going to be doing that work that the human would normally be doing. Right? So, normally, even though you start with giving ChatGPT or other large language models an end goal, it still takes a lot of finessing from humans to get them to that end goal. Some, you know, some subtle, guidance and some, and some correction. Right? So with these new reasoning models, the thought is, okay, these models are actually going to be kind of going, you know, doing the actions in all of those incremental steps, you know, kind of step by step, right, is a common, phrase when prompting a large language model, to go through this chain of thought reasoning or thinking more like a human. So Apple researchers here, you have to think. Right? Don't get me wrong.

Jordan Wilson [00:36:56]:
All these big companies have some of the most brilliant minds, but you have to look past the research. Right? Why was this study even greenlit? Well, you have to remember, reports last year was that Apple was spending 1,000,000, 1,000,000 with an s, 1,000,000 of dollars a day on developing its own internal large language model. Yet when they announced Apple Intelligence, which has not had a very successful rollout so far, right, Apple was not able to kind of release its own flagship model. So they have a small, Edge AI. Right? So they have a kind of a small large language model that runs on device or on computer, right, on on your computer, on your smartphone, and doesn't require the cloud. But for the most part, for most complex tasks, it's actually tapping in to OpenAI's GPT 4 o. So I always like to read through the, you you know, the marketing speak here. So you have to think if there's some other, implications here of Apple researchers.

Jordan Wilson [00:38:09]:
And, again, Apple spending 1,000,000 of dollars a day trying to develop its own large language model seemingly did not accomplish its goal in time for Apple Intelligence, therefore, having to rely on a third party. And there's been other reports that Apple is still looking at, working in the future with other large language model developers, whether that's Google, or Claude, you know, Claude from Anthropic. So you have to think if there's maybe some other, other things going on kind of behind the scenes or for the reason why this model or this study is coming out with Apple researchers saying, hey. These super smart models, yeah, they can't actually think. Don't worry about them. So pretty pretty interesting study. You know, if you're a dork like me, you might you might read, you you might enjoy, reading it. The other thing, and we're gonna close with this before I recap everything, there is a bigger argument here, right, for even how large language models work versus how the human brain works.

Jordan Wilson [00:39:14]:
Right? I think there's been so much of a focus on saying, hey. Yes. Large language models, even though as they get more and more powerful, you know, keep in mind, they're just next token predictors, which on the surface, yes, very true. But, also, you know, it's it's people are saying they they they make this argument as to why AI is not super innovative, as to why AI is not going to be impactful. So, the the deflectors always say, oh, it's just next token prediction. It's only as smart as its dataset. Right? Isn't that humans? Isn't that how our brains work? Similar to similarly to a neural network. Right? That's what we are using.

Jordan Wilson [00:39:58]:
We are using the dataset that we've been trained on. Right? Your own life experiences, what you learned in books, what you've read, what you've been taught. Right? Reinforcement learning. That's what humans go go, you know, go through. Right? When I'm little and I break a lamp and my mom says stop breaking the lamp, eventually, my actions that I make from then on are probably going to reduce the likelihood that I'm going to break that lamp again. Right? So, I mean, is there that much difference in the end with how a large language model processes data? It's trained by humans. Right? Reinforcement learning with with human feedback, And then the data is always updated. And, yes, it's a next token prediction, but, that's essentially what human brains do.

Jordan Wilson [00:40:43]:
Right? So, you're you're gonna continue to hear a lot of this y'all, through the rest of 2024 and into 2025, you know, as companies like Google, you know, reportedly, we talked about this last week on the AI news that matters. They're reportedly working on their reasoning model. Right. So we're gonna be going really shitting hard into talking about reasoning models, models that can actually, quote, unquote, think like humans, and agentic system. So that's really what we're gonna be focusing on a lot more here, in the last couple of months of this year and going into 2025. So I think it's important to read studies like this, but to also know why they're coming out at the time that they're coming out. Alright. Yeah.

Jordan Wilson [00:41:24]:
Someone said here, big bogey face app just Apple Intelligence face palm. Yeah. Not super impressed with Apple Intelligence so far. Alright. So that is it for the AI news that matters this week. Let's do a very quick recap here. So first, we talked about Apple is facing up to $44,000,000,000 in losses according to a report from the information before it may achieve profitability by 2029. The Nobel Prize, in physics was awarded to, groundbreaking research in artificial intelligence from John j Hotfield and Jeffrey Hinton.

Jordan Wilson [00:42:03]:
Next, Google is in hot water as the US Department of Justice has proposed some remedies for Google that could reshape the online giant's, kind of share in the search and AI landscape. Tesla undealed a lot of AI powered driving capabilities in its new cyber cab reportedly, to be released without steering wheels, without pedals by 2027 and for under $30,000. I don't think anyone actually believes that, and Tesla's stock went into the tank. Next, Microsoft has unveiled some new AI tools aimed at reduce reducing burnout in health care workers. Next, OpenAI reported a rise in AI generated fake content targeting US elections. And then last but not least, Apple researchers released a new paper challenging the logic of leading AI models and their inability to reason, including OpenAI's latest o one model. Alright, y'all. That was a lot.

Jordan Wilson [00:43:13]:
I hope this was helpful. If it was, please share this show with someone. Yeah. I know it might be your little secret at your company because now if you listen all the time, if you read our newsletter, you're probably one of the smartest people in AI at your company. So everyone's looking at you for the answers. Right? Everyone's looking at Michael for the answers and looking at Marie for the answers and doctor Scott. Right? This might be your secret, but, please, we'd appreciate it if you'd share this with others. If you're listening on the podcast, Spotify, Apple, thank you so much for tuning in.

Jordan Wilson [00:43:42]:
Please make sure to follow the show. Leave us a rating and a review as well. And please go to your everydayai.com. Sign up for the free daily newsletter, and join us back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI