Episode Categories:
Resources:
Join the discussion: Got something to say? Let us know here
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course!
Apple’s AI Research: A Critical Examination for Business Leaders
There’s a digital storm brewing around Apple's latest AI research paper, “The Illusion of Thinking,” which challenges the supposed capabilities of advanced AI reasoning models. This controversy raises crucial questions for business leaders about the reliability and interpretation of corporate research, especially when technological advancements are at stake.
Examining the Core Arguments
Apple posits that large AI reasoning models encounter significant limitations when faced with complex tasks, suggesting that these systems might not truly understand or think as humans do. However, an analysis of the research reveals concerns about its methodology. The study reportedly employs flawed logic, cherry-picked data, and an all-or-nothing grading system that might have discredited even Einstein.
Understanding the Flaws
Delving deeper into the details, the research casts doubt on AI models' reasoning abilities by confining them within artificially restrictive parameters. Crucially, Apple did not allow these models to utilize coding, a fundamental tool for solving complex problems. Additionally, tasks assigned to the models exceeded their token output limits, rendering them unsurmountable under the predefined conditions. Thus, rather than a failure of AI reasoning, it may be a failure of the study design.
Strategic Motivations Behind the Study
The timing and context of this paper suggest a calculated release. Arriving just days before Apple's Worldwide Developer Conference (WWDC), the research seems less about advancing scientific understanding and more about managing expectations in light of Apple's own struggles with AI. With AI advancements being a lucrative sector and competitors like Microsoft leading, the strategic dissemination of such findings could be interpreted as a move to distract from Apple's AI shortcomings.
Lessons for Business Leaders
This scenario underscores the importance of critically evaluating research, especially when it emanates from corporate entities with vested interests. Business leaders should:
- Scrutinize Methodologies: Ensure that research methodologies are sound and unbiased.
- Consider Motive and Timing: Evaluate the strategic intent behind the release of such studies.
- Seek Diverse Perspectives: Reference independent research and multiple viewpoints for a balanced understanding.
Navigating the Information Landscape
In a world increasingly driven by AI advancements, discerning business leaders must navigate these waters judiciously. Analyzing the intent and validity of research papers is as essential as leveraging the technology itself. As the conversation around AI continues, staying informed and critical will separate the leaders from the followers.
Topics Covered in This Episode:
- Apple's Viral Illusion of Thinking Paper
- Critique of Apple's AI Research Methodology
- Apple's AI Deception and Flawed Logic
- Strategic Corporate Propaganda in AI Research
- Apple's $2 Trillion AI Market Loss
- AI Reasoning Models Tool Use Restrictions
- Tower of Hanai and Token Limitations
- Apple Research's Industry Skepticism Strategy
Keywords:
Apple's illusion of thinking, Apple's AI research paper, AI reasoning models, large reasoning models, strategic deception, cherry picked science, weaponized research, flawed logic, cherry picked testing, all or nothing grading, Apple's marketing tactics, Apple vs. Microsoft, Apple's AI failures, WWDC conference, Apple's intelligence, Ajax model, Apple's AI spending, Apple's competition, generative AI, Microsoft, Google, OpenAI, code usage restriction, token output limits, reasoning collapse, AI's reasoning limitations, Tower Of Hanai, reasoning lab, Claude 3.7 SONNET, DeepSeek, thinking models, chain of thought processing, corporate propaganda, premeditated media strike, fear, uncertainty, doubt, strategic media strike, data contamination, Apple's research credibility, research methodology, scientific integrity.
Podcast Transcript
create. Apple's latest AI research paper has gone viral. So viral, actually, it showed up in my wife's nightly business newsletter. She reads that usually has absolutely nothing to do with AI. So Apple's the illusion of thinking paper shows evidence. Well, they say that large reasoning models slam into a wall. The moment that tasks get too demanding. So sounds pretty fatal for AI, right? Maybe, but if you dig deeper, you'll find flawed logic, cherry picked testing and in all or nothing, a grading rule that would flunk Einstein.
Jordan Wilson [00:01:30]:
In other words, if you take enough time to deconstruct this study, you'll find it's not much of a study at all. It's marketing from apple and you shouldn't fall for it. So stick with me for the next thirty minutes or so, and I'll expose this quote unquote research paper for what it is. It's strategic deception. It's cherry picked science and it's weaponized research at best. I hope you're excited for this one. I am, if you're new here, welcome to everyday AI. My name is Jordan Wilson.
Jordan Wilson [00:02:11]:
I'm the host and we do this every single day. This is your daily livestream podcast and free daily newsletter helping us all not just keep up with the world of AI, but how we can use all this information to get ahead to grow our companies and our careers. So sometimes the information like today's show can be a little confusing, and that's what we do. We break it down, whether it's myself or bringing on world class experts. We do it every single day, and then we break it down in our free daily newsletter. So if you haven't already, please go to youreverydayai.com and sign up for that free daily newsletter. We're gonna be recapping today's show and a whole lot more everything you need to stay in the loop and stay ahead and be the smartest person in your company. And if that's what you're trying to do, then you are definitely in the right place.
Jordan Wilson [00:02:53]:
So some days, we start off by going over the AI news. I don't wanna make this, accidental, like fifty minute podcast. I'm actually trying to keep these things under thirty minutes, but we'll see because, at least for today, it's hot take Tuesday, and I got takes. Last year, audience, it's good to see you. Thanks for tuning in. Let me know. Should I take a nice, should I ramp it up? I cut like, I'm kinda feeling spicy. I hope that's okay with you.
Jordan Wilson [00:03:24]:
But let's, let's just get into it. Alright. Let's deconstruct, this paper, the illusion of thinking. Alright. Like I said, it's been grabbing a lot of headlines recently. And let me also put my cards on the table. Right? Because you may be thinking, okay, this, this Jordan guy, you know, he's obviously very pro AI. Am I sure? Yeah, you could say that, if I'm being honest, I'm pro AI because I feel there's no real choice.
Jordan Wilson [00:04:02]:
Right. I believe in large language models, the power of large language models, just the the the way the entire world is investing in them. There's really no other solution. And and one other thing, I I I wanna talk a little bit very quickly about my background. So I mentioned it a couple times in our, you know, 05/1940 plus episodes together, but, I started my career as investigative reporter. I did okay. I was a Pulitzer fellow. I won ACP story of the year.
Jordan Wilson [00:04:32]:
So when I look at these things, I don't just read the study. Yes. I read the study manually twice. I fed it to3 separate large language models. I I combined a a bunch to have conversations with the paper. So I did an old school manual, you know, and then using AI as well. So, I want you to know, I'm I'm not just blindly ever following any study, what it says or what it doesn't say, whether I have a preconceived notion on if I agree with it or not. But essentially, what this, study said is, Hey, these large language models that reason or what they call large reasoning models, they don't really think, right? This whole thinking thing, it's an illusion.
Jordan Wilson [00:05:15]:
And I'm, I'm actually very excited to break this one down, but I'm hoping to do it in a very concise way. So, if I go on fewer tangents today, that's probably why. So let's just start at a glimpse. All right. So maybe you don't have, thirty minutes. Maybe you have five. Well, let me spend the next two minutes just giving this to you at a glimpse. Here's what's happened.
Jordan Wilson [00:05:37]:
And then for the rest of the episode, I'm gonna lay it all down, lay it all out for you. So, you know, I'm not gonna, keep you captive here just, for twenty minutes just to get to what's actually happened. Okay. So Apple Apple, just released, about four days ago. It's illusion of thinking paper, and this was three days before their big conference, the worldwide developer conference. And in this paper, they publicly claim that advanced AI reasoning is more or less fake. They're saying that these large reasoning models don't actually think, their entire experiment was fundamentally ripped. Alright.
Jordan Wilson [00:06:12]:
And I'm gonna show you why and show you how. But, essentially, they said that AI reasoning models couldn't use code. What? Okay. Which is the single most effective way to solve the problems that the researchers were giving them. So, you know, the researchers are like, hey. Here's all these problems. And normally, a large language model would be like, yeah. I'm gonna use code.
Jordan Wilson [00:06:33]:
Right? I'm gonna use the tools at my disposal. But the Apple researchers said, meh, nah. You can't. They also misrepresented, the AI's intelligent decision to give up, on impossible brute force tasks as a reasoning collapse, which I would say, is not the case. And they were treating that as a feature, versus a bug. Next, Apple failed to disclose that their hardest tests were technically physically impossible for the AI to pass due to its handcuffed token limits. Yeah. More on that in a bit.
Jordan Wilson [00:07:11]:
And y'all y'all know I bring receipts. Alright? Also, the paper's timing reveal its true purpose. A strategic media strike to distract from Apple's own AI weakness right before WWDC and their lack of AI because everyone knows it's no secret Apple has failed, and I think this will probably go down as the biggest failure in business history. Apple's absolute failure to put together anything resemblance of artificial intelligence. Maybe that's why they called it Apple intelligence because they couldn't actually figure out artificial intelligence. Right? And this also this wasn't a good faith scientific study. It really wasn't. It was a calculated act of corporate deception disguised as research.
Jordan Wilson [00:07:58]:
And I'm not blaming this on the researchers. Right? I'm sure that there were some higher ups that were pulling some strings or maybe that, you know, passed this down. Like, hey. We need to, you know, get some research that's very, hard on these large language models, these reasoning models. Alright. So that's what we're gonna be going over. But aside from what I just laid out, which we're gonna go into more depth, I wanna talk. There's at least 2,000,000,000,000 other reasons apple is putting out this quote unquote paper.
Jordan Wilson [00:08:30]:
And I have taken my time mainly because I do our hot takes on Tuesday and this paper came out, I believe it was, Friday or Saturday. But there's 2,000,000,000,000 reasons, and no one's talking about this. Why is Apple doing this? Why has Apple put out multiple papers, that literally go against the power and the capabilities of large language models and AI? Well, one, I kind of already answered. They can't figure it out, but here's 2,000,000,000,000 reasons why. Alright. So this is, for our podcast audience. I do have some, visual slides on today's episode. I'm gonna do my best as I always try to do to describe them to you, but you can always check out the show notes, or go to our website and watch the video version.
Jordan Wilson [00:09:17]:
So then you can see what I'm sharing on my screen, but, essentially pre generative AI. And this was, in 2021. Pre generative AI. Apple was crushing the world. This was like '92, dream team, kind of dominance for, Olympic basketball. It wasn't close. Apple had a $2,100,000,000,000 market cap, in 2021, and the next closest company, Microsoft, had only a 1,600,000,000.0. Now I'm, like, I'm not the best at math, but that's not close.
Jordan Wilson [00:09:55]:
Right? Having a, half a billion dollar or sorry, half trillion dollars. Sorry. That was 2,100,000,000,000.0, market cap, versus Microsoft's $1,600,000,000,000 market cap. So they had a half trillion dollar lead on the next biggest company in the world, which is not even close. They were blowing out the competition. Like I said, this is ninety two dream team. This is, you know, 97 bulls, the seventy two and ten right there. No one's close.
Jordan Wilson [00:10:24]:
It is a blow out. Apple is blowing out the rest of the world in terms of we are the biggest, we are the best company and it's not even close fast forward to today. Yeah. Apple is the biggest company in The U S by market cap. Right. And now they are a half trillion dollars behind Microsoft. Right. Which I would say, you know, depending on how you look at it, you could say it's Microsoft.
Jordan Wilson [00:10:50]:
You could say it's Google. I would probably say Google. I would probably say Microsoft is Apple's closest competitor. Right? Because everyone's kind of, changing in terms of, where their revenue is coming from, you you know, where they're trying to compete, etcetera. But you could say historically that Microsoft and Apple have been the two competing with each other. So let me say that again. In 2021, Apple was blowing Microsoft out, alright, to the tune of a half trillion dollars. Now Microsoft is blowing out Apple.
Jordan Wilson [00:11:23]:
And if you took the same growth rates, if you take the growth rates that Microsoft had from 2021 until today, going from about a $1,600,000,000,000 market cap to a $3,500,000,000,000 market cap, If Apple stayed on a similar or that same growth trajectory as Microsoft did, do the math y'all. That means Apple is at about a $5,000,000,000,000 valuation. Instead, they're staggering at 3,000,000,000,000. So, essentially, if they would have made the similar moves that Microsoft did, presumably, they would be a $5,000,000,000,000 market cap company. So they have left. You can make the argument that they've left $2,000,000,000,000 in market cap on the table by not figuring out AI. And it's not for not trying. Right? We've seen a lot of reporting going back multiple years, report from 2023, which I remember covering this report the day it came out on the Everyday AI Show.
Jordan Wilson [00:12:21]:
Yeah. We've been doing this thing for a while. And it said Apple is reportedly spending millions of dollars a day training its AI. And Apple internally at the time said that their internal model, which it was, codenamed Ajax, and it did come out under a similar name. They said it is the most advanced language model, and it is more powerful than ChatGPT. Alright. Imagine spending millions of dollars a day, just on training AI models if you're Apple. And when you finally, quote, unquote, released it, you didn't even say it by name in the main keynote.
Jordan Wilson [00:12:55]:
It was almost like Apple was embarrassed by the language model that they released at last year's w w w w d c. It was a small law a small language model that lived on device. It's Ajax model. They didn't even say it by name in the main keynote. Right? Because if they would've, it would've been embarrassing. Right? So it's it's almost like they didn't wanna claim it because they have reportedly spent many, many, many millions of dollars. And by many millions at that point, I mean, yo, millions of dollars a day back in 2023, do the math as potentially hundreds of millions or billions of dollars that they spent on AI that just didn't work. And like I said, apple is the only big tech company that has failed to produce the most basic, of AI offerings.
Jordan Wilson [00:13:45]:
Apple's produced nada, nada that works at least. Right? So some, some headlines here from some different publications, like payments, Axios, Bloomberg, PC Magazine, Computer World. Let's read some of these headlines, shall we? The Verge. This is a crisis. New Apple report claims will get no Siri, upgrades at WWDC due to AI turmoil. Apple's AI headaches could lead to lukewarm revenue growth. Drama at Apple as AI failures cause heads to roll. Apple sued for false advertising over Apple Intelligence.
Jordan Wilson [00:14:29]:
Why Apple still hasn't cracked AI? Two more class action lawsuits target misleading Apple Intelligence claims. Yeah. Apple's rollout of AI was absolutely so bad that they have are now facing multiple class action lawsuits because they couldn't deliver the simplest version of AI that they promoted. Right? And I'm technically one of those people. Right? I I I have to be honest. I'm recording this on a Apple Mac mini. The the the camera I'm using for the the the the live stream here, it's the new iPhone. And one of the reasons I bought this new iPhone is because they're like, we're gonna have all this new AI on the iPhone.
Jordan Wilson [00:15:06]:
And, here here it is almost a year later, there's not a single thing on this iPhone. That's quote unquote AI. There's not right. Like I'm looking for it. I'm like, Hey Siri, find me the AI and series, you know, ten minutes later, would you like me to use ChatGPT for this query? So yeah, apple has fumbled the bag harder than any company has ever fumbled the bag. I would say from a business perspective, because when you think of the numbers, I think of the numbers, I don't think that's an exaggeration because that's, even though it's a hypothetical scenario I laid out, it was probably a realistic scenario that apple should have traveled that path. They should have grown at the same rate that Microsoft grew over the last four years because of generative AI, but they didn't. They didn't, but they should have Multiple trillion dollar market cap mistake from Apple, which would very likely, I think, qualify that to be the biggest business blunder ever.
Jordan Wilson [00:16:05]:
And it's probably not even close. So all of Apple's competitors have been cashing in on app on AI, and Apple is still failing. So with trillions of dollars on the line, Apple needed a red herring, the claim that AI reasoning is an illusion. Right? Because all these other companies, even though Apple has their own Edge AI, these are small language models that live on device. They don't have a large reasoning model. So, essentially, what's happening here is all these other companies are running away, you you know, getting insane revenue, from their AI offerings. And Apple's like, what if we just throw some deception and doubt and confusion in the ring here right before our big event? Right? So then people will not be mad at us if we don't release anything AI at WWDC this year. So that was, yesterday.
Jordan Wilson [00:17:03]:
On Monday, Apple had their WWDC event where they essentially took a quote unquote gap year. It was reported they took a gap year on AI. They didn't really release anything new. Whereas last year, at their WWDC, they said AI every three seconds. They actually rebranded it because it's Apple. They're like, oh, it's not even artificial intelligence. It's Apple intelligence. Our AI is better than AI.
Jordan Wilson [00:17:26]:
Right? And here we are a year later, and they're like, whoops. We're getting sued. We couldn't deliver, so let's take a gap year. And instead, let's create some confusion. Let's get a huge viral study. Let's throw a, you you you know, a a big smoke, screen in front of everyone, cause some chaos, and then maybe people will, temporarily forget that we stink at and that we haven't been able to deliver on our promises. And maybe shareholders will look at this study and be like, oh, smart apple. Yeah.
Jordan Wilson [00:17:55]:
Look, this, this great research, shows that, these large reasoning models don't work. Good thing. Apple hasn't figured it out. Wrong. So let's actually deconstruct this thing. Let's take it down. All right. On my screen, I'm showing you the difference between Apple's quote unquote study, which is on the left and what I think a real study should look like on the right.
Jordan Wilson [00:18:25]:
All right. And I've, I've, I've talked about the one on the right. Apple's study, quote unquote, is just apple researchers. All right. Which is not abnormal. Okay. I'll say this, and I'm not saying this in, how do I say this? Like, a lot of people are throwing some shade, at some of the Apple researchers because they're technically interns. I'm not gonna do that because that's technically normal.
Jordan Wilson [00:18:55]:
Right? So when PhD candidates in computer science, right, are are looking, to, complete some meaningful research. You know, a lot of times, big tech companies will hire them on as interns or they were already interns there to begin with. So I'm not gonna go down that route because these people are very capable. But one thing I want you to look at, it's all Apple researchers, and that's it. And like I said, that's usually only normal when you are announcing a new model and you put out a paper around a new model, right? Otherwise, good research that changes the conversation, on artificial intelligence would usually look like the paper on the right. This paper is personhood credentials. Alright. This was a pretty, meaningful research paper that changed the narrative or at least tried to change the narrative on, you know, AIs that are trying to, you you know, imitate humans.
Jordan Wilson [00:19:52]:
This is what research looks like. Because on this piece of research, you have researchers from multiple big companies. You have them from OpenAI, Harvard, Microsoft, University of Oxford, you know, a lot of other ones. My my my screen is actually a little blurry here, but, it's, it's from dozens of companies and universities throughout the world. That's what a normal research paper looks like. Right? You would see multiple companies, multiple research institutions on the left. That's what marketing looks like. Only apple researchers, nothing else.
Jordan Wilson [00:20:37]:
All right. Real quick. Got to take a quick break for a word from our sponsors.
Google Gemini [00:20:45]:
This podcast is supported by Google. Hey, everyone. David here, one of the product leads for Google Gemini. Check out v o3, our state of the art AI video generation model in the Gemini app, which lets you create high quality eight second videos with native audio generation. Try it with the Google AI Pro plan or get the highest access with the Ultra plan. Sign up at gemini.google to get started and show us what you create.
Jordan Wilson [00:21:16]:
Alright. Let's get back into it, and let's break down this paper a little more. And I'll tell you this. The paper's out there. It doesn't take long to, to read. And I think enough people by now have already gone through the more technical side of this paper line by line. So I'm just gonna more focus on some big picture, ideologies and methodologies that were, completely elementary, and, just defied logic. Not in a good way.
Jordan Wilson [00:21:43]:
Right? So let's start with their Apple's premise and these kind of flawed benchmarks. Okay. So Apple claimed the need for this test. Right? Like, why would they even come out or or sorry. Why would they come out with this research? They argued that standard AI tests for math and coding are unreliable due to data contamination. Data contamination is like kind of saying like, hey, all these other, you know, studies that all these other researchers do from multiple companies, multiple universities, yeah, they got it wrong because, you you know, their data's bad. That's not good for a researcher. And that's why I also don't think that this is gonna turn out well for Apple because they essentially just kind of slapped a bunch of researchers silently in the face and said, yeah, your research is is is rubbish because you didn't even know that your data is contaminated.
Jordan Wilson [00:22:29]:
Not a good look. Alright. So they said that also all these other all like, every other single, you know, research paper, you know, it's just contaminated data. And and and these models are essentially just memorizing, and they're just, they're they're not even reasoning. They're just remembering. Right? So it's a valid concern here, so okay. We're still fine. And their proposed solution was to create a clean and controllable environment to test what they said was a true unvarnished reasoning, limits of modern AI.
Jordan Wilson [00:23:01]:
Alright. Sure. Let's see what she got, Apple. So they came up with their kind of reasoning lab. They built what they said was a sterile testing environment using four classic logic puzzles, framing them as pure tests of logic. And each puzzle was paired with a simulator, an automated referee that checked every single move the AI made and immediately flagged the illegal one, ending the test with a failure. So if a model got any of these four puzzles, a single move in any of these four puzzles wrong, test over failure. So, not good.
Jordan Wilson [00:23:41]:
That is hyper strict, unforgiving. That's not how large language models, especially reasoning models, would generally work, but, okay, sure, Apple, do your thing. Not making sense, but let's keep going. I do wanna talk specifically about one, of these, kind of logic games that they used, the Tower of Hanai. So this is a very classic game. And also, all of these games are classic, which were already disproving Apple's point that they were trying to prove because they said all these other, you you know, benchmarks out there were contaminated. So they thought, like, oh, well, we can use a game like Tower of Hanae that's, you you know, nondeterministic because it's a game. Wrong.
Jordan Wilson [00:24:27]:
All the solutions, the the the algorithm, everything about this tower of Hanae is on the Internet. It's in it's in the training data. So, like, their their their original even reasoning for creating their games to test these reasoning models was absolutely bonkers. Like, no, it's already wrong. You'd like you're already wrong and we haven't even started. All right. So this was their thought. So, the Tower Of Hanai is a classic computer science problem.
Jordan Wilson [00:25:00]:
You have to move discs between pegs, never placing a larger pay, disc on top of a smaller disc. And there's, kind of three towers. All the discs start on the left tower, and you have to ultimately move them all the way over to the right tower with the largest disc on the bottom and the smallest disc on the top. So, you know, if there's only three discs like this example I have on the screen, it's not terribly hard. Right? But as you add more discs, there's more complexity. So, you know, as an example, they gave games like this, but we're just there's three other ones. Let's just talk about the tower of Hanai. And then they gave a system prompt and then a prompt to different reasoning large language models, and then they had their them output their text, output their answer in text form.
Jordan Wilson [00:25:47]:
Right. And then they had a simulator essentially and double test, you you know, double checked all of the the the the, AI models results. Okay. Sure. They also did checker jumping, river crossing, and blocks world. So let's talk about the actual model. So they tested thinking, versions of deep seek r one and Claude three seven SONNET. They they did a lot of more technical testing.
Jordan Wilson [00:26:15]:
They technically tested some OpenAI models, but OpenAI doesn't show the complete chain of thought, whereas, DeepSeek r one and Claude three seven SONNET thinking two in the API. So, you know, they did the right thing there, right, by making those the baseline models, and they also tested them against the non thinking versions of themselves, which just adds some complexity. That's not even what we're doing here. But like I said, the scoring system is absolutely brutal because if you make one wrong mistake from getting it perfect, it's a zero. So there's obviously when you talk about, when you talk about these puzzles, they're extremely complex. Right? And there's many different ways that you can solve them, but also you have to think of the context window, of these models. Right? And and also the output limit for tokens, which we're gonna talk about here in a So this part is crucial. Alright? Because well, let's actually look at the results.
Jordan Wilson [00:27:19]:
So the results from what Apple reported, they said on easy puzzles, standard models did better or non thinking models. On medium puzzles, these reasoning or thinking models exceeded. And then on hard puzzles, they said all models completely failed. They didn't even try. Right? And and this is what they called the efforts collapse. So this is where you saw and as a former journalist, when I read this on Saturday, I'm like, oh, gosh. The media is gonna get because I saw it literally once it came out, right, because it was trending on Twitter right away because you saw these, you know, headlines like, oh, you you know, reasoning models collapsing. You you know, the AI wall.
Jordan Wilson [00:27:58]:
Right? Like, all these AI doomsday articles. And, like, I'm reading this, and I'm like, oh, gosh. Like, the media is gonna completely fall for this. Right. Being a former journalist, nothing against I like, I go to all these conferences. I meet brilliant tech journalists. And then there's some that, you know, are overwhelmed and you get all these press releases and you're like, okay. This is a salacious headline.
Jordan Wilson [00:28:19]:
Okay. Looks factual. It's a research paper. Sure. Let's go with it. Right? It's gonna click. Right? We're gonna get clicks. Look at these headlines we can put on this Right? So that's kinda what happened.
Jordan Wilson [00:28:29]:
And, you know, they talked about this efforts collapse, and that was their headline finding that on the hardest puzzles, the thinking models essentially would think less or even just give up generating fewer words before failing. So they just said, oh, reasoning models give up. And then also the algorithm failure, their supposed killer blow was in a separate test. They gave the models the step by step instructions or the algorithm, and it didn't help. And they all still failed at some point. So this is Apple's conclusion that, well, they just failed. Right? And and I have a graph here. I'm not gonna spend five minutes to explain it, but this just shows the complexity for the Tower Of Hanai example and the number of discs.
Jordan Wilson [00:29:16]:
So, you know, the more disc in that, example, the much harder it gets. Right? I can solve it with three. I could probably solve it with four, but I don't have time to waste. You know, to solve it with anything more than that, you gotta be either, like, have a computer science, math, like crazy logical brain, or you have to just study this game. Right? It's how some people can do the Rubik's cube, you know, while juggling in ten seconds while, you know, spitting fire, you know, or whatever these, you know, incredible acts of, you know, athleticism and, brainpower people do. But, you know, for the most part, the average human might be able to solve this at, you know, four discs, five discs. If you're a genius, maybe longer, but a human's not solving this at eight discs at nine, ten. Definitely not there.
Jordan Wilson [00:30:07]:
Right? So, essentially, it's not surprising necessarily that an AI couldn't. Right? Because if you get the smartest humans in the world, and give them a 15 disc, are they gonna be able to do it? I I I don't even know if it's possible. Right? Anyways, let's look a little bit here about what this actually means from an output token. That's important because the models they chose aside from the fact they didn't allow them to use code, which come on. They also surprisingly said that they only chose the, models with a 64 k token output limit. Alright. That's important to talk about because one of the requirements that the models had to do, in the output. So, Apple said that they weren't counting thinking tokens.
Jordan Wilson [00:31:00]:
That's not usually how it works. So that, you know, kind of chain of thought processing, which is a lot of what's happening under the hood. But they did require the model to spit out every single move. And to put out a a move, it's actually kind of complex. It's not like b one it's it's not like chess. You know? I don't know chess, but it's not like b two to d two. Right? One move can be very complex and can eat up a lot of tokens. So conservatively, right, I looked at the actual, example moves that they gave.
Jordan Wilson [00:31:30]:
They didn't obviously share their whole findings. It was very little, and saw that most, moves were, 10 to 12 tokens. So conservatively, a 13 disk. Right? A 13 disc, problem of this Tower Of Hanae would require 65,000 output tokens. I'm gonna repeat that. The study was not possible. Right? They did it all the way up to fifteen, twenty discs. Can't do it.
Jordan Wilson [00:32:08]:
If you require the model and Apple, hey, Apple researchers, next time, do what smart researchers do. Yeah. I'm getting I'm getting mad because I read a lot of research papers, and this one I knew was marketing, and that made me upset. Right? Not just right because I I do this every day, but because there's a scientific community that I think is is disgusted by this and rightfully so. This was a a haphazard, terrible study. Let me just say let me just say, like, how it actually is. This is terrible study. They didn't share any of their actual results.
Jordan Wilson [00:32:45]:
They said, here's the system prompt. Here's an example of a prompt, and here's our overall outputs. Right? You need to share. Share exactly. Here's what the chain of thought said. Here was the you know, on the hardest on an eight, on a 10 disc. Here's what the output was. But going by how large language models work and the requirements in their own paper, they would have to output every single move.
Jordan Wilson [00:33:10]:
And if and I'm being ultra conservative here to solve a 13 disc would take more than 8,000 moves. And to be able to spit those all out as required by the system prompt and the example in the system prompt, it's not possible. Sixty, sixty five thousand tokens. Okay. So Apple, you literally designed a test that you knew was going to fail at a certain point of complexity, at least according to, kind of the laws and the math set So their conclusion at least while reasoning models are an illusion. And the thinking that we see is a trick, Right. They're not actually thinking, you know, they're just, they're just doing next token prediction. You know, it's, it's stage one thinking, not stage two.
Jordan Wilson [00:34:10]:
Right. I'm going to have a whole episode on this at some other point. The concept of reasoning. Right. In in like stage, like stage one and stage two thinking. Right. So what is reasoning? Right? I'd like to say it's just connecting stage one thinking anyways. Right? Stage, stage one is quick, intuitive, automatic responses based on pattern recognition learned from data.
Jordan Wilson [00:34:46]:
That's stage one. And then stage two represents more deliberate analytical or conscious approach that involves reasoning and planning. So you could say the same thing about, you know, non reasoning models and reasoning models. Non reasoning models, right, these are the faster ones. These are, you know, pattern recognition. But all stage two reasoning thinking is, it's just stage one, but slower. Right? So I don't know. Like, even the concept of arguing against reasoning models seems a little bit illogical when it's just really made up of stage one thinking anyways.
Jordan Wilson [00:35:24]:
It's like, what is reasoning? It's not a different language. You're just taking more time doing stage one thinking, right? Pattern recognition. That's all reasoning is anyways in my head. Right. I don't touch the stove because it's hot. Right? But I've learned those different things that lead me to make that reason or to, you know, think or plan ahead in a certain way. Right? If I'm planning, for a big show like this, I spent many hours planning this show. I'm using stage one thinking.
Jordan Wilson [00:36:01]:
Right? That's literally what I'm doing. Pattern recognition. I've done this so many times. I recognize patterns. I put them together. Right? That's planning. It's a lot of it's thousands or millions of neurons following in our, firing off in our brain. That's just quick, intuitive, automatic responses based on data and pattern recognition.
Jordan Wilson [00:36:23]:
That's all it is. Anyways, I'll save that for another day, another show. Let's get back to this Apple study. Right? The other thing is apple. Not only did they cook the books beforehand. Sorry. You did. Unless you actually share the data and we can make an assumption otherwise, or we can make, a connection.
Jordan Wilson [00:36:45]:
Otherwise, if we go by the math, if we look at exactly what happened and the fact that they literally decided the way that we're going to measure reasoning. Well, they said the data is contaminated, so they had this brilliant idea. Let's use a non deterministic game. It's already all on the internet. Anyways, it's already in the training data. So you're already wrong to begin with. And you, you, you, you find you cherry pick, Right? This is almost like they got results, and it then seems like they just reverse engineered the entire study. Right? I'm not saying they did, but in theory, that could have happened.
Jordan Wilson [00:37:26]:
It it this study makes no sense. Go read it for yourself two or three times, and then go talk with a large language model. Don't lead a large language model. Just ask, does this make sense? Or ask your own self. Does this make sense? Anyways, I have an example study here, an iPhone study. All right. So the apple researchers, you know, they went through 25 rounds. Let's say I get 25 new iPhones and I turn off cellular data and I turn off wifi.
Jordan Wilson [00:37:54]:
I turn off Bluetooth. I turn off everything, but there's a new feature, on iPhones called SOS and it uses satellite. All right. And then I go on vacation, and I'm on satellite mode and I'm testing the phone I'm testing, but I'm only testing a couple of things, you know, just like Apple did. I'm just gonna test, you know, FaceTime and, phone calls and, getting on social media, and using, you know, ChatGPT and Google Gemini on my phone. That's what I'm going to test. Okay. And then, well, turns out doesn't really work very well.
Jordan Wilson [00:38:41]:
So now instead of coming up with a specific report that says, Hey, I'm reviewing this SOS satellite feature, which just sends messages to emergency response services. Instead, I'm gonna say I'm gonna put out, well, it's factual. Right? I can put out the facts. I can say, hey. Here's what I did. And then at the very end, I'm gonna say, hey. I restricted, you know, wifi and Bluetooth. Right? It's similarly the way that apple set up this study, they restricted tool use, which is the way that any reasoning model would solve this thing.
Jordan Wilson [00:39:22]:
And guess what? I'm going to solve it here. Live in like thirty seconds. I'm not going to solve it. A large language model is going to solve it. You're going to see when you give the model, the tools that it needs, it does the job. So I don't know why apple thought, oh, this is brilliant. We'll just restrict its core capabilities. We'll put it in this super refined box.
Jordan Wilson [00:39:41]:
We'll we'll sprinkle a bunch of, you know, big words on people. We'll send it out to all the journalists and they're going to cover it. Yay. No. Yeah. Just wait until my report, the illusion of iPhone connectivity drops. Right? FaceTime doesn't work. So why does the paper fail? Well, there's a lot of reasons.
Jordan Wilson [00:40:03]:
I'm going to go through this quickly. The data, it is precise, but the interpretation is a spectacular fail failure of logic. Alright. Let's look at my just and I could go on for hours. I'm gonna try to go fast now, but they're clean test. So let's go over. I I I have five critiques here. Alright.
Jordan Wilson [00:40:22]:
So the test is raped. They're clean test used different puzzle games. One was tower of Hanai, but the solutions are plastered all over the Internet anyways. So, the test punishes a creative AI for not being a perfect monotonous calculator. Alright. Critique two. It is designed to guarantee failure on the harder levels of this testing. By doing no tool use, Apple didn't get the models tool use, and they couldn't write code, which is the obvious and the only way that a reasoning model would actually solve the puzzle.
Jordan Wilson [00:40:58]:
Also, they've set these arbitrary limits. They capped, the AI, specifically quad 3.7 thinking, which is the best model that they used, in terms of thinking. It was, you know, that deep seek. They capped it at $64,000 output tokens when there is a 128 k model available. And also the absurd scoring, the one mistake in your out rule pretty much ensures failure. Alright. And, yeah, receipts. Alright.
Jordan Wilson [00:41:29]:
So in in in this, when I'm reading the paper, like, when I'm seeing things that are verifiably false right away, how how can you take the rest of the paper seriously? Right? So, you know, Apple said in their report section a two, we didn't have to go to the bottom for this one. They said for Claude 3.7 SONNET thinking and nonthinking models, we use maximum generation budget of 64,000 tokens accessed through the API interface. And I literally went through and I looked at, the day this was released. The day quad, 03/07 thinking on the API was released, I went to archive.org. I got a screenshot. And, yeah, obviously, there's the max token is a 128 k. If you can't get the basic things right, why would anyone trust your outputs, let alone a flawed methodology? Alright. Critique three, mistaking intelligence for a flaw.
Jordan Wilson [00:42:34]:
So they said, giving up is actually a smarter strategy. So in this case, the AI correctly identified an impossible brute force task and sought a shortcut. That's the reality. And and the algorithm failure is a red herring. It proves the AI is a complex mind, not a simple machine. Alright. Critique four, well, this wall that they're talking about, it's imaginary, because Apple's wall was just an artifact of their own restrictive rules by cutting down the token output and restricting tool use. So, yeah, let's look live.
Jordan Wilson [00:43:10]:
What could go wrong here, doing this live? Alright. So, I built a working tower of Hanai in Claude 3.7. So, I didn't use Claude 4. I use Claude 3.7 with thinking. Alright. So this is a working verifiable tower of Hanae that I just built. Okay. So, again, I'm not gonna take too long to go through this, but the object is you have to move these three, discs from Tower 1 on the left all the way to Tower 3 on the right.
Jordan Wilson [00:43:42]:
And you can never have a wider. So there's it's it's kinda like a pyramid for a podcast audience. So let's just say there's a skinny, a medium, and a thick, right all the way on the left. So you can move them one by one, and you can never set a wider one on top of a skinnier one. So I'll go ahead and well, maybe I'll solve this. I did it earlier, and I could solve it correctly. Right. So there's a certain, certain number of moves.
Jordan Wilson [00:44:07]:
Alright. Luckily here, I was able to solve it. Alright. So I solved it in, seven moves, and that is the optimal number. So according to the study, if you make a wrong move, it's gone. Alright. So now I can reset this, and I'm gonna go to 10 discs. Okay.
Jordan Wilson [00:44:25]:
And this is, well, actually, no, let me go to 13. No, I'll I'll do 10. Alright. And I'm gonna turn the solution speed on it very fast because, I built this thing to have a solve mode. Alright. So I can just click solve and we'll see. It might take a while. Alright.
Jordan Wilson [00:44:39]:
So we'll check back on it, but we're gonna see the number of moves that this does. Alright. So like I said, it might it might take a while, because the, minimal solution, if you get it perfect, is 1,023 moves. Alright. Well, actually, we can wait because it's already at about 400. So you'll see here for our podcast audience, this is literally going through this game Step by step, the one that apple researchers said is not possible. You'll see here. I'm not a computer scientist, right? I'm just a pretty smart person that knows the difference between marketing and research and this apple paper is marketing.
Jordan Wilson [00:45:27]:
It's not research because you'll see silly old me, random guy here. Right. I mean, I'm not a random guy, but you know, I just built something and allowed Claude 3.7 Sonic with thinking to pass this. Right. Just for fun, I'm gonna do this thirteen one. Yeah. Optimal moves 8,000. Alright.
Jordan Wilson [00:45:49]:
We're not gonna sit and watch this, but maybe we'll check-in on it at the end, and then we'll talk about the number of tokens. This is the example I gave. The number of tokens exceeds, 65, 64,000. So let's go. Alright. So we'll check on that maybe at the very end. So let's go through our critique, our the end here. And I think the real motive here I'm sorry to use this word, but it is what it is.
Jordan Wilson [00:46:18]:
This is corporate propaganda. This paper isn't about science. It's a brutal corporate strategy. This is a textbook case of weaponizing research to confuse the pub to confuse the public and to hopefully distract stock analysts, enough to where you don't lose, hundreds of billions of dollars in market cap because you're not pursuing AI at the rate at which you should. That's exactly what Apple's doing here. Because if Apple truly believed that they had sound research, which they don't in this case, this isn't sound research. They would have invited researchers from other companies, competitive companies. That's what people do, or, you know, partner companies from multiple outside research organizations.
Jordan Wilson [00:47:07]:
Apple didn't do this because this is marketing, right? And no researcher at a prestigious university would have ever put their name on this study. It's not sound. There are more holes in this thing than Swiss cheese. And, this is really just this fits Apple's pattern of cynical research, because this is not the time they've done it. They've done it multiple times. Right. They put out these papers, that are essentially downplaying AI's impact while they're still scrambling to figure AI out. And this provides cover for their own AI weakness and their long standing failure of Siri.
Jordan Wilson [00:47:49]:
I mean, let's talk about this. How much has Apple invested in Siri? Countless amounts yet OpenAI, Google, and, and other companies and even little startups have smart AI, assistants like Siri that run laps around Siri. This is just Apple's pattern of failure. And then we, we, we, we, we have to talk about this premeditated media strike, Right? Are you gonna tell me who's actually believing this, right? That this comes out hours before, Apple's big WWDC announcement where, oh, we're actually not really announcing anything revolutionary, when it comes to AI. Oh, well, hey. Did you see our research paper? This whole AI thing, we're not sure about it. Look at this Reese. Oh, the research paper is useless.
Jordan Wilson [00:48:44]:
Right? Go read my, my iPhone reports where I take my phone in a cave, you you know, in a dark room and I write how the camera doesn't work and I cover the flash. Right? No. Anyone with a brain and who knows the basics of AI and takes the time to analytically read this report knows that this thing, ultimately, this is damage control. This is PR. This is marketing. This is a pre bottle designed to discredit the entire field right before they were underwhelmed. The, the FUD, right? The FUD strategy. So the whole point was to dampen, you know, with the the fear, uncertainty, and doubt is just to dampen the competitor hype and to, make breakthroughs from Google and OpenAI seem like an illusion, and Claude, you know, anthropic and to lower expectation for themselves.
Jordan Wilson [00:49:42]:
And this is, to posture themselves as skeptics, not laggards and a desperate move, I think, from a company that is playing from very far behind. So final verdict as we wrap this up. Right? Darn. My tower of, Hanae 13 disc may not finish in time unless I really draw this out, which I'm not going to do. This isn't a research paper. It's not. This is cherry picked science at best that is meant to deceive the public and apple accomplished that. Right? I, there's gonna be clap back, right? Because the scientific AI research community, they're pissed.
Jordan Wilson [00:50:25]:
Go take a look on Twitter, go take a look on Reddit. Like researchers are not happy about this because essentially what apple did by, you know, kind of more or less saying that all previous research, the data's contaminated and then they do this, you know, little board game test where all the results are online anyways. Like researchers are not happy about this because this was a slap in the face to them, and, like, essentially, Apple was kind of invalidating some great research or trying to anyways by saying, oh, all prior research was invalid. And actually us at Apple, we're just gonna now, set the tone and say, hey, these large reasoning models that are actually revolutionary, that are out there, literally curing diseases, finding new drug discoveries. And they're not that good. They're actually bad. It's it's an illusion. Don't worry.
Jordan Wilson [00:51:26]:
Trust us. We're apple. Right? But it's scientifically illogical. Like it's, it's a bias test with a predetermined outcome. I could be wrong there, but I think there's a reason that Apple didn't show their work. Right? Because then people would have quickly picked this apart before their WWDC announcement. And I'm telling you, this is not over. Whether it's in three months or three years, this thing is going to unravel and there is going to be well rounded scientifically sound research that takes this quote unquote study, this piece of apple marketing and just puts it in a shredder.
Jordan Wilson [00:52:07]:
Right. And apple will take a huge black eye from it. And rightfully so this is strategically deceptive, a ruthless and cynical market play and the real illusion inside this illusion of thinking paper. The real illusion is the paper itself. Period. All right. I hope this is helpful. Y'all that's it.
Jordan Wilson [00:52:36]:
I lied again. I said, I said like thirty ish minutes. We went 50. I'm sorry. Y'all I really wanted to do as much as I could to provide you depth because unfortunately, what happens a lot of times in this space when a big company comes out with a paper, you know, right around, an important event, sometimes the media just blindly writes about it, and that shapes the public discourse. And not everyone has a super watchful eye and can break this down at a at a really granular level, with context and tell you what it actually means. So I hope this was helpful. Alright.
Jordan Wilson [00:53:15]:
Little hot take Tuesday, extra spice for you. If you haven't already, please go to youreverydayai.com. Sign up for the free daily newsletter. We're gonna be recapping the short version, of this show. So, maybe you weren't able to listen to it all. That's okay. It's gonna be in the newsletter. So make sure you go to your everydayai.com.
Jordan Wilson [00:53:34]:
Sign up for that. Thank you for tuning in. Hope to see you tomorrow and every day for more everyday AI. Thanks, y'all.
