Episode Categories:
Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
Anthropic Claude Fable 5 and Mythos 5: Strategic Implications and Risks for Business Leaders
Anthropic’s latest AI model releases, Claude Fable 5 and Mythos 5, have set a new bar for AI capabilities available to the public—at least for now. The models’ potential to rapidly transform operational workflows, software engineering, and complex knowledge work is evident in highly specific, consequential launches, such as Stripe’s reported migration of a 50 million line codebase in a single day—a task that would typically consume months with traditional processes. Yet, alongside these advances are material operational constraints, cost barriers, and compliance concerns that shift the paradigm away from democratized AI access.
Below is an in-depth analysis of the capabilities, business value, and four critical risk vectors that should command the attention of anyone responsible for technology adoption and risk management.
Claude Fable 5 Capabilities: AI Productivity Gains and Limitations
Anthropic’s Claude Fable 5 delivers a substantial leap in model capabilities and performance, representing what was described as a “genuine step change” rather than an incremental improvement. The model executes elongated and highly complex tasks—far beyond the conversational or chatbot context—enabling use cases such as massive codebase migrations, advanced analysis, and long workflow automations 00:35, 03:01.
Access is currently available through Anthropic’s public-facing suite (claude.ai, Claude Code, etc.) and via API, with early-use benchmarks suggesting superiority to previous iterations (Opus 4.8, Sonnet, Haiku) especially in coding, knowledge work, and longer-context tasks 03:15. However, practical business adoption is immediately impacted by model access restrictions, variable performance based on use case, and hidden operational hurdles unique to Fable 5.
AI Subscription Access: Gating, Pricing, and Strategic Spend
Unlike prior releases, Claude Fable 5’s subscription-based public access is temporary and set to expire after June 22. Post-cutoff, non-enterprise users will require significant API expenditure to retain access 01:02, 10:52. The API pricing is set at $10 per million input tokens and $50 per million output tokens, making it twice as expensive as previous models and decisively moving high-power access into the hands of large enterprises with the budget for scalable API costs 10:28, 36:41.
Cost-aware business leaders are advised to immediately benchmark Claude Fable 5 against GPT-5.5 and alternative models during the remaining window of subscription access, as API spending for intensive workflows may increase by 5x, 10x, or even up to 500x compared to subscription models 12:22, 12:30.
Benchmarking AI Models: Capability Gaps and Business Use Cases
Benchmark data released by Anthropic—and third-party analysts—show Fable 5 and Mythos 5 leading in a majority of formal metrics, such as software engineering, vision, scientific research, and general knowledge tasks 13:19. Fable 5 notably enables non-engineers to build interactive 3D environments and prototype complex applications previously requiring entire development teams 22:25, 34:59.
Yet, careful analysis of independent benchmark suites revealed that performance leadership is not universal. In high-value business tasks such as long-context reasoning, multi-document research, and precise instruction following, GPT-5.5 Pro and select models from OpenAI and Google outperformed Fable 5 24:42, 26:15, 27:13. Anthropic’s own admission that certain safeguards and downgrade mechanisms quietly shift user sessions from Fable 5 back to older (less powerful) models further complicates reliability for business-critical tasks 09:50, 10:13.
Anthropic Model Safeguards: Use Restrictions and Compliance Impact
A defining element of Fable 5 is its strict use restriction and safeguard architecture. Anthropic has implemented conservative safeguards that:
Route particular queries—such as those involving cybersecurity, biology, or AI model replication—to an older model (Opus 4.8) without user notification 10:04, 13:45.
Limit or refuse service on “risky” subjects, with up to 5% of user sessions affected by these fallbacks 14:13, 37:54.
Quietly downgrade model capability or provide outright refusals, leaving business users without clarity regarding which model handled their data and potentially affecting workflow reliability 39:18.
Enforcement extends to chat history and organizational memory: activity in one session (even if compliant under previous policy) can trigger blanket restrictions or denied access to advanced features 18:01, 38:43. These guardrails, while designed for safety, may create unpredictable compliance liabilities and workflow disruption, especially in regulated sectors.
Data Retention Policy: Privacy, Confidentiality, and Legal Consequences
Anthropic now requires a mandatory 30-day retention for all input and output traffic on Mythos-class models, including Fable 5—regardless of enterprise custom agreements or prior opt-outs 15:01. There is no provision for zero-retention or HIPAA-compliant operation within these models. All business and personal data entered into the system is accessible for audit and review for at least 30 days 16:01.
This policy shift breaks sharply with industry norms at the frontier of AI development and may create significant legal exposure for any organization required to maintain strict data privacy or compliance (e.g., healthcare or legal sectors) 16:47. Flagged conversations can be reviewed by Anthropic staff, increasing corporate exposure to inadvertent leaks—compounded by previously documented security lapses such as source code and model leaks 18:30, 19:14, 20:07.
Cost-Benefit Analysis: Intelligence, Price, and the End of Subsidized AI
Comparative analysis of intelligence per dollar unit shows that while Fable 5 often demonstrates an approximate 6% intelligence lead over previous models, this comes at more than twice the cost 32:14. For organizations with annual AI spends in the high six-figures or more, this incremental uplift may not justify the exponential increase in cost—a critical consideration now that the era of heavily subsidized AI access appears to be ending across Anthropic, Google, and Microsoft 31:15.
The intelligence-to-cost ratio, especially when benchmarked on real-world business tasks (multi-document research, decision support, or instruction-heavy workflows), often favors sticking with proven alternatives or focusing spend only on use cases where Fable 5’s lead is both pronounced and commercially justified 29:41, 34:59, 36:41.
IPO Context and Business Roadmap: The Strategic Narrative
The timing and design of the Claude Fable 5 and Mythos 5 launch coincide with Anthropic’s pre-IPO revenue ramp, which is openly referenced as part of the corporate strategy 44:17, 44:49. By positioning the most advanced model behind API gating, subscription conversion, and enterprise agreements, Anthropic is making a direct play for large-scale, budget-ready business clients. This puts pressure on procurement and technology leaders to scrutinize both the cost structures and the narrative framing of “dangerous” AI power that only the platform’s gatekeepers can manage 44:05, 44:22.
Summary: Enterprise Adoption, Internal Benchmarking, and Next Steps
The outsize business value of Claude Fable 5 and Mythos 5 lies in specific high-performance use cases—massive code automation, artifact generation, and complex research tasks—if accompanied by enterprises’ willingness to absorb higher costs and accept deeper compliance complexity.
Immediate recommendations for organizations include:
Prioritize rapid, real-world benchmarking of Fable 5 during the subscription access window, focusing on actual workflow and team requirements.
Conduct rigorous legal and compliance due diligence regarding the new mandatory data retention policy, especially for regulated verticals.
Reexamine model cost-effectiveness in light of shifting subscription-to-API economics and prepare for increased spend if scale remains a priority.
Monitor model safeguards’ impact on mission-critical workflows, particularly those involving sensitive or complex subject matter.
The window to make informed, data-driven decisions about Claude Fable 5’s place in the business AI stack is both brief and crucial. The models’ advanced capabilities are real—but so are their new risks, costs, and operational boundaries.
Topics Covered in This Episode:
- Claude Fable 5 and Mythos 5 Overview
- Stripe's Code Migration with Claude Fable 5
- Fable 5 Capabilities vs. Opus 4.8
- Claude Fable 5 Access Restriction Timeline
- Four Major Drawbacks of Fable 5
- Benchmark Comparisons: Fable 5 vs. GPT-5.5
- Usage Limits and Downgrades in Fable 5
- Anthropic's New Data Retention Policy
- Instruction Following and Long Context Performance
- Fable 5's Coding and 3D Simulation Abilities
- Subscription vs. API Pricing Models
- Anthropic's Enterprise Targeting and IPO Strategy
Episode Transcript
Jordan Wilson [00:00:16]:
Anthropic just released the most powerful AI model the public has ever gotten its hands on. It's called Claude Fable five from the Mythos family, and it's a genuine step change in AI capabilities, not just another incremental update for most casual users. In use cases shared at launch, Stripe actually revealed that it used the family of models to migrate a 50,000,000 line code base in a single day. Work that Anthropic says would normally take a team like two months. So that's the good news. But here's the part that no one's saying out loud. The fable and mythos rollout might be the first real step toward AI that's not democratized. Because after June 22, your average Claude users will likely lose access entirely, leaving the true frontier power to big enterprises that have big API budgets.
Jordan Wilson [00:01:15]:
So intelligence for the, but not for we. But that's not the only con to the massive pros that this model brings. What else is there to be worry about? Well, Anthropic can decide where this model is allowed to work. It quietly downgrades you without telling you, and it actually keeps every prompt input and output that you send it for thirty days. So, yeah, we're breaking down today, not just the hype and the capabilities and the biggest upside, but also four big catches that should make every business leader nervous. And we're going to cover everything else in between as well. All right, let's get into it. So here's the big picture.
Jordan Wilson [00:02:01]:
Fable five is yes, definitely a capability shock. So anthropic shipped the most powerful public Claude model yet and the most capable the world has seen. So this is fable five from the mythos family. So if you're like, wait, Fable Mythos five one. Isn't this the first Fable? Yes. Alright. So essentially, Fable five is the made for the public version of the also new Mythos five, but with guardrails. So Mythos five is now available to the project glass wing, you know, elite there.
Jordan Wilson [00:02:38]:
But you have to be essentially invited to use the full Mythos five model, which is also released. But for everyone else, you have the Fable five, which is the made for the public version with more guardrails. But even so, Fable five is a genuine step change. It's the biggest capability leap that Anthropic has shipped. So the upside is clear and it's enormous. It executes real long complex work, not just better chat bots. Right? And you can not only use this in, you know, claud.ai and claudcowork and claud code and all of those things, but also on the API side. So the flexibility and the capabilities are a big step forward.
Jordan Wilson [00:03:21]:
It's enormous, but it arrives with four big catches. Gated access restrictions, a lack of transparency, and required data retention. For the most part, these are four new bottlenecks, roadblocks, or obstacles that we really haven't seen from a frontier model before, which is why today's show is gonna be less about showing you hands on some of the capabilities and more about actually discussing the pros and the cons and the facts behind this new model. So on today's show, stick with me for the next twenty five ish minutes, and you're gonna learn why Fable is the most capable public clog model ever released. You're gonna know what the step change actually does across coding, science, and workflows. You're gonna know those four catches, the shrinking access, usage restrictions, hidden downgrades, and data retention. And I'm gonna let you know how business leaders should adopt this frontier model without getting burned come, well, two and a half weeks from now. Alright.
Jordan Wilson [00:04:17]:
Let's get into it. Welcome to Everyday AI. My name is Jordan Wilson. We do this thing every day, and it's for you. This is your daily livestream podcast and free daily newsletter helping business leaders keep up with the nonstop feature upgrades, model drops, and everything else. I tell you what's important, what's not, and help you grow your company and career. So if that's what you're trying to do, it starts here. But make sure you go to your everydayai.com.
Jordan Wilson [00:04:41]:
That is your free generative AI university. Now almost 800 episodes, you can go listen, watch everything else, but make sure you go sign up for the free daily newsletter. We're gonna be recapping the highlights from today's show, as well as all of the other breaking news that you need to know that's gonna impact your career. Alright. So let's talk Claude at fable five in mythos five, which I think are anthropics, boldest and riskiest launches yet. So here's some of my thoughts. Alright. And notice this model came out Tuesday.
Jordan Wilson [00:05:11]:
I decided not to do the show Wednesday because I wanted to actually spend some time on this. Right? You know, maybe you saw on YouTube people within thirty minutes of it being released that didn't have early access or, like, here's a full review. That's not me. I think especially with the model of these type of capabilities, that's probably irresponsible to put out something within an hour or two. Right? So any hot takes that you saw, unless it was from people who had early access, which a lot of people did, I think it was probably premature. So I'll quickly even walk you through some of my first impressions before we break everything down in the show. So, is this Fable five the best model I've ever used? In some use cases, absolutely. But for me, I'm not personally blown away by Fable.
Jordan Wilson [00:05:58]:
And here's the reason why. And I think this actually an important distinction because we're gonna be looking at benchmarks, both that anthropic released and from artificial analysis. And for the most part, anthropic does obviously very, very well with the benchmarks. But one thing that literally no one talks about is almost every single benchmark out there is not against OpenAI's most powerful model. So what do I mean by that? So if you see something called like x high, that is the thinking version of OpenAI's g p d 5.5 with extra high reasoning, and that is what almost every single benchmark uses. Sometimes, Anthropics and competitors will even just take the default five five model, which isn't very good. Alright. So that's important to keep in mind.
Jordan Wilson [00:06:41]:
For me, I use five five pro. That thing is not benchmark. It is expensive. Right? So if you're on the $200 a month plan and you're someone like me, I use Five five Pro at least 20 times a day, usually usually more. Right? That's why I pay a lot of money, for some of these subscriptions. So when I'm doing some of my, kind of anecdotal testing, between Five five Pro and Fable, I mean, it's hit or miss. Sometimes Fable is better, but in many instances, Five five Pro is better. So if you're just looking at this from a pure intelligence standpoint, at least if you are someone that uses something like Five five Pro all the time, I don't think you're gonna be blown away.
Jordan Wilson [00:07:22]:
But I do think that is the mid that is the very minority of our audience. I would say the majority of our audience is not using five five pro for every single thing. But that's what I'm saying. Is this a huge capabilities gaps in terms of what was available? Not necessarily. Is it a big capability jump in terms of what people, the majority of users use on a day to day basis? Absolutely. So I'm I'm not trying to take away. Right? But, I've had a lot of people asking me, what do you think of Fable? And that's my honest answer. Right? It's it's better, you know, I I'm getting better results right now in Claude Code, with Fable than I am in codex with five five x high.
Jordan Wilson [00:08:00]:
But that's because in codex, you can't use five five pro, which is a much more capable model. You know, there's also some anecdotal things like I just saw, Dylan Patel of semi, semi analysis. Right? One of the most, kind of trusted, sources in the whole compute and, you know, technical side of AI. So he actually just shared a tweet. I'll read this. He said, you should usage share of OpenAI grew versus Anthropic yesterday despite the Mythos Fable five launch. Multiple power users at SemiAnalysis tried Mythos or Fable. They got refusals for nonsensical reasons, got pissed off at Anthropic, gave codex a legitimate try.
Jordan Wilson [00:08:40]:
Now they actually prefer codex to four eight opus. And I do think that's actually gonna be, an ongoing issue that anthropic is gonna have to, address. Anyways, let's get into it. Let's go over the Fable essentials. So, when it comes to any benchmark, Fable five is a big step up from where we were with Opus 4.8. So essentially, even let me break this down for our nontechnical audience. Now you essentially have four ish tiers of clawed models. So you have your base.
Jordan Wilson [00:09:13]:
So we're going from least powerful and fastest to most powerful and slowest. So you start with your haiku. That's your fast, not very smart. Your sonnet, that's kind of your, little slower but smarter. Then you have your Opus, which previously was the most powerful yet slowest model. And now you have your Mythos family, which includes Fable. So, you know, might use kind of Fable and Mythos interchangeably in certain respects, but that's just to let you know that Fable five at least is a big step forward, specifically from Opus 4.8. But you might be using Opus 4.8 even if you think you're using Fable.
Jordan Wilson [00:09:50]:
That's because some of the new strict safeguards block risky topics, and you might not even know. So if you're doing things like cyber in biology, you know, those queries might actually fall back to Opus 4.8. So Mythos five though is with fewer restrictions, but right now that's only available to selected partners through the project Glasswing or the trusted access. So, you know, they are kind of the same in terms of Mythos five and Fable five, but Fable five is kind of what's available for anyone in the public. And then, you know, your, Mythos five is essentially your trusted partners, version. Both models though are priced at $10 per million input tokens and $50 per million output tokens, which makes it twice as expensive, as Opus 4.8. And that does bring it much more in line with the five five pro model. Alright.
Jordan Wilson [00:10:44]:
So Fable right now is included on all your paid plans, but here's the caveat, only until June 22. And then essentially, Anthropic said, yeah, we're gonna pull it. Maybe we'll bring it back in the future. Maybe we won't. But essentially, if you want to keep using Fable five post June 22, yeah, you're gonna have to fork over the API cost, which are gonna be astronomical. Alright. So my recommendation is start using Fable five as much as you can right now while it's included. Right? Reality, you might if you try a very complex prompt even on a $200 a month Pro Max plan, you might not even be able to finish it.
Jordan Wilson [00:11:25]:
You know, I get tired though FYI of people saying, oh, you're anti anthropic. Right? You say all these things. No. I'm not. There's plenty of other people, and I have a list of people that did one. They tried one complex prompt with Fable five because they really wanted to push it on its capabilities on max plans, paying $102,100 dollars a month and said, I hit my rate limit. Right? The first prompt I tried, obviously, on a $200 a month plan, hit the rate limit. It couldn't do it.
Jordan Wilson [00:11:51]:
So I would recommend you take some of your most complicated problems, run them now through Fable five when you can still have subscription access. Right? Because it might be taken away after that. And then compare it to GPT 5.5. Right? Which is for the foreseeable future going to be still open for all. Because I think as a business leader, many people are gonna have to start making that decision. They're gonna have to start saying, are we gonna start, you know, maybe five x ing or 10 x ing our AI spend? Right? Because that's probably what it would cost. When you talk about the the difference between what is included in a subscription, it's subsidized, right, versus what you're paying via the API. I mean, many people have said it's 100, 200, 500 x more expensive.
Jordan Wilson [00:12:35]:
So that's why I would start using Fable five now while it's still included in the subscription. I would make sure that you have, good access to g p d 5.5 that might, you know, force you to pay the $200 a month plan. And then I would throw your most complex problems at it and then work with your team to see which ones are the best. Because after June 22, yeah, Fable five is leaving. And I do say, maybe it's gonna be Fable five is gonna be be the best model for you. And maybe you can justify the cost, but I think you don't have too long to, you know, go head to head. So reading from some of, anthropics released here, a couple couple things I think that are worth pointing out. So they say today, we're launching Claude Fable five, a mythos class model that we've made safe for general use.
Jordan Wilson [00:13:19]:
Fable five's capabilities exceed those of any model we've ever made generally available. It is state of the art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas. The longer and more complex the task, the larger Fables five lead over our other models. I would have loved to put that to the test, but hit my hit my, my my usage limit. Alright. Next, here's what they said, about the fallback on Opus four eight. Releasing a model, this capable comes with risks. Without safeguards, Fable fives capabilities in areas like cybersecurity could be misused to cause serious damage.
Jordan Wilson [00:13:58]:
We've therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next most capable model, Claude Obis 4.8. To release the model both safely and quickly, we've turned these safeguards, conservatively. They'll sometimes catch harmless requests, though they trigger on average in less than 5% of sessions. With more capable models arriving in the coming months, we're working to improve our safeguards and reduce false positives as quickly as we can. But here is the thing that I think all business leaders need to pay close attention to. This is the new data retention policy, and this is extremely troublesome. And I hope this is something anthropic reconsiders. I do assume they'll start to get a lot of flack as the mainstream media reads the fine prints or starts paying attention.
Jordan Wilson [00:14:51]:
So they said, finally, we're making a change to the way we handle business customer data for Fable five, Mythos five, and future models with similar or higher capability levels. We will require thirty day retention for all traffic on Mythos class, models on both first and third party surfaces. We don't use this data to train new clogged models or for any non safety related purposes. And we've instituted new privacy protections, including logging all human access to the data and ensuring its deletion after thirty days in almost all cases. The data will help us defend against complex and novel attacks, including new jail breaks and attacks that operate across many requests, as well as help us identify and reduce false positives. Alright. This is the first ever FYI. This is the first time a Frontier AI company has decided to not make data retention optional.
Jordan Wilson [00:15:49]:
So if you want to use this, you have no choice. Even if you there is no turning off, you you know, the concept of turning off model training means, you know, big company can't access my data. Nope. Not anymore. If you're using Fable five, Mythos five, or any future capable models, Anthropic is keeping your inputs and the outputs for thirty days. Here's what that means. So Fable five forces the thirty day retention on everyone. So even if you previously had a no retention deal that you sign, which is most enterprises.
Jordan Wilson [00:16:21]:
Nope. Nope. You are by using this, you are giving anthropic permission to access everything. And it's the first time a lab has chosen this. And I saw some people not really, paying attention to what this this means, and they're like, oh, no. This is part of the new, the president's new executive order on AI. No. It's not.
Jordan Wilson [00:16:40]:
This is something that Intropic chose to do. And there's huge ramifications here. I mean, talk about health care. Right? So that lets, Anthropic legally handle patient data that requires no retention. So the the the whole, you know, zero day retention, that's out the window. No such thing anymore. HIPAA, zero day retention, no longer thing, at least according to, what anthropic released. So what does this actually mean? Humans can review flagged chats.
Jordan Wilson [00:17:11]:
Yeah. Alright. So this is just risky. It's not just risky for legal, confidential, privileged client material. It's risky for just about anything. Giving a human the ability to review your enterprise chats, I think sets an extremely dangerous precedent. And I think and hope it's something that gets called out, among other, leaders in the AI space. And I hope and assume that OpenAI, Google, and Microsoft will not follow suit, because I do think that this is an extremely dangerous precedent to set.
Jordan Wilson [00:17:50]:
And I'm gonna tell you why. Well, another thing that's worrisome is memory actually pulled old chat into fable sessions. So this has been confirmed anecdotally. I confirm this myself. I've seen other people that have confirmed this that they literally can't even use the model, because if they've been working on things that were, considered, you know, dangerous or risky, by today's precedent, but not by yesterday's precedent, they literally can't use the model. And I do have some examples of that. But here's why this is worrisome. Anthropic y'all and please don't don't be like, oh my gosh, Jordan.
Jordan Wilson [00:18:26]:
You're so anti anthropic. No. I'm not. I am pro facts, pro truth. Right? And the fact and the truth is anthropic more than any other lab has had this little problem with leaking. Okay. So, you know, that Claude code thing, one of their most successful products ever. Yeah.
Jordan Wilson [00:18:43]:
The source code. Whoops. Got leaked. K. You know what else? This mythos model, this, the, the, the most powerful model that Anthropic has spent months hyping and, you know, putting all this marketing around as they prepare for their, IPO to go public. Right? This Mythos thing is so good, so powerful. Guess what happened? Whoops. It was leaked.
Jordan Wilson [00:19:08]:
Right? Literally, Discord users guessed and got access to it when they weren't supposed to have access to it. So I'm not saying that your chats, your company's personal data is at risk of leaking. I'm not saying that. But if you look at this new rule that you can't opt out of, and if you look at anthropics recent history, these things are in the last couple of months, right? This isn't when they were a young startup four years ago and didn't have everything in place. This is happening like now. Again, I'm not showing bias. I am trying to inform you of making the right decision for your company. That's all I'm saying.
Jordan Wilson [00:19:57]:
Because Anthropic does not have a stellar record, especially recently of keeping private matters private. All right, back to the fable show. So anthropic is kind of framing, both of these as mythos class model. So fable is not necessarily a step down from mythos, but it is in certain ways. Right? I I did post about this online that I didn't necessarily like how Infropic put out their benchmarks, which I'm gonna show here in a second. Because what they did is they labeled Fable five and Mythos five in the same column. And presumably, you know, when they put them in the same column, they took whichever benchmark was better and applied it unilaterally to both models. It's not the way this works.
Jordan Wilson [00:20:43]:
I would like a little bit more transparency, from Anthropic when it came to posting on their, release site on the benchmarks. Because a model with more safeguards is obviously going to have lower benchmarks. Because in some instances, it is going to default or fall back to Opus 4.8, which is now, you know, a good generation behind, right, because of mythos and fable. So I think that's an important thing to keep your, your eyes on that. Even though fable is part of the mythos family, it is in theory, less powerful, but overall, the capabilities are undeniable of this new model. Right. In in some of my initial testing, right. I was able to, kind of breakthrough on some, individual use cases where five five x high and Opus, you know, four eight max reasoning couldn't get through, especially when it came to those longer running tasks.
Jordan Wilson [00:21:52]:
So, on agentic capabilities, right, obviously, big capability leaps. You know, and traffic says fables lead really grows on the longer, more complex tasks. But again, it's tough. Right? Yeah. If you're on a $20 a month plan, good luck, you know, really realizing the utility here, which I think is unfortunate. You do have to unfortunately beyond that $100, $200 plan, to really, I think, realize the true benefits of this class of models. But the shift shows up first in anything coding, apps, you know, one thing that I think has been going very viral, in the first, you know, forty eight hours of release, its ability to create, interactive and immersive three d worlds via code, stitching together, you you know, hard and long workflows, things like that. So from a benchmarks standpoint, and here's what I was talking about, previously, you know, kind of, sticking mythos five and fable five in the same column even though they would if if you went through each of them individually.
Jordan Wilson [00:22:58]:
And you did see this a little bit more in their model card. So I know they did test them individually, so I don't like that they put them on the same pain. Anyways, you know, when it came to a lot of the big benchmarks, you know, big big steps big steps, not just, you know, from their previous model, Opus 4.8, but also g v d five five, Gemini three one pro, Gemini three five Flash. Right? Essentially, in almost not almost every, but in a majority of meaningful benchmarks, Mythos five and Fable five shot ahead. But not every single one. Right? There's a handful of pretty important benchmarks where you would have thought previously. Right? I think a lot of people, especially the way that Anthropic was hyping this, and they've spent months hyping it. You would have thought that every single benchmark that existed aside from, speed and cost.
Jordan Wilson [00:23:52]:
Right? But when it came to raw capabilities, you would have assumed that Mythos was gonna wipe everything out, but it wasn't like that. Right? So if you look at, like, GBQA diamond, not not the top. Right? It's it's behind most models from OpenAI. It's behind Gemini 3.1 pro. So GPD QA diamond is a difficult science reasoning benchmark made up of expert level biology, physics, and chemistry questions. So it's often used to test whether top AI models can handle PhD level reasoning, not just general knowledge, which to me is interesting. Right? Because Anthropic made all this hoopla about its use on bio phys biochemistry, all of these things as one of the safeguard reasons. But OpenAI's in Google's models are far ahead of them in this benchmark.
Jordan Wilson [00:24:42]:
Same thing, with the artificial analysis long context reasoning benchmark. So this is the benchmark that tests whether a model can reason across very long inputs, right, and multiple documents, not just retrieve a single fact from a long prompt. And I think this is the type of, for a lot of general, knowledge workers. I think this is an important benchmark that tells you something, and this is where I think I've seen anecdotally where g b d five five pro and, Fable, I'm like, no. Five five pro is still much better because these are the type of types of prompts I'm using. Multiple document dumps, having it, retain context, go through multiple rounds of research, creating multiple artifacts. And at least for me, when I can even get it in on a single prompt on a $200 plan, I'm like, Fable five, not that good. Right? If I'm being honest, not that good.
Jordan Wilson [00:25:35]:
You know? It it may be if if if I'm saying Five five Pro gives me a level on this, you you know, Fable five gives me c. Whereas before, you know, Opus gave me a d, but it's still not not close. And, again, that's just not my personal opinion and, that I've confirmed anecdotally. That's according to the artificial analysis long context reasoning benchmark, which I think is an important one. So, yeah, there you go. Fable five, very far behind, Gemini three five Flash, Gemini three one Pro, even some, you know, open models, Mini Max, m three, you know, and then the the the best there is obviously, GPT five five. So when you talk about using, you know, different models, like inside of codex, inside of quad code, right, those are the types of things when you're giving it access to all of these files and folders, and you're having it run a longer task that requires it to maintain that context with a very long, hard input. It's not frontier.
Jordan Wilson [00:26:39]:
Right? It's not top class. It's not state of the art. Similarly, and this is one that I've talked about recently. Again, maybe some of these benchmarks just hit my personal use cases, but the if bench. Right? So this is the instruction following bench from artificial analysis. So this test whether an AI model can follow specific instructions and constraints, especially ones it has not seen before. So this helps show whether a model is truly good at instruction following or just optimize for common benchmarks patterns. And this is another thing I'm not even talking about, like, SuiteBench.
Jordan Wilson [00:27:13]:
Obviously, Infropic's models always do good there, but Infropic admitted that Claude cheats on SuiteBench. It looks up the answers, and then it tries to give the answers in a way that disguise it that it actually cheated. Right? So something like instruction following for me is extremely important, and I think for a lot of knowledge workers are. So if you are someone that, you know, doesn't just give these open ended, you know, here's some documents, go solve this. I take a lot of time because when I'm running longer longer running tasks, especially using things like goal mode, using planning mode, I'm putting a decent amount of time into giving a model proper instructions. And I saw this with four eight. Opus four eight was absolutely terrible at instruction following. It just refused to follow simple instructions that I laid out because right away, it came up, you know, kind of this new honesty thing that Anthropic says it's pushing.
Jordan Wilson [00:28:02]:
You know, it's like, Jordan, honestly, what you're asking for doesn't exist. And I had to go through back a couple times, tell it to browse the web. I'm not going to, Jordan, because you you don't know this, Jordan. You're a silly human, but this doesn't exist. And I had to go through multiple times to finally make it do follow instructions. So in tropics, instruction following are absolutely horrendous. So, you you know, on this chart here that our livestream audience can see, its models are consistently the absolute worst. Alright? I'm not saying models like, you know, OpenAI's and, Google's are the best.
Jordan Wilson [00:28:34]:
Actually, some of the best ad instruction following are, you know, Grok is really good. Mini Max is good. Mimo, Quen three seven. So, actually, some of the open source ones are really good at instruction following, but Infratix are a huge step down from not just those, but also the OpenAI and Google models. It's just not good at following instructions. So before you just look at INTROPICS blog post and say, oh my gosh. This is the best model ever, and it crushes every single benchmark. Any company that puts on a new model, they're obviously gonna cherry pick the best benchmarks, but at least for ones in my use because when I started first using Fable, I'm like, okay.
Jordan Wilson [00:29:08]:
Some things I needed to code a game, state of the art. Amazing. Right? I need to just create a three d world. Fantastic. Right? All these, you know, very specific singular use cases, great. When it comes to maintaining long context and following instructions across general knowledge work, compared to five five pro, I'm like, I don't know if I'm gonna use this. Right? Even up in between now and June 22, I don't think I'm gonna use it very much aside from testing it. At least for now, for me, g p d five five pro is better.
Jordan Wilson [00:29:41]:
There There's certain instances where it's like, oh, I need to do this in codecs because it involves more automations. It involves computer use, browser use, etcetera. I might be using Fable for some of those instances because in my testing, Fable is a little bit better than five five x high on some of those things, but not across the board. And then, obviously, it's so freaking expensive. Right? Actually, I had someone you know, some shared the artificial analysis, here thing on intelligence versus price. It completely reset the quadrants. Right? So now it's like business leaders when they're looking at these things, these certain graphs. Right? You have your quadrant.
Jordan Wilson [00:30:18]:
On the left side, you have, tasks that are cheaper. This is essentially the, amount of compute money that it takes to complete the artificial analysis test, which is a series of, like, 15 benchmarks. So, you you know, you always wanna be in this upper left hand quadrant, which means it's the, it's the smartest, the most intelligence, but it's the cheapest. Right? And there used to be, like, three to five, models that were in that quadrant. And now because Claude Fable five is so freaking expensive, it completely reset the graph. So actually, someone from artificial analysis reached out to me, like like, hey. What suggestions do you have? You know? And I'm like, I don't know. It's yeah.
Jordan Wilson [00:30:57]:
Because Fable five is so, so expensive, right, to get the same level of intelligence. We've been talking about that a lot here. And when we went over it on our start here series talking about token maxing to token efficiency, I think you have to start using as the subsidy era kind of starts to fade out. Right? Anthrapic is stopping it. Google is stopping it. Microsoft stopped it with, you you know, GitHub for now. We'll see what else they do. OpenAI luckily hasn't really stopped it too much yet.
Jordan Wilson [00:31:25]:
Right? But, essentially, they're starting to restrict where you could get, you know, thousands of dollars of of inference or, you know, tens of thousands of dollars of use if you're just looking at the API side for, you know, a $100 a month if you're on a, you know, higher tier. That time is going away. So I think, business leaders need to pay more and more attention about the cost per intelligence. You know, there's certain units you can look at, like, a cost per intelligence unit is extremely important. And now Fable five, yes, it's, you know, beats everyone else on the artificial analysis, but not I don't think it's worth the price. Right? It is more than twice as much as the next, tier, which were technically expensive. Right? The five five high series, Opus four eight. Right? Those are kind of the leaders in terms of the most intelligence.
Jordan Wilson [00:32:14]:
So you're paying more than twice as much for only a roughly 6% increase in overall intelligence. So I don't know, business leader, if you're out there spending, I don't know, a million dollars a quarter on AI or even $500,000 a year on AI, Are you gonna wanna double that for incremental increase? I would say probably not. But again, that's for you to decide. So let's talk about some more areas where they did really good. So in anthropics biology exercise, Mythos, assistant generalists beat world leading specialist teams. So, yes, in a setup, they gave, you know, essentially generalists, you know, people that had a baseline understanding of biology, but were not experts. They gave them mythos family levels, and they had them go against world leading specialist teams and the generalists with mythos, you know, crush the specialist. And I think that's really indicative to what having a state of the art model can do.
Jordan Wilson [00:33:18]:
If you speak the language, there's no right and we talk about that with, GDP valve all the time on this show. There's no more there's no more, comparison. Right? If you know what you're doing, if you understand the language, and if you know how to work with a large language model, you're gonna beat the world's leaders every single time in in tropics, kind of experiment here show showed that. So graders said that a six in this, experiments, graders said that a sixteen hour output normally needs forty to ninety five working days. So what the generalist will able to accomplish with mythos in a sixteen hour output compared to what would normally take expert teams multiple months. Okay? So that shows you the step change here that a model like fable five and mythos five unlocks. So this is the very power though that Infrappy called dangerous days before shipping it, which is interesting. And we're gonna get to here in a minute.
Jordan Wilson [00:34:18]:
In the viral, like the viral nature of this model can't be denied. Right? If you'll look on Twitter, LinkedIn, wherever, you're gonna see some extremely impressive examples that were just literally not possible before. Right? And I'm not just talking about pelicans riding a unicycle or whatever that SVG is. Right? The demos now that you're able to write whether you're creating full games, three d worlds, simulations, websites. I do think for individual artifacts, especially on the creative end, Fable five is now in a league of its own, even compared to, I think, GBD 5.5 pro when you're talking about those single artifacts. Right? Coding a game, you you know, creating a three d world. But, I mean, here's my question to a lot of people. Maybe that's part of your job, but I'd say so many of these things that you see online that people think that means that a model is good.
Jordan Wilson [00:35:14]:
What percentage of people are creating three d worlds for their job? Right? If you're a video game designer, sure. An artistic director. Sure. But I don't know. I would venture to say, like, point 01% of The US population is using AI models for things like that. So, you know, a lot of, outputs were, you know, people that shared about where mythos or fable, you know, tested, fixed, shipped, and documented the entire thing. So non engineers can now prototype work that once needed teams of engineers using AI. Right? So even if you look in the AI, phase, if you go back and look pre reasoning models, what a nontechnical person, can accomplish with a, something like Fable used to take a team of AI enabled engineers.
Jordan Wilson [00:36:07]:
So, Fable's ability to kind of run those long tasks, with singular outputs that are very focused and don't require a ton of instruction following. You know, people have been sharing, like, you know, go clone this Pokemon game. Go clone a, b, and c. Right? Go clone replic. Go clone anything. It does a wildly impressive job doing those things. But it's not all amazing demos, benchmarks, and roses, because I think we have to now tap into the four downsides. And I'm gonna go through these quickly, because the show's already running long.
Jordan Wilson [00:36:41]:
So Fable, obviously, on the cost, $10 per million input and $50 per million output. So it is more than double the cost of Opus 4.8, and now it puts it more in line with g b t five five pro, which is the world's most expensive model. So subscription access though ending soon. So, yeah, this whole, democratizing AI, yeah, could be out the window. We might now be on the AI for the versus AI for we side of AI. So beyond who pays, though Yeah. So number one, good luck because from an affordability, I'm not paying for g p d five five pro on the API. Absolutely not.
Jordan Wilson [00:37:20]:
I'd go broke. Right? Same thing. I'm not gonna be paying for Fable five, especially if g p d five five pro is continued to be offered in a subscription. So, beyond the democratization, there's another big downs. Actually, let let me talk about this. So the downside number two is anthropic decides where Fable can work. So things like cyber biology, chemistry, and model copying prompts can be rerouted to Opus 4.8 without telling you. And Anthropic says that over ninety five percent of Fable sessions avoid this fallback, but still 5%, that's a decent percentage where you're just not going to be able to use Fable even if you're paying for it.
Jordan Wilson [00:38:09]:
And some of these limits trigger with no notice to you at well or just get outright refusal. So I talked about, you know, how Dylan Patel shared something about semi analysis, some of their engineers getting, refusals. You know, here, very well known, and, hopefully, I get his his name right. Daria Unatmas, is a well known professor and biomedical researcher working on AI capabilities. He couldn't even use right? He shared this online. He couldn't even use, Fable because it brought in the memory and the history of his chats. Right? That's another thing. So even if you've worked on things in biochemistry, testing models before, and if you try to start, a new, chat using Fable, you might get outright denied.
Jordan Wilson [00:38:57]:
So that's downside too. Improvict decides where Fable can work and who can't, and you also have to think about your chat history and your memory if you're using it on the, quad on the website might be impacting your ability to even use it at all. Alright. So downside number three is the invisible downgrades can quietly bake break trust. So those hidden safeguards can silently weaken Fable on AI and machine learning research work. I do assume eventually I will run into this because I'm constantly doing some meta prompting inside large language models, showing them my input, showing them the output, and then copying and pasting the chain of thought to look at the tool calling, and then trying to make my prompting approach more token efficient. I assume once I do that a couple of times on Fable, it's not gonna work. I'll share that if I actually have the limits to do that testing.
Jordan Wilson [00:39:45]:
But Infropic estimates that these silent limits touch about point o 3% of traffic. But critics say that this quiet weakening is more dangerous than an honest refusal. So, yeah, they're just gonna roll you back and not even tell you. That, again, unprecedented move from Anthropic. So it's not like it's saying, oh, you've been downgraded in this case to using, you you know, Opus 4.8. It's just not gonna tell you. Right? And Anthropic says in its model card, this is more about, anti, or sorry. It's it's it's about stopping, model distillation or competitors from gaining an unfair advantage by using their models.
Jordan Wilson [00:40:25]:
But to me, it's absolutely asinine. Obviously, you want to do everything you can to protect against, distillation. But do an outright refusal. Say, you can't use our model for that versus giving them intentionally, you right. You're in theory, you are intentionally making the outputs worse and not telling the user. That is a huge miss like, a huge trust barrier. Right? And this is again what semi analysis, you know, shared about when they were trying to do these things. So they said, anthropics model will not help you if it thinks your ML research, ML engineering is interesting and or will secretly degrade its IQ so that the average engineer won't notice.
Jordan Wilson [00:41:11]:
We are already seeing Anthropic's latest models, moderation filters, our GPU inference research and programming. Alright. And downside number four, the last one, data retention. We already touched on this, but I think this becomes a deal breaker. I do hope Anthropic does the right thing and revisits this because I don't think this is, a good idea, especially with Anthropic's recent history of keeping private data private. So Anthropic does retain all Mythos class prompts and outputs for thirty days. So that breaks through any, zero data retention or HIPAA style setups that you had previously, but Infropic does say it doesn't train on those input and outputs. But, I mean, those four catches right there are huge.
Jordan Wilson [00:41:59]:
But the model itself actually adds a fifth risk, and that is Fable confidently faking its own work. Right? So, Anthropic, to their credit, does share about these things in the model card, but most people don't read the model card. I created, like, 20 different notebook l m audio overviews on different aspects of the model card. Yeah. I'm a dork. But in an audit of 886 real internal uses, Anthropic studied how mythos falls short in this regard. So in 16 of those sessions of the 886, mythos or, Fable claimed untested work had been fully tested. So not just, skirting around the issue, but said, nope.
Jordan Wilson [00:42:43]:
This has been tested. You're good. It's correct. This is verified. We tested it. So, yeah, even though it didn't so intentionally lied. So we could skip a human review phase. That's wild.
Jordan Wilson [00:42:56]:
That's wild. Right? So human verification and not raw capability, I think becomes your team's real bottleneck. Also, as we wrap up here, you have to keep in mind about Anthropics IPO. Alright. And I think you have to keep this in mind with all companies that are about to go public, but surprisingly, we haven't seen this from SpaceX as they go public here in like a couple of days. We haven't seen this from OpenAI, but we've seen a lot of this from Anthropic, which is surprising. But if you don't understand the storyline, I think your business and your budget just becomes another character in Anthropic's story. I think you have to follow the money.
Jordan Wilson [00:43:37]:
Right? Because Anthropic just came out with this big blog post, this big statement about how, AI, development needs to slow down. Yet they launched what is the most powerful model in the world, and they've shared their progress on recursive self improvement. So if you have models that are, more capable than anything else and they're improving themselves, but you're the ones calling for a slowdown, but you're also trying to go public. Right? I think the mythos and the fable really just weaves into this long term narrative that anthropic has been painting this picture of, oh, mythos is so powerful. It's so scary. The world is not ready. Don't worry. We are the answer.
Jordan Wilson [00:44:22]:
Right? They painted the story of the villain and how dangerous it is so they could be the hero. And then they came out with Fable five, and they said, don't worry, enterprise. Don't worry, Wall Street. We are here to protect you. This is Fable five. Right? So you can go have a little taste, but we're gonna start charging you through the ears for this model, here in two weeks just in time to re ramp up our revenue right before we go public. So hype sells capability, safeguards sell control, credits sell their future revenue. But all I'm saying is maybe Fable five is gonna change your business.
Jordan Wilson [00:45:01]:
Maybe you have access to mythos five. I'm not saying that that's not the case. All I'm saying is don't believe anyone's hype, Not anthropics, not open AI, not Google, not Microsoft, not anyone else. When a new model drops that promises world breaking capabilities, you have to do your due diligence. Is it safe for your company to use? What's the ROI? What are your use cases? What are your internal benchmarks? So before we all go, just see everything that's going viral online, take the time, understand exactly what these new models and the new capabilities mean for you, your company, and your career. So that my friends is a wrap on a longer show. I hope this one was helpful. If so, please let me know, by number one, subscribing to the podcast if you're listening.
Jordan Wilson [00:45:50]:
So, if you're listening on Spotify, appreciate that. Apple Podcasts, make sure to follow, the show. Leave us a rating if you could. Leave feedback. Alright? I I I always go through the comments when I can. And then when you're done, make sure you go to youreverydayai.com. Sign up for the free daily newsletter. Thanks for tuning in.
Jordan Wilson [00:46:06]:
This was a longer one, but I still hope to see you back tomorrow and everyday for more everyday AI. Thanks, y'all.
