Episode Categories:
Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
Managing AI Cost Control: Why Your Chatbot Bill Is Now a Board-Level Concern
The rapid evolution of advanced AI—from metered utilities to sudden budget overruns—has fundamentally altered the economics of intelligent automation in business. Recently, critical changes in how large language models (LLMs) like Anthropic’s Fable Five, Microsoft Copilot, OpenAI GPT, and others are priced, accessed, and used have forced organizations to rethink deployment at every level. For those seeking actual answers on precise cost-control measures and strategic operational shifts, the lessons emerging in real time matter more than ever.
Subscription AI Cost Trends: From Unlimited Usage to Metered Utilities
Frontier AI models, once available through flat subscription fees, have quickly moved to metered billing. In practical terms, business users previously enjoyed seemingly limitless runs of agentic AI—generating expert-level output and automating hours of workflows for a predictable monthly price. For at least three years, this model enabled explosive adoption: $20, $100, or $200 monthly plans took center stage as teams used state-of-the-art models such as ChatGPT, Claude, Gemini, and Copilot, often running agents for hours without worrying about additional costs.
This model is ending. Anthropic’s Fable Five—the world’s most powerful model at the time of this analysis—shifted out of included subscriptions. Major vendors like Microsoft (Copilot), Google (Gemini), and X AI (Grok) followed suit, introducing AI credits, strict usage limits, and shared pools. The era of “use what you want and pay once a month” has been replaced by tracked consumption and tiered meters, even at the highest tier plans.
The Chasm Between Token Pricing and Actual Spend
While the per-token cost for AI fell by as much as 98% over three years, total spend soared. This is due not just to more sophisticated functionality but to the vast expansion in the number of tokens consumed per task. Agentic models now work for hours, processing longer context windows, leveraging multiple subagents, and orchestrating complex workflows automatically. Workloads that previously demanded simple chat interactions now use billions of tokens weekly.
For context, maintaining a $200 monthly subscription for OpenAI Codex yields approximately 2 billion tokens of usage per week—spend that would equate to a staggering $200,000 per month under new API pricing for Fable Five. This cost spike, invisible for those on all-you-can-use plans, becomes a critical financial consideration the moment usage moves to metered access.
Enterprise AI Budgets: From Uncapped Growth to Strict Spend Limits
A key outcome of these new billing regimes is direct spending controls at the board and C-suite level. Across the tech landscape, companies have responded with firm guardrails:
Uber exhausted its entire AI coding budget for 2026 within four months, underestimating usage growth and the velocity of token consumption.
Tesla capped employee AI tool spend at $200 per week—a fraction of what even individual developers at major tech firms were accustomed to.
UBS surveys reveal at least 60% of enterprises now throttle AI spending, with projections suggesting that figure will rise.
Spending restrictions are no longer an isolated practice—they have become standard for both resource allocation and risk mitigation. Notably, some small- and mid-sized businesses may temporarily find strategic advantage as larger organizations restrict experimentation and high-volume utilization.
Token Efficiency vs. Token Maxing: The Shift in AI Strategy
A dramatic operational pivot has taken place: from token-maxing behavior—encouraging as much AI usage as possible—to a new emphasis on token efficiency. Research from Stanford’s Digital Economy Lab highlights that modern agentic models require up to 1000 times more tokens for complex coding or operational workflows versus traditional chatbot interaction. Variability is significant, making forecasting and cost control challenging for finance leaders.
Some models, particularly those from Anthropic, display significantly lower token efficiency yet remain preferred for quality. This dichotomy has resulted in well-publicized incidents, including an alleged $500 million accidental overrun on API billing, highlighting the risk when usage goes unchecked.
Practical Solutions: Model Routing, Open Source Adoption, and Spend Optimization
Business-focused organizations have demonstrated multiple tangible techniques for containing and optimizing AI spending without sacrificing outcome quality:
Model Routing: Companies like Coinbase have shifted their orchestration layer from premium models to a mixture of open source options (GLM 5.2 and Kimmy 2.7), reducing spend by half at peak while increasing token throughput and maintaining output quality.
Third-Party Routers: Growing adoption of services such as OpenRouter Fusion, Perplexity Computer, Merge, and others allows businesses to plug in multiple models and automatically route tasks by cost and complexity, further slashing API bills.
Fine-Tuning and Specialist Models: With offerings such as Microsoft Foundry and Thinking Machines Lab, organizations can now have tailored models fine-tuned for specific domains or tasks—beating state-of-the-art frontier models at a fraction of the cost (as much as 14x lower, according to recent studies).
Strategic Caching and Session Management: Automatic compaction, chat resets, and reuse of prompts ensure reduced overhead and avoid runaway consumption from repeated or redundant calls.
The Seven-Step Framework: A Playbook for AI Cost Governance
The operational reality now demands a disciplined, systematic approach to AI expense management. Seven steps have proven effective:
Metering Usage: Achieve granular visibility into token spend by team, workflow, and model—building or adopting tools for precise measurement.
Setting Budgets: Cap expenditures with alerts for overages; always relate limits back to mission-critical outcomes, not blanket cuts.
Defaulting to Cost-Efficient Models: Prioritize open or low-cost models for routine and non-sensitive workstreams, ensuring due diligence on IP and privacy.
Routing by Task Complexity: Assign advanced (and more expensive) models solely to high-difficulty or high-risk tasks; keep everyday automation on leaner models.
Maximizing Caching: Reuse repeated prompts and manage session sprawl with smart resets and compaction.
Implementing Fine-Tuning as a Service: Adopt fine-tuning strategies to create specialist models tailored for organization-specific workflows.
Controlled Escalation: Reserve premier model calls for genuinely complex or business-critical issues, ensuring the board that costs remain justified.
Outlook: The Future of AI Spend is Smarter, Not Harder
Unlimited AI access for flat fees is vanishing, replaced by precision management, rigorous efficiency, and a marketplace for task-appropriate models. This shift is not a future threat—metered pricing is now standard for major platforms, and large enterprises are already restructuring to adapt. Those who invest in advanced routing, fine-tuning, and analytic oversight today will not only contain runaway costs but also enable scalable, dependable business value from AI across all levels of the organization.
Forward-thinking companies have already moved from experimentation to disciplined practice. Cost control in AI isn’t an option; it’s a strategic imperative.
Topics Covered in This Episode:
- End of Unlimited AI Subscription Plans
- Anthropic Fable Five Subscription Removal
- Copilot and Grok Switching to Pay-Per-Use
- Enterprise AI Cost Control Challenges
- Token Consumption in Agentic AI Models
- Board-Level AI Spending Concerns
- Strategies for AI Spend Optimization
- Fine-Tuning and Multi-Model Routing Solutions
- Seven-Step AI Cost Reduction Playbook
Episode Transcript
Jordan Wilson [00:00:18]:
Most AI tools at work don't know anything about you. Every time, you're stuck reexplaining your projects, your team, what actually matters. Slack just fixed that. The all new Slack bot is a personal AI agent built into Slack, and it starts with your context, your messages, files, and channels, translation, and AI that already knows how you work. Check it out at slack.com. For the past three and a half years, AI has almost felt too cheap to be real, especially for power users during early twenty twenty six. I mean, you pay a monthly fee. You open chat GBT or Claude or codecs, Gemini Copilot, whatever, and then you have agents run for hours completing your actual work with your contacts without paying an additional penny on top of that subscription.
Jordan Wilson [00:01:13]:
But that version of AI is ending, and it's gonna hit even harder later today. Why? Well, that's because Anthropic's fable five only has hours left until it's no longer included in paid subscription plans. And that's not just an anthropic story. And it's the clearest signal yet that Frontier AI is becoming a metered utility. I mean, Microsoft GitHub moved Copilot into AI credits. Google changed Gemini limits around compute complexity features and chat length, and Microsoft started pricing Copilot co work through credits, not included subscriptions. And even NeoCloud x AI's Grok has been moving users toward share weekly pools and extra usage credits. So with the explosion of agentic models that can now work for hours with expert level artifacts, the question is no longer, how how do we get everyone in our company using AI? The board level question is, how do we use more AI without bankrupting our budget in the process? And that's exactly what we're gonna be tackling today.
Jordan Wilson [00:02:21]:
So here is the big picture. Frontier AI has definitely now become a meter utility. And the wild west days of using whatever model you wanted around the clock and just paying a $20 or $200 a month subscription fee and nothing more, those days are pretty much gone. Like I said, for three years, we've been doing this seemingly unlimited use of AI. And if you go back to literally back in 2023, even when the models were preagentic, I even said, eventually, we're gonna see $202,000 subscription plans. And I think even at that time, it's still gonna be a steal. And now when I throw some numbers out, you're gonna see that even these $200 a month plans or if there isn't even a plan on top of that someday will still be a steal considering what you're gonna start paying tomorrow. That's because even though the price per token has fallen by literally, like, 98% over the past three years, depending on what tier you're looking at, the agentic capabilities and the amount of tokens that these models are using has legit skyrocketed.
Jordan Wilson [00:03:29]:
And over the past two or three months, AI has quickly gone from, let's use as much as we can in token max to putting strict spending limits. And sometimes, you know, developers at huge companies aren't even to use AI as much as everyday nontechnical people like you and me. And I think the Fable issue today, Infropix Fable five leaving subscription, in a couple of hours only heightens that chatbot's bill problem. So on today's show, here's what you're gonna learn. You're gonna learn why Fable five subscription exit is the cleanest and loudest warning shot for subscription AI and your team's AI strategy. You're gonna know why GitHub, Copilot, and Grok swapped unlimited access for pay per use meters. You're gonna understand why companies are capping spend, but still trying to use more tokens and how they're doing it. And I'm gonna leave you at the very end with our seven step playbook to control runaway AI spend.
Jordan Wilson [00:04:29]:
Let's get into it. Welcome to everyday AI and our start here series. This is the essential podcast series to both learn the AI basics and to double down on your AI knowledge. So if this is helpful and if that's exactly what you're trying to do, make sure you go to starthereseries.com. That is gonna give you exclusive access to our inner circle community. You're gonna be thrust straight into our start here series space where you can go listen, to every single start here series, the simplest way to do it on the playlist that we have there. As well, you can go read about it, connect with other people who are trying to do the same thing. And if you missed our last start here series episode, we talked about the desktop agent lingo simplified, going over goals, loops, plans, sub agents, and how it works in codex and cloud code.
Jordan Wilson [00:05:18]:
And well, if you're doing all of those things, chances are, yeah, your bills might start getting more and more expensive, which brings us to today's episode, AI cost control one zero one, and why your chatbot bill is now becoming a board level problem. So over the past couple of months, we've seen this. And it's not just Anthropic's fable five. I mean, let's go back to May, right after Google's IO conference, their big yearly conference. For the first time, Google started putting actually pretty strict restrictions, on their Google Gemini. So it's actually the first time I'd ever, run into Google Gemini limits on their $20 plan. Up until then, it had been pretty generous. And, you know, after they made this shift of actually starting to count longer, you know, taking things like complexity, the model you use in chat length, all of a sudden that $20 plan that used to feel pretty great started feeling not super useful.
Jordan Wilson [00:06:22]:
Then in June, you had Microsoft's GitHub Copilot replaced their premium request with token based AI credits. So the unlimited plan was gone for Microsoft's GitHub Copilot. The same thing with Copilot Cowork. Microsoft is starting to go with task level credit pricing and no longer including essentially unlimited cowork usage after they moved it to general availability. Then couple days ago, we also saw Grok, from x a I space x a I, I think, is technically their new name as of a couple hours ago. Right? They started rolling out, the credit based system as well, which is pretty telling considering, well, they're kind of like a NeoCloud leader now. Right? They are, one of the companies that has technically extra compute, and they're even charging for. And like I've already said, today is maybe the day that all of this gets thrust more squarely into focus because of Infropic's fable five.
Jordan Wilson [00:07:25]:
And why is that important? Well, right now, at least, I mean, we'll see what happens over the next couple of hours because we've heard, rumors that we might get a GPT five six from OpenAI today, sometime between today and Thursday. But, at least for the next couple of hours, Fable five is the most powerful AI model in the world. And at least for a couple more hours, it is still available in a subscription. But after today, you are going to be paying the full API price for that. So we've gone over the, Fable five drama kind of enough, but, you know, they put it out there. They had to take it away because of some problems with the US government, and they brought it back a couple of days ago. But they said only included in plans through July 7. And they said, hey.
Jordan Wilson [00:08:13]:
Maybe one day when we have enough compute, maybe we'll bring it back, but no guarantee. And why does this matter? Because like I said, for the past three and a half years, we've been able to you know, whether you have a $20 plan, $100 Pro Max plan, $200 Pro Max. Right? But for a very low price, and I know that's all relative. Right? For many people out there, I understand a $200 a month plan, especially if you're paying for it yourself, doesn't seem like the cheapest thing. But let me just show you why it actually is. So now starting tomorrow, technically, 11:59PM Pacific time, you're gonna be paying, $10 per million input tokens and $50 per million output tokens for, Fable five. Let me just put this into comparison now. So you're probably saying, alright.
Jordan Wilson [00:09:02]:
What's the big deal? Alright. Like, Jordan, if you're paying, you know, $200, you know, a month for, you know, which I am. I'm paying $200 a month, for the Claude 20 x max plan. I'm paying $200 a month for the chat g b t pro plan. You know, I paid the $200 a month previously for the Google, Ultra plan, so I paid for all of these. So you're saying, okay. What's the big deal? You know, if it just moves to API only, can't you just use that money for the API credits? Yeah. No.
Jordan Wilson [00:09:31]:
K. So granted OpenAI has still ridiculously generous limits. Alright. So let me put that out there. You're normally not getting that level of spend on these other plans, but I've been averaging, over the past couple of weeks about 2 ish billion tokens a week on the $200 a month, OpenAI codex plan. Okay? Which is included with chat g p t. Right? So it's the same plan. So that at a at a very generous blended rate, right, of, a blended rate of input and output tokens, that is a $200,000 a month, an traffic bill.
Jordan Wilson [00:10:18]:
If I were to take everything I'm doing inside of codex and I'm only paying $200 a month. And if I were to say, in theory, move all of that into, fable, that is $200,000 a month. Okay. Now you see why this has been a such a big discussion in the AI world about no longer having this democratized AI, access to AI. We'll see, you know, what it means with, OpenAI's models moving forward. They did say that their, you know, new g p t five six, kind of the different tiers, Sol, Terra, Luna, and, you know, Sol Ultra would be included in subscription plans. So maybe OpenAI, maybe one of the only few players left that has yet to, you know, kind of knock knock out that subsidy, but that is still the reality, because you have to assume that one day, that's gonna be the status quo. And maybe OpenAI might be able to do it for another three months, three quarters, three years.
Jordan Wilson [00:11:27]:
We're not sure. But eventually, AI is going to get more and more expensive, because the models themselves, in theory, are gonna be eating up more and more tokens. And this is exactly why we're seeing that token maxing to token efficiency shift. And we did cover that on the start here series volume 27 or episode seven eighty nine, so go check that out. But we've seen it in some of the biggest companies in the world. So as an example, we saw the very, you you know, popular story. Uber reportedly burned through its 2026 AI coding budget in just four months. Tesla, we recently saw capped employee AI tool spend to just $200 a week.
Jordan Wilson [00:12:08]:
And then UBS says that 60% of interviewed enterprises are throttling AI spent already. Right? So I do assume that number will probably go up if UBS does the same survey quarterly. I don't know if they do. I would assume that's gonna go up to about 70 to 80%, by the end of the third quarter because I don't think the, kind of the token efficiency, wall has hit the majority of enterprises just yet. So and this is actually everywhere. Right? I even had a conversation with a friend who is in software development at one of the Mag seven companies, you know, and he said his spend is a $100 a month. Right? So that $200 a week at Tesla, although when you're looking at API credits, you're like, that's like nothing. Yeah.
Jordan Wilson [00:12:58]:
Imagine having, like, a $100 a month. So that is the reality. So I also want to tell this to people out there who maybe don't talk to a lot of people at Fortune 100 type companies and, you know, big Mag seven type companies. We are seeing actual real spend limits on the amount of AI that these employees can use. And that actually, I think, for SMBs and medium sized enterprises might create a strategic advantage. Right? If the biggest players are maybe, you know, having to reel in their spend and maybe part of that was their own fault for pushing these token leaderboards and all these other things, I think there is an advantage to be had for everyone else. Alright. Let's take a quick break for a word from our partners.
Jordan Wilson [00:13:45]:
Here's what most AI tools still can't do, work outside their own little box. The all new Slack bot just changed that. It's your AI teammate inside Slack, and now it can read, write, and act across the other apps your team already uses. No more tool switching. No more reexplaining yourself every time you open a new tab. One ops team at Engine says the summary feature alone saves them fifteen to twenty minutes of use. See what Slackbot can do at slack.com. Okay.
Jordan Wilson [00:14:18]:
So why now? Right? Why are we seeing this, you know, your chatbot bill becoming a boardroom problem now? It sure wasn't at this time last year. It wasn't really at the end of twenty twenty five. And I think to make a very short story of the last six to eight months, of AI development, if you listen to the show, it's no surprise. Maybe if you're newer here, let me fill you in. Infropic set the pace at the end of twenty twenty five, with Claude Code bringing that to the desktop, as well as Claude Co work in early twenty twenty six. And then I think the rest of the way has been blazed, by codex, cursor, and some others. But we are now having extremely powerful AI models even for nontechnical work. Right? When I talked about spending, you know, 2 ish billion tokens a week, I'm not doing a ton of traditional coding.
Jordan Wilson [00:15:10]:
Yes. I'm building a bunch of programs that I use all the time, but a lot of my token spend is just doing right scheduled work that I wouldn't normally be doing. Right? Doing a bunch of research, creating a bunch of just documents. And the reason why we are seeing this explosion is because these capabilities are becoming more powerful, and they're becoming easier for nontechnical people to use. So you have more and more people starting to automate their everyday work with these desktop agents that can work twenty four seven and can work for hours. Right? I remember even in the earlier days in 2026 with Claude Code, I would have to really push the models, to work for maybe, like, an hour or so. But now with the combination of sub agents and plan mode and goal and and all those things that we talked about in the last start here series, every single day, I'm usually having multiple threads running for eight to sixteen hours and creating amazing outcomes. So it's all about the amount of tokens being spent.
Jordan Wilson [00:16:10]:
So a Stanford digital economy lab study this year showed how much more impactful and how much more token hungry today's models are than the old chatbot only non reasoning models. So, essentially, they found that agentic coding tasks used a thousand times more tokens than a code chat or just a normal reasoning chat. So the same task runs varied up to 30 x making forecasting brutally hard. That's the other thing. And I think that, you know, say what you want about your model of choice, but certain models aren't exactly to co token efficient. Right? If you look at the artificial, artificial analysis leaderboard, the anthropic models are extremely token inefficient. Yes. They are some of the best in the world, but if you are not keeping a close eye on your AI spend and you're just giving these very powerful models access to all your data, your repos, your, code, you you know, your code base, whatever it is, and you're giving it this big goal and then just not really watching it.
Jordan Wilson [00:17:13]:
Right? We've seen models like Fable, like Opus be two to three x, more token inefficient than some of their closest competitors from the other labs. Right? From, you know, Google, OpenAI, you know, and then even some of the open source models as well. And then there is the rumored story, right, that a single enterprise wasn't keeping a close watch on their API bill and accidentally spent $500,000,000, with Anthropic. I still don't know if that is true or not, but there were some reputable, you know, news organizations reporting on it. But that is the risk, and that is why this essentially AI bill has gone all the way up to the boardroom. And it used to be a couple of months ago. Right? The, the message being pushed down was let's increase our seats. Right? Hey.
Jordan Wilson [00:18:10]:
We're paying for, you know, 5,000 seats, but only 2,500 are being used. We need to increase that. And now that has shifted because a lot of those seats, quote, unquote, that they're paying for before now have meter billing maybe on top of a very limited, you know, monthly or weekly quota. So now, you know, there was all of this push over the end of twenty twenty five and early early twenty twenty six. So let's get people using this AI as much as possible. Right? We all got this taste of the gateway drugs. Right? The 2025 models were fantastic. The harnessing improved.
Jordan Wilson [00:18:46]:
The tool calling improved. Models reasoned and thought by default. Right? So all of a sudden, you had companies that were smart and doing things the right way. They showed measurable ROI, and they're like, let's get this thing. Let's use this as much as humanly possible, and then the whiplash effect is hitting us all in the face. And that's why we have Jovan's paradox, and that explains kind of this concept of cheaper cheaper tokens, but larger bills. So if you don't know what that is, that's essentially the idea that as a resource becomes more efficient to use, its consumption often increases rather than decreases. So Juman's paradox means cheaper units can drive more total consumption.
Jordan Wilson [00:19:26]:
Simple. Right? But cheaper tokens also created those longer context in more agents, more retries, more automation, and sometimes more loops of tool calling in a bad way. Right? Just models calling tools when you're not looking at them and they're going over and over and over, and you're maybe not monitoring that. And that's why these bills are growing because of demand expanding faster than the unit prices fell. And also, sometimes, the human in the loop being a little lazy and not keeping an eye on because they're like, hey. I can give this thing a goal and give this all this data, and it's gonna work for sixteen hours, and it's gonna get the job done. Maybe not looking at how much that job is actually costing or how many, you know, times a certain model is trying to go fetch a website and failing, and it's just trying over and over and over. And maybe each of those tries, maybe there's hundreds or thousands of tries, and it's costing you maybe a couple of dollars every single attempt.
Jordan Wilson [00:20:26]:
So there's use cases already we've seen of companies maybe doing this the right way. And I think one of them is Coinbase. So what they did and they shared about it a little bit online is they swapped their default models, from essentially anthropic to GLM five two and Kimmy 2.7. And they shared that the combination of that, right, not every single model, right, but a lot of their, you know, orchestration. So, you know, still handing off, some of the heavier lifting to maybe an anthropic model. But a lot of that, you know, middle grunt work, the orchestration, the summarization, the writing, some of those things that you maybe are fine with a, you know, one b tier model. Well, they were just going open source with that and using something like z a i's GLM five two or Kimmy two seven. So combining that with caching, proper caching, and then difficulty based routing, well, they were able to cut spend to nearly half of what it was at its peak.
Jordan Wilson [00:21:30]:
Yet they increased their actual token usage. And in terms of output, it didn't fall off at least according to the company. And that's also led to a surge of a new breed of model routers. Right? And what's funny is, well, I told you guys about this more than eighteen months ago. That's why, especially those, you know, year end, prediction series. That's why you gotta listen to these and put these things into practice because I told you guys this was coming. Right? The mixture of models and some, some other things. But we've seen a lot of, I think, great innovation.
Jordan Wilson [00:22:09]:
Some of them like open routers fusion, perplexity computer, computer. Although, I wish it was a little more affordable. You you know, and you've seen some other ones as well. Merge, another great example. But, essentially, now you have these services that are kind of plug and play and will do this for you. So similarly, how Coinbase went in and did this all manually. Right? You are having some quote unquote third party services that are going to do a lot of the heavy lifting for you. Right? You entering your API keys, you know, connect your data as you normally would via, you know, easy one click, easy, you know, porting over connectors.
Jordan Wilson [00:22:50]:
And then, well, you can just hopefully reduce your API spend. There are obviously things you have to keep in mind and do your due diligence on in terms of privacy, data sharing, and all of those things. Right? I'm not gonna be vouching for all of these other companies and how they do that. You have to go do your own, kind of digging on those. But routing is just gonna pick the model, but tuning is, I think, where we're gonna start to see a lot of motes being built. Yes. That's right. Fine tuning models.
Jordan Wilson [00:23:22]:
I don't know if we'd say it's making a comeback because it didn't really go away. Right? But I think fine tuning was all the rage back in, like, you know, late twenty twenty three and 2024 before we had a handful of, you know, Frontier AI labs to choose from. And before, you know, bringing your data was difficult. Right? Back when, you know, we were talking about rag pipelines a lot, also fine tuning models was a huge deal. And I think that, again, this is something that we talked about eighteen months ago. Right? Several models kind of teaming up, not just mixture of experts. Right? Where you're looking at as far as first dense models, and giving a query to a big model and a big model just activating the, you know, the experts that it needs within those parameters that it carries. No.
Jordan Wilson [00:24:13]:
I think we're looking at the mixture of models as being something that is gonna become increasingly popular in the near future. And I actually hope that we see that from the actual frontier labs themselves. I would love to see that from, you know, anthropic, OpenAI, Microsoft, Google, etcetera. Right? Where even within something like cloud desktop or, you know, codex where you can go into a certain mode, and it will do it for you. We've kind of seen previews of this that I don't think were super well done, where you had these auto modes, and it decided how much reasoning it would want to use. But that was really a little bit more, at least for our non developers, nontechnical people. You know, it it it was more of, I think, just an an experiment, to maybe, you know, keep compute in check around new model launches. But I do think that the combination of mixture of models and fine tuning models, is going to help companies not just, keep spend in check, but actually build the modes.
Jordan Wilson [00:25:16]:
So this kind of flew under the radar. You you know, I saw this out when I was out at Microsoft build conference, but now Microsoft Foundry offers fine tuning as a service. Right? So as an example, you know, Foundry has a ton of models from all of these different providers. Right? If you just need one model and if it's 80% of your usage, let's just say as an example, is doing some financial modeling around a certain sector instead of taking one of the big state of the art trillion parameter models. Right? You can work with now Microsoft in Foundry. They do this as a service. They can find the right maybe a smaller or medium sized model and then just fine tune that or essentially create a version of the of a medium sized model that's more cost efficient, faster, and, well, just better because they're essentially going to train a version of the model just around your specific tasks. Speaking of that, that seems to be a, key area where thinking machines, labs is starting to play.
Jordan Wilson [00:26:16]:
Right? Former, you know, OpenAI CTO and co founder Mira Marathi, with Thinking Machines Lab. They just came out with a study that showed Bridgewater and Thinking Machines Lab reported a tuned specialist beat a frontier model at about 14 x lower cost. So that does seem something that Thinking Machines, is doing as well, essentially offering fine tuning API as a service. But when you have the big labs, right, and sure, we'll throw thinking machines in as a big lab in terms of, market cap valuation money raised. Right? You they're at least in that consideration. Microsoft, obviously, one of the biggest and best in the country. Sorry, in the world. Right? When you are starting to see fine tuning as a service, that is not, about model specialization.
Jordan Wilson [00:27:06]:
That is about the, the impending, reality of subscription models going away in enterprise companies having to look at that. Right? If you were paying for, as an example, if you were paying for 10,000 licenses of Microsoft, GitHub Copilot, and then all of a sudden they switched over to use it billing. Right? Are you gonna continue to do that? I don't know. Right? This is where you're gonna have to start looking at these alternatives. So here is your Monday move. You have to build the spend it router. Right? Whether you look at some of those other services, whether you start doing this and duct taping it yourself internally or whether you are just preparing for when the big buy apps offer this, you have to start putting this in place. You don't just cut AI blindly.
Jordan Wilson [00:27:55]:
That is the absolute worst thing you can do, and unfortunately, a lot of companies are doing this. So don't do that. Right? You have to start classifying work by value risk and complexity and start using some of these cheaper open source models. Again, go through all your due diligence on the privacy data, you know, IP side, all that. Right. But you should be looking at these tune specialists, routers, caps, and frontier escalation Because the winners ultimately are gonna be the ones, I think, like Coinbase that are finding ways to actually spend more tokens and to use more tokens. I don't think there was anything wrong per se with the token maxing era, but it was more of doing it irresponsibly on the API side and using the most expensive models. Right? I do think that you have to have some experimentation on these models that are a fraction of the cost, and you still get maybe, you know, 85 to 95% of the performance of the state of the art models.
Jordan Wilson [00:28:55]:
So here is your seven step playbook, to control runaway AI spend as we wrap today's show. Number one, you have to meter first. So you have to get token spend visibility by team, task, and model. Right? So similarly, how I said, hey. Right now, I'm going through personally 2,500,000,000 a week. I can break that down by project. I actually build a tool inside of codecs that helps me do that. Right? You can be doing these things as well.
Jordan Wilson [00:29:21]:
You have to start seeing where your tokens are going, who is spending them, what task, what model, and what are the outcomes as well. Number two, set budgets. After you do that, you have to be able to cap spend per team, user, and workflow with overage alerts. That doesn't mean you cut use. Right? So let's say you first meter and you say, here's what we're doing on our subscriptions. What is that, what is that spend that you're gonna have to do to get the same outcomes? Right? So don't cut budget without protecting your outcomes. Alright? Number three, you have to start learning to default to cheap. Right? I'm not saying the cheapest models, but now we do have models that are extremely capable, open weight models like GLM 5.2.
Jordan Wilson [00:30:08]:
You know, like, Kimmy, two seven. Right? We're gonna I'm sure have some other things from Quinn. They have some great models. Deep seek, I'm sure we'll come out with another good one. Again, go through all your proper, IP privacy security, but you have to start looking at these maybe open source, open weight models, whether you're running them locally or paying for them on the API side. They're normally a fraction of the price as the closed, frontier models. Number four, you have to route by difficulty. You shouldn't be, as an example, having Fable five on extra ultra reasoning, whatever it's called inside cloud code.
Jordan Wilson [00:30:45]:
You shouldn't be doing that to help you write better emails. Right? You have to send the right type of tasks to the right models. Then you need to understand your get proper, caching. You need to trim and reuse repeated prompts and start lean sessions per task. Right? Even simple things, like, if you're working on the front end using things like projects. Right? Certain, you know, certain providers do things better than others. Right? As an example, codex is great at compaction. Auto compaction is really good.
Jordan Wilson [00:31:20]:
You you know, Claude recommends as chats get longer to start new chats, to save on cost, but you have to start doing those kind of nontechnical best practices as well as caching and trimming. Number six, fine tune the repeats. Yes. If especially, I think if you are a, Microsoft organization, I do assume that we're gonna start seeing fine tuning as a service. Right? I think Thinking Machine Labs and Microsoft, offering this leads me to believe we're gonna see this probably within, nine months from all of the big players. You have to start looking at and understanding what it's what are we gonna find to. Right? What type of model is going to do 80% of our AI use for certain teams. And, well, if you fine tune a model, it might be a little bit of work and a little bit of cost upfront, but you have to be able to math the math because chances are it's gonna save you a lot in the long run.
Jordan Wilson [00:32:13]:
And then last but not least, escalate on purpose. Call the Frontier model only to close hard problems, especially if that Frontier model is not included in a subscription. And I say that as a very timely reminder. Yes. Fables, Fable five, couple more hours, and it's gone. We'll see what happens with OpenAI, g p d five six, Sol, and Sol Ultra. They did say it's gonna be included, inside codex. So who knows what's happening? But all I know right now as we wrap up today's show, the future of AI is token efficiency, and it is no longer, if you are the decision maker six months ago, regardless of what tool, what system, the right move was use as much AI as possible.
Jordan Wilson [00:33:05]:
And that's still me, maybe the case, especially if you're an open AI organization, because they are pretty much the only player right now that hasn't drastically, either cut, the subscription limits or that hasn't gone to true metered access. But chances are, some of your teams, some of your organizations most needed, most use AI tools have converted to metered pricing. So that doesn't mean stop. That just means be smarter. Don't stop experimenting and well, follow us daily because we're gonna be helping you every step along the way. I hope this one was helpful. If so, please make sure to go to starthereseries.com. Go sign up for our community.
Jordan Wilson [00:33:48]:
Go network with other people who are leaders in the AI industry, figuring it all out there. Also, if this was helpful, please subscribe to the show on Spotify and on Apple Podcasts. Thank you for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks y'all. Here's the thing about AI at work. You can't use what you don't trust. The all new Slack bot runs inside Slack security boundary, only sees what you've already allowed it to see, and never trains on your data.
Jordan Wilson [00:34:25]:
Personal AI agent with the trust to actually put it to work. Slackbot from Slack. Learn more at slack.com.
