Ep 789: Tokenmaxxing is over: The New Era of Token Efficiency and how Your Company Should Adapt

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


Token Efficiency in AI: Why Smart Companies Are Abandoning Token Maxxing

Enterprises that once equated massive AI token usage with productivity are facing a sharp reckoning. After months of treating AI token consumption as a corporate status symbol, significant financial waste and inefficiency have forced a new paradigm: focusing on token efficiency over raw volume. Businesses seeking AI ROI now measure impact per token rather than tokens spent—and this shift is rapidly separating sustainable AI strategies from costly misadventures.

This article explains the pitfalls of token maxxing, defines what tokens are, highlights four specific ways AI models use tokens, and details how businesses can extract real value by prioritizing token efficiency.

The Downside of Token Maxxing for Enterprise AI

For much of the past year, the practice of "token maxxing" dominated the enterprise AI landscape. Companies strove to demonstrate their AI innovation by pushing up token usage, such as running long agentic loops and tracking employees by tokens used, sometimes ranking staff with internal leaderboards.

One reported case included an internal Meta leaderboard where an engineer ran 281 billion tokens in a single month. This approach encouraged empty agent loops, generated non-strategic activity, and sometimes even prompted employees to create AI tasks with little business value simply to protect their jobs or climb leaderboard rankings.

Token prices have dropped significantly—GPT output token costs, for example, have fallen from $60 per million to just $8 per million. Despite this, overall usage has exploded by 100–200x as models become more agentic and reasoning-intensive, setting the stage for out-of-control spending. Notably, one unnamed company reportedly accrued a $500 million AI bill in a single month by not monitoring or capping usage, demonstrating the risk of unchecked token consumption.

Understanding AI Tokens: The Basics Business Leaders Need

A token is a segment of text processed by AI models. One token roughly equals three-quarters of an English word—so, a million tokens represent around 750,000 words. All AI costs, whether via API or subscription plan, ultimately convert to tokens, which are the real currency of intelligent output.

Four core types of token usage are essential for budget-conscious AI adoption:

  • Input tokens: Every word or file input increases usage. Larger files and longer prompts have a direct impact on cost.

  • Output tokens: All model responses—text, code, images, video—are charged as output.

  • Reasoning tokens: Modern models now "think" by default, consuming tokens internally for each step in their reasoning process. These steps can go unseen but dramatically increase costs.

  • Tool use tokens: When models use plugins, web search, file reads, or run code, each tool invocation eats up further tokens, especially in automated or repeating agent workflows.

AI Spending: The Cost Structure Shift

Initially, companies—especially the major model providers—heavily subsidized token usage under flat subscription plans. Individuals could burn through up to a billion tokens per day under certain plans, sometimes running $12,000–$17,000 worth of API usage, entirely covered by their subscription.

Today, those open-ended subsidies are being phased out. Providers like Google and Microsoft have begun enforcing hard usage caps, especially on business and enterprise plans. Once internal token allocations are exceeded, companies pay standard API rates—which means automated, unsupervised agents can quietly rack up formidable overages unless tightly managed.

Context Windows, Caching, and Unintentional Waste

Models process information in fixed-sized blocks called context windows, defined by tokens. If a process is not architected to cache prior context efficiently, recurring loops or long-running conversations can reprocess and charge for previously-seen content, compounding costs for no additional business value. Complex workflows set to loop over dynamic enterprise data—such as updating dashboards every hour—are especially prone to ballooning expenses.

Measuring ROI: Token Efficiency as the Critical Metric

The shift is toward token efficiency—extracting maximum business value and measurable output per token spent. This strategy involves comparing the cost and deliverables of AI-driven processes to legacy, human-driven outputs. Businesses now benchmark:

  • The value of the AI-generated output (e.g., a completed client deliverable or artifact)

  • The per-token or per-dollar cost to produce it

  • The overhead in terms of time and rework compared with previous (pre-AI) methods

A responsible approach means selecting the right model and configuration for each job and quantifying new value creation—not simply tolerating or encouraging runaway usage.

Model Selection: Intelligence vs. Cost

Token efficiency requires understanding two core variables: model intelligence (measured in scores like the Artificial Analysis Intelligence Index) and the cost per intelligence point delivered. For example:

  • On recent tests, Anthropic’s Opus 4.8 outperforms OpenAI's latest model on intelligence, but at a 50% higher token cost ($5,100 vs. $3,300 for the same business task).

  • Google's Gemini 3.1 Pro achieves top-tier results at less than 1/5th the token cost of Anthropic Opus 4.7.

  • DeepSuite and other agentic coding benchmarks consistently show certain models (e.g., OpenAI GPT 5.5) delivering better output at half the cost or lower compared with legacy high-cost models.

Building modular AI flows that can switch between models—and monitoring cost per intelligence point on a per-task basis—is now key to sustainable enterprise use.

Practical Steps for AI Cost Management

Enterprise leaders looking to optimize AI investments can take several clear actions:

  • Quantify AI output per dollar: Measure what deliverables cost to produce with AI versus traditional methods.

  • Build modularly: Avoid vendor lock-in by configuring systems to pivot between models as costs or capabilities evolve.

  • Choose the right model for each job: Small models may suffice for parsing PDFs and simple summarization, while pricier models should be reserved for complex creative or analytical tasks.

  • Actively monitor agents: Detect and shut down inefficient loops or unnecessary tool use.

  • Review costs regularly: Analyze ongoing API and subscription spending alongside business value generated—adjust agents and model selection accordingly.

Token efficiency is rapidly replacing brute-force token maxxing as the driver of true enterprise AI ROI. Companies that monitor value per token and relentlessly optimize their AI architectures will adapt fastest as the age of open-ended AI subsidies closes.

Conclusion: Activity Does Not Equal Value in AI

Raw token consumption does not correlate with business success; value is in measurable, efficient transformation. The companies that treat AI as a rigorously measured productivity asset, not a badge of technical busyness, will position themselves for competitive advantage as AI costs become a line item impossible to ignore.

For more resources and actionable frameworks, explore series and community platforms focused on practical AI mastery tailored for business outcomes.



Topics Covered in This Episode:

  1. AI Token Maxing: Rise and Fall
  2. Defining AI Tokens and Tokenization
  3. Four Main Types of AI Token Usage
  4. AI Agentic Loops and Token Consumption
  5. Corporate Token Leaderboards and Meta Example
  6. Risks of Unmonitored Token Burn in Enterprises
  7. Token Subsidies and AI Pricing Trends
  8. Measuring Token Efficiency versus Token Volume
  9. Benchmarking Models: Cost per Intelligence Output
  10. Shifting from Model Selection to Harness Efficiency
  11. Best Practices for Enterprise Token Optimization
  12. Monitoring AI Agents for Token and Cost Control




Episode Transcript 



Jordan Wilson [00:00:16]:
There's been this pseudo badge of AI honor over the past six months that makes no sense, and I think it has led a lot of enterprise companies astray when choosing the right AI strategy. What is that thing? It's token maxing. So token maxing is the corporate practice where companies push AI token consumption and correlate that AI usage as a direct metric for employee productivity or pointing to the fact that they're doing something positive with their AI spend. Like, if you're using billions of AI tokens a day, that means that you're living inside of large language models, and clearly, you are creating huge value for the company because more prompts has to mean more ROI. Right? But that strange AI power flex of early twenty twenty six has quickly revealed itself to be a false and sometimes dangerous metric by mid twenty twenty six. Because it turns out, all that token maxing leads to is a competition to see who can run the longest agentic loops regardless of if it actually accomplishes a company's goal. So what's the new metric? What should your company be looking at? That's token efficiency. In other words, getting the most bang for your AI token usage buck means not just burning AI usage for the sake of using it, but instead aligning your company's AI strategy with the right model used the right way for the right job.

Jordan Wilson [00:01:46]:
Weird. Common AI sense pays off more than a leaderboard showing who is using the most tokens. So on today's show, we're gonna break down the rise and fall of token maxing, not only showing you what tokens are and how they work, but also uncovering the proper best practices to put your tokens toward actual measurable outputs versus just lighting millions of dollars on fire. So here's the big picture. For the past six ish months, companies have treated heavy AI token use as proof of productivity. That's because token prices themselves have fell, like, 200 times. Right? If you remember back, you know, the earlier versions of the GPT technology were, like, $30 per million, you know, input and, like, $60 per million token output. Now it's, like, three and eight.

Jordan Wilson [00:02:43]:
Right? So token prices have fallen significantly, but the token usage has also gone up, like, 100 to 200 x as models get more agentic. And now that models are reasoning by default, AI agents make the number of tokens per task absolutely explode. So I think now in mid twenty twenty six and through the rest of the year, I think the goal, if it hasn't already in your organization, shifted away from token maxing in toward token efficiency. So stick with us on today's show, and you're gonna learn what token maxing is and how it led companies to waste a ton of money. You're gonna learn even more fundamentally what a token is and the four main ways that AI models gobble those token ups tokens up. You're gonna know why companies are shifting from token volume to token efficiency and how to turn AI spending into real measurable results. Let's get into it and start here with the basics. Yes.

Jordan Wilson [00:03:44]:
This episode is part of our start here series, the essential series to both learn the AI basics and to double down on your AI knowledge. So if that's what you're trying to do, make sure you go to starthereseries.com. That is going to give you free access to our exclusive inner circle community. It's free. There's no other way to get access aside from going to starthereseries.com. And once you sign up, you will be inserted in to our start here series space where every single episode is there for you. You can digest them all very easily. We have a playlist there as well that we keep updated.

Jordan Wilson [00:04:22]:
So start here series. If you miss our last episode, we talked about the 2026 LLM cheat code and 10 essential steps to get the most out of any AI chatbot. And if I'm being honest, that's probably one of the best shows we've ever done here out of 790 episodes. So make sure you check, that one out. That's volume 26 and episode seven eighty six. But today, let's talk token maxing. So token maxing is, well, thinking that however many tokens you spend means something. And in rare cases, it might.

Jordan Wilson [00:05:00]:
But I would say in most cases, it just means that employees have been given a green light to use as much AI as possible. They're probably not closely monitoring what their agents are doing. And in some cases, it's more of, you know, people trying to pound their chest and talk about how many tokens they're using. That's because, yeah, companies, I think, especially when it came to, you know, Claude Code, in Claude Co work of late twenty twenty five and early twenty twenty six, that was kind of the ushering in of the new era of these agents that can run all the time. So, you know, you saw people sharing, oh, I had my agent run for thirty minutes, for sixty minutes, for five hours. And then there was this rush for enterprises too. Well, let's do that. Because clearly, that has to mean something.

Jordan Wilson [00:05:46]:
And it was the assumption that if you use a lot of tokens or, you know, if your agents were constantly running around the clock, that means that the company was making money. But, obviously, those two things do not correlate. And I think some of the big companies set a bad example. One of the most prominent examples, well, we saw some, some reporting that Meta reportedly built an internal leaderboard ranking employees by tokens used. So one engineer reportedly ran 281,000,000,000 tokens in a single month. And some employees, well, they wanted to climb this internal leading board, leader board because, you know, not just that meta, but a lot of other companies that were doing something similar. You know, it if you weren't using a lot of tokens, developers and, well, you know, people that were being pushed to use AI models felt if they weren't using enough tokens, then maybe their job was at risk. So you had stories come up that employees were just running empty agentic loops.

Jordan Wilson [00:06:52]:
Right? Having it, you know, intentionally trip up, you you know, do some of their personal errands, but giving it impossible pass just to make sure that they were using, those tokens. So they weren't actually creating new business value for their company. They were just trying to light tokens on fire. Alright. So now you know the trend. Let's double click into the basics. Right? This is the start here series. This is for both beginners and for, you know, experts trying to double down.

Jordan Wilson [00:07:28]:
And depending on when you're listening to this, right, as of the recording, it's June 2026. So if you're listening to this in December, maybe the tokenization story has changed. But in simple terms, here's what a token is. It's essentially a small piece of text that an AI model reads or writes. So, essentially, if you, you know, put a big prompt to JGBT or Gemini or Claude or Copilot, and you insert some of your company's documents. Right? These models don't technically read it. They convert everything into tokens. So they think in tokens.

Jordan Wilson [00:08:00]:
They technically read in tokens, and they output or talk back in tokens. So one token is about three quarters of an English word. Right? So, you could say a a million tokens is about 750,000 words, give or take. Right? And models charge a different way. We're gonna talk about some of these model subsidies and how they're kind of going away. Right? But your usage or what you're actually paying for is ultimately converted into tokens. So now that you know that that is kind of the currency of a token, let's break down how these models use tokens. Right? And that's obviously changed because, you know, two or so years ago, it was a very straightforward transaction.

Jordan Wilson [00:08:43]:
Right? It was literally you could measure it in those words. Right? Hey. I, you know, I pasted a 100 words to, you know, chat GPT, 70 five tokens. It gave me a 100 words back, another 75 tokens. Now it's not as easy because the models are getting more and more cable. So let me just give you a, an example. Right? So, Codex and OpenAI recently ran a promotion. I'm on the $200 a month plan.

Jordan Wilson [00:09:15]:
This part's kind of important. And they had a couple incentives going where, you know, essentially, they gave some multiples to some people. And I was one of those people. So I was able to essentially go through maybe about a billion tokens a day, right, on some days where I was running certain processes a little more heavily. Not even a ton of coding projects. I've just been automating a lot of my manual knowledge work, which has been great. Right? So think of a billion tokens, and it depends on your input output and how much reasoning you're using, you you know, how agentic your task is. But think of about a billion tokens, somewhere from, like, 12,000 to maybe $17,000 if you were paying for those straight via the API.

Jordan Wilson [00:09:59]:
Alright? And this part's important because early on in the early days of AI, and we're getting out of those nice days. Essentially, companies would subsidize those tokens via a subscription plan. So I'm wasn't paying 12 to $17,000 a day for these tokens, obviously. Right? We don't have, you know, 15 ads on our daily podcast like other podcast do. Right? We're not making $17,000 a day in that. So a lot of companies have been subsidizing those tokens and, you know, as a way to attract and retain customers, you know. It's kind of the dance, especially, I think OpenAI, Google, and Anthropic have been taking different stances on it, especially as of recently. Because many companies, specifically Anthropic, they've been kind of the strictest on those usage limits.

Jordan Wilson [00:10:48]:
Google, after their IO conference on their $20 a month plan, they started to, implement some actual usage limits whereas before they really didn't. And Microsoft, even with their Microsoft GitHub Copilot, they're starting to do that too. So before, on certain paid plans, you really had some wide open usage. Now they're really starting to meter it. Right? So, you know, before, it's kinda like, well, taking a Lyft. Right? It's almost like you had a a monthly subscription to Lyft, and you can just take it from any point. Now some of the companies are more taking it to be the metered taxi approach. Right? OpenAI, I think, is still probably the most generous in subsidizing those subscriptions, but eventually, a lot of those subscriptions or those those subsidies are going to change because someone's paying for it, and that's essentially a a customer acquisition or customer retention costs that companies are ultimately paying to keep their millions or, you know, eventually billions of customers happy.

Jordan Wilson [00:11:49]:
But many tier so that's a personal plan. Right? Personal plans work a little bit different than, you know, big business or enterprise plans because many enterprise plans, you know, you're you may get a certain allotment of tokens, per seat. But once an individual surpasses that, then they're charged at the normal taxi rate. Right? So it's like, alright. Well, you get, you know, 10 free Lyft rides a month or 10 free Waymo rides a month. But then after that, not only are you paying the metered taxi rate, but it the meter is gonna be running all the time because if you're running agents, those things are running around the clock. So let's talk about how these tokens are fully spent. I wanna break this down in, hopefully, a very, easy and simple to understand way.

Jordan Wilson [00:12:35]:
So there's depending on how technical we wanna get, I'm gonna keep it easy. There's essentially four different types of tokens. Right? Because everyone's like, well, how do I know how I'm using tokens? Well, basically, anytime you or the model is working, tokens are essentially being, counted up. Right? So you don't normally have to worry about it, from the cost side. So if you are paying strictly via API, so if your company is, you know, working just on the back end via the API, you're paying strictly API costs. If you're on a personal subscription, for the most part, you're just, you know, you just have your allotment and your usage. And if you're on a team plan, it might be a hybrid of those things. But every input that you send, every message is converted to tokens.

Jordan Wilson [00:13:21]:
Right? So longer prompts and larger files obviously raise those inputs, those input costs. And even how different models tokenize or understand things really vary from model to model, or sorry, not just model provider to model provider, but also the actual model to model within the same provider. You know, some, you know, models cache things a little differently. Right? So, like, a browser cache that remembers things so it doesn't have to reload that web page every time, every time. Sometimes, you know, depending on how you have things set up, how you have your connectors set up, if you have a very long chat thread going, it might rerun those tokens every single time. Right? Like, if you have a, browser that crashes and you had, you know, 50 windows open and you reload it, it might load every single one of those 50 windows and it could be very slow. So depending on how a model caches things, it might actually impact how many tokens are being used. And that's important because the token usage also impacts the context window.

Jordan Wilson [00:14:25]:
So that's essentially how much information a model can remember fairly accurately before it starts to lose, information at the top of the context. Okay? So that's, that's driven by tokens. Right? So the most common way and the easiest way to explain is your input tokens. The second type is output. So that is what a large language model gives back to you. Sometimes that's text. Sometimes it's code. Right? Sometimes in theory, it's a video.

Jordan Wilson [00:14:56]:
Right? Sometimes it's an image. So, you know, in those cases, it's a little bit harder, to to count and understand the token usage there. But AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead. But you don't have ten hours a day to understand it all. That's what I do for you. But after 700 plus episodes of Everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward.

Jordan Wilson [00:15:49]:
In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode six ninety one. Number two, tap the link in your show notes at any time for the start here series, or you can just go to starthereseries.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same. The start here series will slow down the pace of AI so you can get ahead. This is where it starts to get expensive. Right? And I remember being at the, NVIDIA GTC events, might have been last year. Yeah.

Jordan Wilson [00:16:32]:
I think, you know, NVIDIA CEO Jensen Huang kinda had a graph up there and saying, yes. The price of tokens is going down, but the amount of tokens used is going up. So, yes, we've seen roughly a 10 x decrease in the cost of tokens. So you may say, oh, well, what's the harm with token maxing then? Well, that's because of token type three that is reasoning. And because of these models that now reason by default, whereas before, they were kinda just like, you know, super smart next token predictors. Right? It was text in, text out. Now not so much. Obviously, the type of inputs can cost more to tokenize the type of outputs.

Jordan Wilson [00:17:16]:
Right? If it builds audio, video, an interactive website. Right? But now these models think by default, and that's where we sometimes see what Jensen Wong and others have talked about, a 100 x or more, kind of output that an agentic model requires versus a, you know, traditional non reasoning, non thinking transformer model. And these new models reason internally. Sometimes you can see these models reason by looking at their chain of thought, and you can observe and see all these different tools it calls and things like that. Sometimes it doesn't, but these models actually think. And even though you might not be able to count, in theory, the words of those thinking, it is eating those up. Those tokens are being consumed. And then last but not least, I kinda just alluded to this because you can sometimes trace this, in the chain of thought, and that is the tool use.

Jordan Wilson [00:18:12]:
Right? That's the harness. That's, you know, as an example, you know, if you're running something like codex or quad code, you know, it's not just the model that thinks. It's thinking and it's using tools. You know, if you give it special access to skills, sometimes those skills will run multiple apps. They'll run, you know, different plugins, you know, calling different tools, searching things on the web, using computer vision, using computer use. All of these different tools are obviously very token intensive. That's because AI agents can use tools, they can read files, they can run code, and then they can repeat steps automatically. And then when we talk about scheduling agents, right, and you talk about running agents on an automation, so many times, business leaders aren't even always looking at these things.

Jordan Wilson [00:19:05]:
Right? As an example, if you have a scheduled agent, you know, and it's pulling from your dynamic data in one of its jobs, and I do this all the time, but luckily, I'm on, you know, codex. I love it. I'm on the $200 a month, plan. So I don't usually have to really worry about my limits. Right? But a lot of times people, they'll set up powerful agentic loops and it's working with your dynamic data. So maybe you have something every single day that, you know, triages your company's, you know, SharePoint, OneDrive, maybe there's a team inbox, and then it creates a live dashboard. Right? And it updates every day, or maybe it updates every single hour. Right? That's an agent going in there and ingesting a lot of input tokens.

Jordan Wilson [00:19:51]:
It's, exerting, I guess, a lot of output tokens. It is using a lot of thinking tokens, and it is calling a lot of tools. So it is doing all of those things on a loop schedule and let alone being efficient, which we're gonna get to here in a minute. Right? Because especially as humans get more hands off, right, which I think is the wrong approach. As models get more autonomous and more powerful, the expert driven loops should be expert driven, not the lazy human in the loop that I absolutely despise. Right? But, well, humans take their hands off. So a lot of these processes are happening on repeats in a loop, and human just look at the output, and they don't necessarily say, oh, well, maybe this it cost us a billion tokens a day. Right? And cost like that can really add up.

Jordan Wilson [00:20:45]:
So that was kind of the story, the token maxing story, because it became this bragging rights. You know, you saw these leaderboards, people would post, you know, oh, I used, you know, this many tokens this week. And a lot of times, it was the big labs. Right? The people actually building these things kind of bagging or showing how many tokens they're using. And well, a lot of times, you know, your fortune 500 type companies are looking at what big tech is doing and assuming that they just have to follow suit. And that doesn't always turn out that well. There's actually an Axios report, that just came out talking about an unnamed company. You know, there's a lot of, thought out there or talk out there on if this story is actually true or not, but Axios is generally a pretty good media outlet, better than a lot of the others.

Jordan Wilson [00:21:37]:
And they said that a company reportedly wasn't monitoring or setting spending limits on Claude, and they reportedly used $500,000,000 in AI usage. One company, one month, by not properly monitoring, this specific model is very token hung. Right? Claude is like me at 4PM if I haven't eaten. I'm angry and I wanna eat everything. Alright? So here's kind of what I want to talk about. Shifting toward token efficiency. So don't get me wrong. I am not anti using a lot of tokens.

Jordan Wilson [00:22:21]:
I'm not. Right? Because I think that there is obviously a certain amount of experimentation in token burn that needs to be done. Right? A good example, if you're in marketing or advertising, let's say you're starting a new campaign on Google Ads or something like that, or, you know, an advertising campaign on Meta. A lot of times, you have to just burn a lot of money, right, to, you know, kind of, train the algorithm on what you want, what you don't want, you know, types of conversions for your company to have enough data to know what works and what doesn't. And, technically, using using tokens is the same. Right? Because you don't know as models capabilities are constantly changing, you know, where your company or your department is going to be able to derive the most value unless you are experimenting, unless you're, well, maybe burning through a lot of tokens. But if your strategy if your company's strategy is token maxing, right, it's not gonna work. Right? You have to start looking at a certain point at token efficiency because even if you've been able to afford it, right, even if your company you know, everyone's on a a highly subsidized plan and you're not having to pay a lot of API overages, the thought is all these companies eventually, the free ride is gonna be a little less free.

Jordan Wilson [00:23:48]:
And as these models become more capable, they're going to cost more to run because they're gonna think more. They're gonna become more and more powerful. Right. So token efficiency means getting the most useful work per token spent. Right? My thought is I like to think that there's a certain price that you would pay for a certain output. Let's say you're gonna hire a consultancy. Right? Maybe you have to think, okay. Are we gonna hire a a big four consultancy? Are we gonna hire a small team of full time employees? Or are we maybe gonna hire an agency or a group of contractors for a certain project? Right? That's a normal decision that a lot of companies have to go for.

Jordan Wilson [00:24:35]:
And they probably say, this is the output. The output is x. We have three different routes that we can go. But no matter what, we have to monitor our cost and, you know, essentially, that output is a certain piece of intelligence or an artifact that our company needs. You need to start looking at your AI usage in the same way, Not just, alright, let's give everyone clawed desktop. Let's give everyone Odex and go to town. Don't get me wrong. Your company should be doing those things, but doing it in a responsible way, because the future of work, I think, is in these extremely powerful desktop harnesses that can be always on and take advantage of your company's dynamic data.

Jordan Wilson [00:25:18]:
But you have to understand that it's not just about running through. It's about what those tokens are actually doing, measuring what they're creating, knowing, and being able to equate what those tokens create in new economic value for your company versus how much it costs a human to do the same thing without any AI. Right? What did that artifact what did that output what did that client deliverable? What did that RFP cost in 2020? Right? And then being able to put a price on what a human would cost pre AI, you know, knowing how your, whatever AI system you're using, how much it costs from a token or a usage perspective to do the same thing. You have to start looking at it in a very pragmatic and measured out way. We've actually had a couple of episodes, including one in the start here series, specifically talking about measuring ROI. But that is the accepted output. So the measure is the accepted output against the total cost, time, and rework. Right? So you have to measure it, what it costs for a human to do it, no AI, what it costs for an augmented, you know, kind of a best human with the best non agentic AI, and then you need to measure it with the agentic AI with the expert driven loop.

Jordan Wilson [00:26:38]:
So you have to know there's gonna be a different mix, but you have to know the right model, use the right way, just beats simply using more AI every single time. So couple of things. I'm not gonna say use this model for this, but I think there's been a lot of, I don't know if it's disinformation necessarily, but a lot of misinformation floating out there around what model is the best to use for what reasons. Right? And I think we've seen a shift, and I'm probably gonna do a start here series episode on this very soon about the shift away from models to the harness. Because I don't think models are as important as they were anymore. Because especially these, you know, point upgrades, you know, going from, you know, GPT five three to five four and Opus four seven to Opus four eight. Are the models better? Yes. Are the benchmarks better? Sure.

Jordan Wilson [00:27:34]:
Right. But, ultimately, it's the harness, I think, that makes the biggest difference. Anyways, you still have to look at the price per intelligence. Right? Are those tokens worth it? Or if you give a a model a certain task, is it efficient? Is it just gonna run-in loops all day? Or right? If you give a model, a task and it gets it done in one pass without misfiring 20 times, without having to, you you know, unnecessarily call seven tools that it didn't really need because it was rushing to the conclusion instead of breaking it down in a proper way. Is that the right model? And I think there's a couple important metrics to look at. We're gonna tackle two here. So if you listen to the show at all, you know, a lot of times they talk about the artificial analysis. So there's the artificial analysis intelligence score.

Jordan Wilson [00:28:26]:
So that's essentially a series of tests, and then models get a score. And one of the most important, you know, metrics to look at undoubtedly is just that score. So that's ultimately how intelligent is a model. Nothing else considered, which is a good starting point. Right? Especially if you're on a highly subsidized plan that you don't have to worry about anything else. But that's not the reality. And, well, that, you know, fantasy AI setup that we all had in mid twenty twenty five through, maybe, like, five months ago is gonna slowly start to go away. So you have to now start looking at what does that model actually produce, per token.

Jordan Wilson [00:29:11]:
Right? So that's why the artificial analysis, that same exact test that I talked about, right, where it's essentially a conglomerate in artificial analysis. The company runs these, you know, these, you know, roughly 10 or 12 different tests, and they count how much money because they're running this via the API. So they know not just the score, the intelligence, because that's important, but how much money did they have to spend to get that level of intelligence. So artificial analysis also has this, in intelligence versus cost. And so one thing you'll see, the new Opus four eight, did beat out the new the, OpenAI 5.5. Not by a lot, but by a solid point. One point is not a big jump in the artificial analysis intelligence index, but it's, you know, big enough to say yes. When it comes to just strictly intelligence and intelligence only, Claude Opus 4.8 is the best across all of these benchmarks, which is an important thing to talk about.

Jordan Wilson [00:30:11]:
But the cost, it is absolutely bonkers because the anthropic hope the anthropic models are and you can see this by test. And people always this is annoying y'all. Like, people always accuse me of being, oh, anti anthropic. You you love Google. You love OpenAI. No. Cloud is a great model. Right? But it is incredibly token inefficient.

Jordan Wilson [00:30:36]:
And if you are on a a paid plan, anthropic's plans are the least subsidized. Right? That's maybe why they're the most most profitable, AI company per user right now, presumably. But in terms of if you are out there, you know, on a subsidized plan, not only is it you're gonna get the least for your dollar, but if you're paying on it through the API, all you're doing, you, your company is saying, let's just go ahead and flush this money down the toilet because it doesn't matter. Because the facts say, the stats say, the unbiased third party say that you are literally lighting your money on fire if you are choosing to use the anthropic models via the AVI. Right? Yes. Maybe there's a capability that the hardest, you know, cloud desktop has that no one else has. Sure. Right? But if you have built your company's AI strategy in a modular way, which is the right way to do it, right, and you are still using Infrafx models, you are essentially saying by all measures that you just wanna waste money.

Jordan Wilson [00:31:40]:
Right? Here's another, example also from artificial analysis. This is the total cost. Right? The total cost to get to this. So in tropics models are the most expensive. So to run that test, about $5,100 for Opus 4.7. Comparatively, OpenAI's latest model, $3,300. So I'm not the best at math, but that is 50% more expensive to get that same level of intelligence. Right? That's like if you're saying, okay.

Jordan Wilson [00:32:13]:
I'm am I gonna take an Uber or a Lyft? Well, the Lyft is 50% more expensive, and they're gonna get you there in roughly the same time. Well, not actually because there's also a speed factor, but I don't wanna focus on speed today. But from a cost per intelligence ratio. Right? And this is where when we talk about token efficiency, this is why you have to keep up. This is why you have to build in a smart way. You don't want to lock your company into a certain vendor if you don't have to. Right? Obviously, if you're using something on the desktop side, but if you were building on the API side, or, you know, maybe you're using Cursor, right, where you can use a number of models. You're using, you know, Copilot that can use the, GPT models.

Jordan Wilson [00:32:59]:
You can use the, clone models. You can use the, you know, the Microsoft models. Right? You have to know these things. You have to understand how much does the intelligence cost, not only which model is the most intelligent. Another great, metric to look at, and I'll probably be talking about this a little bit more, is DeepSuite. So DeepSuit is a new, and I'd say it's probably one of the better, agentic, coding benchmarks. So there's an older benchmark called SweeBench that I think is actually antiquated. And DeepSue is actually a really good benchmark.

Jordan Wilson [00:33:38]:
Maybe I'll do an entire show on that. Let me know in the comments if you want me to do it. Just leave DeepSue, so that's deep s w e, if you care about this. Right? But this is a great benchmark, that shows something similar to the artificial analysis. But the reality is, at least when it comes to token efficiency. Right? I'm not gonna tell you use this model for this. Right? But in the same way that the artificial, analysis actually, Gemini three one Pro, one of the best models out there in terms of, you you know, getting for what you pay for because it costs even a fourth, just about a fourth of what GBD 5.5 costs. Right? And it costs a fraction, you you know, not even, 20%.

Jordan Wilson [00:34:30]:
You know? Yeah. It's it's, let me let me do my math here. So that is about more than five x, more to use Opus 4.7. You're gonna pay five x more, not 50%, five x more to get the same level of intelligence out of Opus force of Inverse's Gemini three one pro. So if we get back to deep suite, the same thing shows true. Right? GPT 5.5 gets a not only a higher score, much higher than the Opus models, but it does so at, again, a fraction of the cost. Right? So, it's more than, double. Right? So you're getting not only a higher level of intelligence, but you're about half the cost.

Jordan Wilson [00:35:20]:
So when it comes to token efficiency, it's not just about what your company gives you access to. Because like I said, you always need to be building modularly. Because what happens if a model you know, if you've, set up a big chunk of your company, whether it's running somewhat autonomously, you know, a bunch of, you know, cron AI jobs, whatever it may be. If GBT five six breaks your flow, right, if if, you know, Mythos comes out and nothing works, you have to have a fallback. Right? So you should always be testing these things, but not just what is the most intelligent model. That's a good place to start. But the cost per intelligence is an extremely important thing, and that is the prudent and the responsible and I think will be the rest of the story. The telling takeaway through the rest of 2026 is setting your setting your company up to not only be AI native, but to be token efficient because the AI subsidy days are gonna be over fairly soon and you have to be ready.

Jordan Wilson [00:36:25]:
So your company needs to measure the accepted outputs and the AI cost, not just the number of tokens used. You have to match the model to the job. You can use smaller, less expensive models for certain things, parsing PDFs, you know, doing simple text summarization, and you can use those bigger, more expensive models like, you know, Opus and g p t five five for those more complex tasks. You need to start monitoring, you you know, these agents that are stuck in a loop. You need to start maybe kicking out certain models that just constantly are calling tools that they maybe don't need to. Right? You have to be paying attention to the chain of thought. And then you need to review your AI cost because these long running agents, it is so easy. It's so easy as these models become autonomous like you can in a in a Waymo or, you know, in that version of Tesla's cars that I don't know if they ship, but they keep saying.

Jordan Wilson [00:37:21]:
Right? It's almost like your hands off the wheel. Right? You can completely hands off the wheel doing something else. You can't do that right now with agentic AI. You can't just be a passive human in the loop because that is gonna cost your company probably millions of dollars in the long run. You have to have an expert driven loop that is constantly monitoring not just the inputs and the outputs, but monitoring the efficiency of those tokens as well. Start doing that now, and you are not gonna be behind for the rest of the year. You are gonna be ahead, and that is where I want you to be. You know, one quick tidbit to wrap this up.

Jordan Wilson [00:38:03]:
It was actually, you know, just got out of a a great session, last night with the mic, Microsoft CTO, Kevin Scott. And he kind of ended, you know, one of his you know, this just a little little, gathering at a bar. But, you know, really just talking about how activity is not equal value. And I think that's a great takeaway to end this on. Because just because you're using the best models and you have all of these scheduled prompts and you have all of these agents running around the clock. Don't confuse activity for value and especially pay attention to what model you are having be active for you and the value that it is actually creating. Alright. I hope this one was helpful.

Jordan Wilson [00:38:48]:
If so, let me know if you could like and subscribe to the show. If you're listening on Spotify or Apple Podcasts, I'd really appreciate that. And then make sure you please go to starthereseries.com to get exclusive access to our inner circle community so you can start here series along with everyone else that's doing it. So thank you for tuning in. I hope to see you back tomorrow and everyday for more everyday AI. Thanks y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI