Ep 725: Measuring AI ROI: Why you’re doing it wrong and the 7 Steps to fix it (Start Here Series Vol 11)

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


Measuring AI ROI: Why “Vibes” Aren’t Enough and the 7 Steps Every Company Misses

The debate around AI return on investment (ROI) is filled with noise, outdated measurement playbooks, and misleading studies. Yet, leading enterprises quietly pocket substantial returns while others argue over unproven metrics. Cutting through the clutter, there are concrete, measurable steps companies can apply to rigorously quantify—and meaningfully increase—their AI value, right now.

Why Most AI ROI Measurement Fails

Businesses typically approach AI ROI using legacy digital transformation frameworks—models that simply do not hold up against today’s AI capabilities. Many track output “feels faster” or “seems better,” but lack hard receipts, instead relying on pre-AI benchmarks and vanity metrics (like tool utilization rates) that obscure true performance. Most studies even show that 95% of enterprises have adopted AI in some form, eliminating the superficial advantage early AI adopters once enjoyed.

The real disconnect is a lack of meticulous measurement tracing true before and after impact, especially across quality, cost, reliability, and risk. Without this rigor, workers are often “pocketing” AI-driven productivity for themselves, especially in remote and hybrid models, causing invisible gains that do not translate into company-level performance.

The Data-Driven Evidence for AI ROI

Despite viral claims that “95% of AI pilots return zero ROI” (based on a widely-cited but deeply flawed MIT study using just 52 qualitative interviews and a six-month P&L window), the most robust quantitative studies tell a different story:

  • IDC (International Data Corporation): $3.70 return for every $1 invested in AI.

  • Wharton Three-Year Study: 74% of enterprises report positive ROI on generative AI.

  • Google Cloud 2025 AI Report: 74% of executives achieved ROI in the first year of deploying gen AI.

  • Deloitte AI Survey: 84% of AI investors saw measurable ROI.

When AI is deployed, quality and speed jump: leading AI models now tie or outperform expert humans in 70% of blind trials across 44 job types, completing tasks 100x faster—if humans are prepared to use them well.

The Hidden Challenge: Outdated Job Roles and Training

Nearly 89% of organizations have not updated job responsibilities or workflows to reflect AI’s impact, per Workday research. Enterprises often measure output by pre-AI standards, failing to capture the “invisible productivity” when employees quietly automate substantial portions of their job responsibilities.

True, scalable ROI requires new job architecture, focusing on outcome-driven metrics and standardized, ongoing skill training. Education isn’t optional; effective AI measurement and deployment demand robust documentation of workflows and real-time process adjustments—well before any technology rollouts.

Pinpointing ROI: The Only Formula That Matters

ROI isn’t a mystic calculation. For direct cost measurement, follow this formula:

(Fully loaded hourly rate × hours saved with AI) – AI tool costs = Net AI ROI

This formula focuses on three levers:

  • Time Saved (the largest driver)

  • Costs Reduced

  • Risk Avoided

It’s not about total message volume or dashboard “utilization.” Boards and CFOs care about cost per task, throughput, error rates—metrics that connect directly to business goals.

The 7 Steps to Measuring AI ROI—A Repeatable Blueprint

Organizations need operational rigor, not guesswork. The following seven-step framework quantifies AI’s impact and prevents common pitfalls like chasing vanity metrics, running endless pilots, or abandoning initiatives before full adoption.

1. Define Rubric and KPIs

  • Rigidly set success criteria for each workflow, before any testing begins.

2. Establish Baseline “Pre-AI” Performance

  • Time several employees performing the process manually. Track error rates, time, and cost per completed task.

3. Build 20–40 Realistic, Messy Work Cases

  • Create a range of real-world and outlier workflow challenges—not cherry-picked best cases.

4. Configure a Consistent, Production Workspace

  • Use the same AI models, platforms, and permissions across all tests for replicable results.

5. Run Each Test Three Times (With AI and Human Teams)

  • Disable persistent model memory. Require proof artifacts for validation.

6. Grade Outputs Blindly Against Standardized Criteria

  • Remove bias by grading outputs without knowing if they came from a human or AI.

7. Retest Each Month and Track a Rolling Average

  • Model and tool updates can raise or lower performance. Continuous validation ensures ongoing accuracy.

Plug the time and cost differentials into the ROI formula; savings and gains will become explicit.

Redefining the ROI Question for Business Leaders

AI ROI is not an elusive, abstract future gain—quantitative evidence is overwhelming now. The only meaningful question is, “How much is lost by not training teams and measuring impact, starting immediately?” Perpetual hesitation means competitors, and even your own workforce, accrue advantage below the surface.

Success comes from disciplined measurement, process documentation, and capacity to adjust as models and workflows evolve. The gravity of AI is inescapable: treating its performance with the same rigor as financial metrics is the new baseline for business leaders seeking sustainable, large-scale value.

For companies ready to tap measurable outcomes, moving beyond “AI vibes” requires nothing more (and nothing less) than operational discipline and refined measurement—starting today.


Topics Covered in This Episode:

  1. Proving and Measuring AI ROI in Companies
  2. OpenAI’s GDP Benchmark: AI vs Experts
  3. Debunking MIT’s Viral Zero ROI Study
  4. Quantitative AI ROI Data from Business Studies
  5. Invisible Productivity: AI Savings Pocketed by Workers
  6. Limitations of Pre-AI Job Roles and Metrics
  7. ROI Calculation: Time Saved and Cost Reduction
  8. Five Main Reasons AI ROI Isn’t Measured
  9. Seven-Step AI ROI Measurement Blueprint
  10. Importance of Ongoing AI Model Retesting
  11. AI ROI: Training, Education, and Implementation Gaps


Episode Transcript 


Jordan Wilson [00:00:17]:
Can you actually prove AI is working at your company? Sure. Everything feels faster and hopefully outputs have improved and maybe more work is getting done. But is there ROI and don't just give me vibes, give me receipts. What's the actual return on your investment for AI at your company? Chances are you can't answer that. And the reasons are actually very simple. It's because businesses are using the same digital transformation playbook they've always used when it comes to AI and, well, that playbook is useless in 2026. So there's a 99% chance you're not actually measuring ROI, but I'm gonna simplify it for you on today's show and tell you why the ROI on AI debate is absolute nonsense. And I'm gonna give you seven simple steps on how your company can fix it today.

Jordan Wilson [00:01:19]:
Alright? Let's get measuring. So, here's what you'll learn on today's show. You're gonna learn the basics of measuring Gen AI ROI and why most companies never actually do it. You're gonna learn why a piece of MIT marketing disguised as a study misled markets and what five larger studies actually show. You're gonna learn how to use a four metric scorecard covering quality, cost, reliability, and risk, and how to run a seven step evaluation sprint that actually solves your ROI. So what the heck is this? Well, this is everyday AI, number one, but this is our start here series. I didn't have a good answer. When people first subscribe to the podcast and they're like, Jordan, there's 700 episodes.

Jordan Wilson [00:02:08]:
Where do I start? I I'm like, I don't know. Well, now I do know. You start here with the start here series. This is the essential podcast series to both learn the AI basics and to double down on your AI knowledge. So make sure to go to starthereseries.com. That's gonna give you free access to our inner circle community, and then you can go in the start here series space and listen to now all 11 episodes of this series. So last episode, which was a banger FYI, we talked about from AI chatbots to autonomous workers and how consumer AI has changed and what's next. So make sure you go, check that out.

Jordan Wilson [00:02:49]:
It's episode seven twenty three, but, volume 10 of the start here series. And today we're gonna be measuring ROI and tell you why your company is doing it wrong and these seven steps to fix it. Let's start here. Talked about this once or twice, but let's talk about this benchmark from OpenAI. It's called GDP BOW. So this benchmark tests AI on real world deliverables across 44 jobs in nine different GDP sectors. Essentially, this is a unbiased benchmark that pits AI models against expert humans, experts in their fields. Right? They each have the same task, which is generally, to produce something of value.

Jordan Wilson [00:03:35]:
Right? That's why it's called GDP Val, to produce something of economic value. And then there's expert judges, expert human judges, and they don't know which output came from the AI. They don't know which output came from the expert human. And right now, top AI models either tie or win 70% of the time in these blind evaluations against professionals that have an average of fourteen years of experience. I mean, my gosh, we should just end the episode now. Right? There's your ROI on AI. Right? Done in under four minutes. There's a lot of issues here.

Jordan Wilson [00:04:08]:
Number one, even what I just said. Right? Do you know how to use today's models to, you know, one shot personalized synthesize information and create, you know, spreadsheets, PowerPoint decks, and use your own company's information without you having to do anything and verify outputs? Well, probably not because companies aren't educating their people, and it's hard to keep up with, today's capabilities of AI models. But what this did show us is that the whole debate on ROI is absolute rubbish. Yes. There we go. For, listeners across the pond, look at me. I'm, multilingual here. It it's garbage.

Jordan Wilson [00:04:50]:
This whole discussion on does artificial intelligence give you a return on your investment? It's it's honestly a very dumb question if I'm being honest. Right? Go go read the GDP valve study. I keep saying, oh, I'm gonna do a show soon. I I will actually do a show soon, because it's worth it. So not only did it show that, well, today's AI models are exponentially better than expert humans when judged by other expert humans. Right? But it said that in many tasks, they do it a 100 times faster. So, again, I am not a mathematician, but if the AI model is the same or better 70% of the time and it's a 100 times faster, that's the math. So why are we still even debating this concept of is there return on investment? Well, that's because of an infamous piece of marketing from MIT.

Jordan Wilson [00:05:47]:
So, and yes, I'm calling it a piece of marketing or study in very heavy quotes. Right? Like the, what is that? The the Austin Powers, you know, a billion doll yeah. Huge quotes on the word study. Right. So MIT claimed, in their wildly viral study that ninety five percent of enterprise AI pilots delivered zero ROI, and this was in August 2025. Right? And, I mean, the market's moved. Right? NVIDIA stock tanked three and a half points. Palantir went down, nearly 9%.

Jordan Wilson [00:06:23]:
There's actually billions of dollars of losses across the, not just the tech sector, just the business sector because, right, AI has essentially propped up an otherwise struggling, US and at times global economy. And then, you know, this study came out and said, oh, AI's you know, there's no ROI. Well, most people right? As a former journalist, I understand this. Right? There is one study that came out, and then essentially, every other, journalist hit the copy and paste, you know, more or less, hey. That headline sells, you know, 95% of AI pilots don't, produce any ROI. Right? That's great. Let's go ahead and run that. But I don't think any journalist actually read the story or very few actually did.

Jordan Wilson [00:07:07]:
Right? That's because this quote, unquote study was based on 52 qualitative interviews. There's no quantitative piece to that 95, and it was, directionally. Right? This was a directional study. So it's more or less this was a vibe. This was a vibe study. I'm gonna talk to 52 people, you know, nothing quantitative. I'm just gonna give my vibe on this. But the stat came from 52 qualitative interviews that measured p and l impact in under six months, which there's no such thing.

Jordan Wilson [00:07:36]:
Right? When you talk about digital transformation, go back and look at the Internet. Right? You could have shown a p and l impact in under six months on using the Internet. Right? Because if you could have, you could apply that exact same methodology to AI. Anyways, well, why this is a piece of marketing and not actually just a bad study. I would call it a bad study, but it wasn't. It was marketing because in the end, they were selling their own product. They said, hey. Right.

Jordan Wilson [00:08:03]:
All these AI pilots fail because you're not using our agentic AI product, and they were, you know, selling access to it. Yeah. I'll stop there. If you really want the the down low on that one, make sure to go listen to episode five ninety seven where I, like a human, read it, like a human with a brain, read the study multiple times, and it's, you know, laughable. Anyways, if you look at the real data, right, companies that talk to thousands and got quantitative data, obviously, the ROI for AI is there. Right, the IDC, the International Data Corporation, found a $3.70 return for every dollar invested in AI. Wharton, a three year study found that 74% of enterprises report positive ROI on GenAI. Google Cloud's 2025 annual AI study found 74% of executives reported achieving ROI just within the first year of generative AI deployment.

Jordan Wilson [00:09:03]:
Also, Deloitte's 2025 AI survey found 84% of those investing in AI said that they were gaining ROI. This I think that this this debate is still going on because humans are innately lazy. Right. That's the reason why we don't want to actually sit in measure. Right. Again, my my background, I have a little bit of marketing and advertising background. So I understand the importance of measuring something before and after to see if there's an impact. Right? Running millions of dollars of ads for clients over the years, there's this thing called return on ad spend.

Jordan Wilson [00:09:47]:
Right? You see how much, you you know, how much money was spent on ads. You have to be able to attribute and track the revenue that it brought in, and then you do some simple math. Right? Return on ad spend. Well, there is also a simple equation that I'm gonna give you guys here in a little bit, the same thing, return on AI. So here's the dirty little secret because you're probably thinking, okay, Jordan. Well, if every company is using AI, well, why isn't everyone just seeing exponential growth? Well, I think you would see that a little bit more if only 5% of, businesses were using AI. Right? But most studies say more than 95 or 97% of enterprises. So, the the playing field has just gone up.

Jordan Wilson [00:10:39]:
Right? The minimum entry is no longer like, oh, you're using AI. That's great. That's such a differentiator. No. No. It's you're using AI. Okay. I've been saying since 2023.

Jordan Wilson [00:10:51]:
It's the same thing as, oh, our company is using the Internet. So, right? But here's what's happening. Workers are pocketing it. That's where the ROI is going. Right? Because true productivity and true ROI, well, it requires results driven metrics, not time based management. Most companies haven't, you know, gotten this figured out yet. And I think a big portion of this is that remote and hybrid workers are using AI, whether it's approved AI or, you know, whether it's, you know, AI sprawl, you know, dark AI, whatever you wanna call it, and they're just pocketing it. It's this invisible productivity.

Jordan Wilson [00:11:34]:
Right? I've literally talked to countless people. I won't say hundreds because I don't know if it's actual 200 or more, but I've easily talked to more than a 100 people over the past few years in very successful. Right? Even one of my, you know, one of my good friends, told me this, like, maybe one or two years ago. Pretty pretty high up, at at a company that was, public company when he worked there. Said he automated about 95% of his job with AI. Right? No one knew because he's at home. I'm like, what are you doing with all this time? He's like, you know, chilling, golfing a lot. Right.

Jordan Wilson [00:12:13]:
That's the reality. You know, so many companies, so many employees, I think are just pocketing this time savings. Right? As as we've gone to this, you know, remote hybrid workforce, and I think that's one of the reasons why last year, so many companies had this, you know, big return to work push. And right now, according to Workday, 89% of organizations haven't even updated their roles to reflect AI. Right? So you're still using, you you know, twenty twenty six tools, inside of job structures made ten years ago. Right? So why does that matter? Well, again, hiring in roles in today's enterprise are still pre AI. So what that generally means, well, expectations or output is measured in a pre AI way, which isn't necessarily the right thing. Right? And in the same way that I've encouraged, you all for for years to not upskill or reskill, that's that's a waste of time when it comes to AI.

Jordan Wilson [00:13:22]:
Right? You need to unlearn and rebuild been saying that for a long time. Departments and companies need you to do the exact same thing. Right? You can't just, oh, let's add a little bit of AI to this job role or, you know, hey, let's let's try to think a little more, AI natively here at this company. No. You you gotta tear it apart. In the same way that job roles haven't changed, right? That means outputs aren't going to change either. But if you're using AI, if you've trained your people on AI, if they have the right AI, they're just banking that productivity. So here's how finding ROI is done.

Jordan Wilson [00:14:02]:
It's simple. ROI can mean a lot of different things. It can a lot of people just think, oh, that means you're bringing in more revenue. No. I think mainly it's time saved. Right? But I think ROI also means just cost reduced, revenue increased, or risk avoided. It's not just prompts sent. Right? I think that's what a lot of people they're looking at, the number of hours, you you know, your utilization rate if you're using, you know, an enterprise tool like Chegg GPT.

Jordan Wilson [00:14:29]:
Right? Everyone's like, oh, you know, we need to increase our utilization from, you know, 8% to 20%. Right? And then companies, you know, hire us to, you know, train hundreds or thousands of employees. And it's like, okay. Utilization is not, the metric that pushes it. It's are you saving time? Are you, increasing revenue and avoiding risk? The formula, well, simple math. It's the time saved or, increased revenue, but let's look at time saved. So it's this time saved multiplied by fully loaded hourly rates minus any any monthly AI subscription costs. Simple formula.

Jordan Wilson [00:15:08]:
Right? And right now, boards and CFOs just care about cost per task, cost per task, throughput in error rates, right? Not usage dashboards, not utilization, but we're seeing the same reasons for failure still years later. Right. Through my own conversations and conversations with others, I can say when it comes to measuring or not measuring ROI, I still think that there's five patterns or five main reasons why either individuals or departments aren't able to do this. So number one, not having a pre AI baseline. Number two, you you know, doing these slow year long pilots. That's a recipe for disaster. Number three, celebrating one lucky run and making that your AI strategy. Number four, when vanity metrics become the focus.

Jordan Wilson [00:16:06]:
And number five, shiny object syndrome pulling focus, away before real company wide implementation can actually start. Right? Now you get a little you get a little ground, you show ROI, and then you're like, oh, what about this model? What about this? What about this? What about this? Right? Before you actually, you know, implement it company wide. And these aren't technology problems. Right? This isn't standard digital transformation. These are we as working human beings as knowledge workers, we don't have a playbook to follow. Right. So all we've been doing is following the same playbook that we've always done. Right.

Jordan Wilson [00:16:49]:
Those slow pilots. You know, you get one thing right, and you're like, oh, this is what we do now. It's the wrong way to do this. The other big reason, well, training. Right? The training actually leads to those five other failures, so it's more of a of an underlying, foundational issue. Right? But right now, 49% of enterprise leaders cite recurring, sorry, cite recruiting AI talent as their single biggest challenge. Alright. AI, you need documented workflow steps and company data.

Jordan Wilson [00:17:22]:
And that process documentation is part of the ROI investment. Right? Education, training, ongoing learning, it is required. Right? So many I think when, when AI, when an AI strategy or actual AI implementation looks that looks like, you know, pushing top down, you know, either c suite or board pushing, you know, a certain tool top down and not training or educating, that's why it's so easy to fail. Right? I use AI all day, every day. And I'll, I'll, I'll tell you this. If I take two weeks off, I'm gonna fail. That's the, that's the pace, right? People are always like, okay, well, Jordan, you clearly don't know what you're saying here. If AI is a 100 times faster, well, shouldn't we just cut all these jobs? No.

Jordan Wilson [00:18:21]:
I think jobs are gonna look very different. I've been on record as saying this since the very first episode. AI is going to, I think, quote unquote, take away more jobs than it will create, but it will create many roles that we don't even understand. I've always said, most enterprise organizations, they need a lot of people like me that all they do all day is they keep up with AI and they apply it to that business. Right? Very few businesses have that because they're like, okay, that's a crazy role. We're not gonna let, you know, Bill over in IT or, you know, Deborah over in marketing, just spend all day, you know, tinkering with AI and seeing how to apply it. No. Right? No.

Jordan Wilson [00:19:01]:
But, yes, that's exactly what you should be doing. And here's. Alright, this is usually something I think we only, give away to companies that, pay us a couple dollars. But let me just go ahead and give you some of the secrets here. Alright? Here's an acronym, BASE. You need a BASE. This is the baseline rule. So this is base stands for baseline assessment of standard execution pre AI.

Jordan Wilson [00:19:27]:
Alright. So before AI touches any workflow, you gotta get that base. Right? That's your time multiple employees completing that exact task without it. You need to time it. You need to record the average time, the error rate, the rework cycles, and cost per completed task as the baseline. You can't go back and collect this retroactively. Right? Because usually what happens well, if you've already, you know, sprinkled AI in the process and you've been doing it for a a year, right, let's say it's a 10 step process and, hey, number two and, step two and step six, we, you know, we molded those a little bit around AI. Right? It's it's too late.

Jordan Wilson [00:20:11]:
Right? You need to do it before you implement AI. You need to measure how long it takes humans to go through and do these certain projects or do these certain tasks. Right? Think of it like an internal GDP valve. Right? You are gonna go through and have a human with no AI go through and do this task, document every single step of the process painstakingly because you need to see, okay, what other humans need to be involved in this? Are there other meetings? What about the person checking the work? What about the communication going back and forth? You need to measure it all and document it all every single step. And then what's the completion rate? What's the error rate? What does quality look like? You need to establish those baselines and then redo the entire process. Right? Blow it up. Measure it first. That's your base, right? Baseline assessment of standard execution.

Jordan Wilson [00:21:15]:
That's your base. Blow it up. Do the same thing with AI, right? Not your first iteration, iterate on it. And then you measure in that number right there. You need that because that is the base quite literally for how you're going to ultimately measure your ROI. So now let's give you that seven step guide. All right. I'm delivering.

Jordan Wilson [00:21:41]:
Here we go. Well, let me just, bring my notes up here. Cause you know, in typical live stream, unedited, unscripted fashion, I didn't put my slide up here with my, seven steps. But don't worry. I got my notes. Alright. There we go. Notes on the screen.

Jordan Wilson [00:22:02]:
So here's the seven step blueprint. Right? Step one, we already talked about it. That's the, well, actually comes a little bit before the base. So step one, you have to define you have to define what the heck it is you're doing. Right? And you have to be very rigid. You can't be flexible because you need to get, an accurate before and after. So that's defining the rubric, rubric, defining the success criteria and KPIs before you even begin testing. And then step two, like I talked about, that's the base.

Jordan Wilson [00:22:33]:
You need to measure the human baseline, and then you're gonna, you know, time multiple employees doing that and then record the averages. Step three, get messy. Right? You have to build 20 to 40 real messy work examples, including drift cases. Right? These aren't easy things for humans or for AI. Alright? And then step four, you need to configure the exact production workspace. It's the same plan, the same model, and the same, permissions. So you're not, you know, flip flopping, you you know, between different people, different accounts. No.

Jordan Wilson [00:23:11]:
Right? It has to be controlled like any experiments, and it has to be repeatable and scalable. The exact same criteria when you're taking this to a production run. Step five, you're gonna run every three times. Okay? So sorry, run each test three times with memory off and also require proof artifacts. So here's what I mean by that. We're doing this on the front end. Okay. If you list if you've listened to this podcast at all, you know I'm a big believer in bringing as many of your processes over, to front end large language models.

Jordan Wilson [00:23:49]:
We actually did a dedicated episode, on that in the start here series talking about an AI operating system. So, yeah, this is all gonna be doing things on, as an example, chat gvt.com, claw.ai, gemini.google.com. Right? But doing it with memory turned off, and doing it usually in a temporary chat, if you were, you know, AI operating system of choice, lends itself to that. Alright. So essentially there, you have steps three through six is that's your internal GDP valve. Right? You have to get the the correct use, use cases. And then from there, you grade blind. Right? So same thing.

Jordan Wilson [00:24:35]:
You have the the the series of people do it, three times. Those 20 to 40 use cases using the same AI model, and then you have humans do those same exact things as well. Multiple humans in the same way, I would suggest, you know, three different times, running the same case and then three different sets of humans. And then you grade blind, you standardize the output format, you have to agree on, you know, what's a pass, what's a fail, what's the grading scale, etcetera. But then at that point, well, you have the input and the output. Right? At that point, after step step six, when you can have your grading criteria, you know, you run through and you do the test three times. You know, those 20 to 40 use cases three times, in a large language model, memory off, temporary chat. You have your your grading rubric.

Jordan Wilson [00:25:29]:
You have your humans do it, right? The same 20 to 40 tasks, three different humans. You have it right there. You have the time. Right? You multiply, the time that it takes the humans to do it on the AI side. Alright? Minus out, you know, so take let's say it's a 100 human hours, take their hourly rate, times it, minus the, cost of whatever AI tools that you're using. Alright. There's your, augmented cost, and then you compare it to your human only cost. Right? And, obviously, the human only cost is gonna be much higher.

Jordan Wilson [00:26:07]:
Right. The same thing, if you are using any, you know, paid non AI tools, for the human only cost. Right? Maybe there's, some, you know I don't know. If you're using the Bloomberg terminal or what right. Whatever. You you know, a certain SaaS, and maybe you don't need to use that SaaS application if you're doing it AI native. The same thing. You need to look at the total human and software costs.

Jordan Wilson [00:26:28]:
And there you go. That's your return on investment. And then step seven, you need to retest this monthly after every model update. Alright? And then track a three month rolling average. Costs are gonna go up and down. Right? You think that, oh, well, they're just always gonna go down. Well, no. Sometimes, you know, the Frontier AI labs will actually roll out an update that's under the radar.

Jordan Wilson [00:26:52]:
It's not like going from a GPT five two to a five three. Right? A lot of times, there might be five, ten different versions of a GPT five two until there is a GPT five three. And, you know, whether it's OpenAI, Anthropic, Google, Microsoft, etcetera. Sometimes, you know, one of those under the hood updates actually might make things worse. Alright? So that's why you need to retest monthly or after any major model update and then keep a three month rolling average. There you go. Alright. I'm gonna go through it quickly now with no commentary.

Jordan Wilson [00:27:26]:
Step one, define the rubric and success criteria. Step two, measure the human baseline. Step three, build the 20 to 40, real messy work cases. Step four, configure the exact production workspace. Step five, run, with three times, AI models in three sets of humans. Step six is grade blindly and standardize the output format and criteria. Calculate your ROI. And then step tab step seven is retest monthly.

Jordan Wilson [00:27:57]:
So there you go. That's how you measure return on investment in AI. But let me just tell you this, and I'm gonna be, very direct when I say this. You don't need to do this. You absolutely should. Right? You absolutely should. You don't need to. This is this is gravity.

Jordan Wilson [00:28:26]:
Right? AI is gravity at this point. It is the all encompassing force and there's no denying it. Right? Real studies with quantitative data that asked thousands of business leaders all overwhelmingly show ROI in AI is real $3.70 for every dollar invested The GDP valve, right? AI models without training, right? Without iteration in a single, in a simple shot, blind test, the AI model is the same or better than the expert 70% of the time and a 100 times faster. So I'm going to leave you with this. Yes. You need to configure and figure your ROI on AI, but the new ROI question is not did AI work. It's how much did we lose by not educating and measuring sooner? That's the reality, And that is the challenge to you today. Dear listener, stop looking at ROI on AI.

Jordan Wilson [00:30:07]:
Like it's something tricky. Like it's something we can't all achieve. It's as simple as being meticulous in your measurement and that's it. And when you are meticulous in your measurement, you will undoubtedly see in insanely high return on investment of your AI. So no longer a question, mark no longer does AI work. The question is, and the way you need to change is how much are we losing by not measuring sooner and educating our teams better? Alright. That's a wrap y'all. That is volume 11 of the start here series.

Jordan Wilson [00:30:56]:
I hope this was helpful. And if it was helpful, well, number one, please, like and follow the podcast. I'd appreciate that. But then when you're done, go to starthereseries.com. That's going to give you free access to our inner circle community, and then you can go listen to the entire Start Here series all right there and connect with others who are trying to grow their company and their careers with generative AI. And, hey, while I have you, make sure if you haven't already, go listen to the 2026 AI prediction and road map series. That's episodes seven twelve and seven thirteen. So thank you for tuning in.

Jordan Wilson [00:31:32]:
Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI