Ep 628: What’s the best LLM for your team? 7 Steps to evaluate and create ROI for AI

The Detailed Blueprint for Selecting and Measuring ROI on Large Language Models

Selecting the right large language model (LLM) and justifying investment in generative AI has become a frequent challenge for organizations. Companies are no longer limited to backend API integrations – AI operating systems delivered through front-end chatbots like ChatGPT, Claude, and Gemini are swiftly becoming the backbone of contemporary workflow management. Yet, with this new terrain comes confusion, excessive options, and the absence of clear ROI.

This article distills a proven seven-step methodology drawn from in-the-trenches experience, spotlighting the precise processes successful teams use to evaluate and extract tangible returns from front-end LLMs.

From API Models to AI Operating Systems: A Shift in Evaluation

The landscape has changed: rather than exclusively deploying LLMs through backend APIs—where benchmarking is relatively straightforward—organizations have begun relocating daily business processes into the front ends of chatbots and AI “operating systems.” This evolution mirrors past organizational pivots, such as standardizing desktop operating systems to streamline workflows in the 1990s and early 2000s.

Modern LLMs in these front-end environments present not only multiple models to choose from, but also modes—features such as connectors for dynamic company data access, agentic workflows, image generation, and advanced research capabilities. These modes are now where work truly happens, offering far more than API connections can support.

Why ROI Often Eludes Even Smart Organizations

Many organizations struggle to pinpoint ROI from generative AI, despite high expectations. Based on direct consultation with hundreds of companies, several recurring pitfalls emerge:

  • Excessively Long Pilots: Year-long pilots in AI guarantee diminishing returns; rapid cycles show clearer value.

  • Poor Change Management: Even large companies often lack comprehensive change management and training during AI deployments.

  • Missing Human Baseline Measurements: Without establishing pre-gen AI time and cost baselines for tasks, organizations have no reference for comparison.

  • Ignoring Variance and Reliability: Single, successful AI runs are celebrated while failing to track performance variance.

  • Succumbing to Shiny Object Syndrome: Weekly industry announcements distract teams, preventing focused progress.

Using Public Evals to Shortlist LLMs

Several benchmarks help organizations quickly filter LLM options before beginning custom evaluation. These publicly available evaluation platforms provide detailed, use-case-specific insights:

  • LMSYS Chatbot Arena (lmsys.org): Offers blind “taste tests,” matching prompts against models and categories (e.g., creative writing, coding).

  • LiveBench: Rates models across reasoning, coding, agentic coding, mathematics, and data analysis.

  • Epoch AI: Provides benchmarking data across diverse categories.

  • Scale’s CLLM Leaderboard: Delivers a broad view on performance, reliability, and utility.

Such evals are essential to identify which models warrant further internal scrutiny.

The Seven-Step Plan for Front-End LLM Evaluation and ROI Measurement

A robust approach means organizing evaluation sprints, garnering buy-in from stakeholders (including IT and legal for data compliance), and freezing model/mode selection for focused comparison. Here’s the process, stage by stage:

1. Define Success Criteria

Create a “job description” for the evaluation. Explicitly state expected outcomes, accepted tools and features, constraints, and “do not do” actions such as excluding experimental features released mid-sprint. Develop a rubric, ideally one-to-five or one-to-ten, with clear pass/fail definitions for each use case (e.g., extracting data from PDFs into structured tables).

Select three to five KPIs per use case. These often include:

  • Human time spent

  • Accuracy

  • Volume of required revisions

  • Value created

Objective, measurable KPIs produce the clearest comparison.

2. Measure Pre-AI Human Performance

Before AI can be assessed, collect baseline data for current processes—typically neglected in failed pilots. Time multiple employees as they perform the workflow without AI. Record averages for time taken, error and rework rates, and costs per task.

3. Assemble a Realistic, Tricky Dataset

Gather 20 to 40 real work examples featuring unstructured, messy data. Deliberately include “drift” cases, such as missing files or dead links, to challenge both humans and AI. Create a checklist mapped to the evaluation rubric for binary pass/fail evaluation.

4. Set Up the Production-Like Workspace

Recreate the environment where AI will run in production. This includes using enterprise or team licenses if those are planned, not relying on free tools. Document every setting—model, tools, connectors, permissions—ensuring that participating employees all have appropriate access.

5. Demand Repeatability and Proof

Since generative AI yields variable results, every test must be run at least three times, both by humans and AI. Disable LLM “memory” features to prevent data leakage between runs. Require proofs—citations, file artifacts—for all accepted answers. Calculate a formal reliability score to capture consistency.

6. Calculate ROI with Objective, Blind Grading

Use blinded grading: those evaluating the outputs should not know if they came from human or AI. Standardize output format and style to remove giveaways; check citations and formatting before grading. Calculate savings by converting time saved into monetary value while subtracting AI subscription costs.

Report on at least these seven metrics: cost, latency, accuracy, reliability, safety, integration, and compliance.

7. Re-Evaluate Regularly and Adapt

AI models and modes change quietly—often without teams knowing. Re-test every month, and after every model update, to prevent quiet regression or discover unexpected improvements. Compare results to a three-month rolling average and investigate any significant shifts.

Beyond the ROI Calculation: New Team Structures

For organizations seeing measurable savings, the question becomes: what next? The most effective path is investing those savings into AI evaluation and implementation teams, whose sole focus is to monitor updates, re-run benchmarks, and introduce new, ROI-positive use cases. This muscle keeps AI investments productive and drives further efficiency.

Conclusion

This framework helps organizations break out of pilot purgatory, avoid ad-hoc or anecdotal evaluation, and establish repeatable ROI for front-end LLM adoption. Success hinges on methodically benchmarking, rigorously comparing, and ensuring continuous monitoring of AI capabilities in the face of constant change.

For organizations intent on turning generative AI into sustainable value, this blueprint provides the actionable, day-to-day playbook needed to select, deploy, and maintain the right LLMs as core business infrastructure.


Topics Covered in This Episode:

  1. Choosing the Right Large Language Model
  2. Evaluating LLMs for Business ROI
  3. Front-End AI Operating Systems Explained
  4. Common Traps in AI Model Evaluation
  5. Public Benchmarks for LLM Evaluation
  6. Seven-Step LLM Evaluation Framework
  7. Measuring Pre-GenAI Human Baselines
  8. Building Realistic AI Test Datasets
  9. Calculating ROI for GenAI Implementation
  10. Monthly Retesting and AI Model Updates


Episode keywords

Large Language Model, LLM, generative AI, AI operating system, front end AI models, AI evaluation, model ROI, model evaluation steps, AI benchmarks, scientific benchmarks, API connection, enterprise AI, ChatGPT, Claude, Gemini, Copilot, team AI adoption, knowledge worker AI, operating system choice, productivity modes, connectors, deep research mode, agent mode, image generation, web search, Canvas mode, advanced voice mode, business process automation, workflow evaluation, change management, AI training, human baseline measurement, time savings, value creation, prompt engineering, model selection, output accuracy, reliability score, AI-generated outputs, citation verification, subscription costs, cost savings, latency, stability, safety, integration, compliance, human-in-the-loop, model updates, retesting AI, monthly evaluation, AI deployment, AI implementation, public model evals, LM arena, Live Bench, Epoch AI, Scale CLLM Leader, use case testing, AI champion teams, team AI pilot, corporate AI adoption, performance measurement, AI-generated documents, synthetic data, data analysis, drift cases, test dataset, data-driven insights, continuous improvement.


Podcast Transcript


Jordan Wilson [00:00:47]:
How do you go about choosing the right large language model for your company and the right model for the right job. And how can you evaluate it and understand if you're actually getting a ROI on gen AI? This is something I've literally talked to hundreds of companies about over the past three years. And today I think it's important now more than ever, as large language models are slowly morphing under our eyes into AI operating systems, it's important we talk about how to do this for front end AI models, because I do think that's the future. That's where work happens as your everyday chat GPT, Claude, Gemini, etcetera, are becoming places of work and where teams go to get work done. So today we're going to go over what's the best LLM for your team and the seven steps to evaluate and create ROI for AI. I'm excited for today's show. I hope you are too. What's going on y'all? Welcome to Everyday AI.

Jordan Wilson [00:01:57]:
My My name is Jordan Wilson. We do this thing every single day. It's a live stream, podcast, and free daily newsletter helping us all make sense of AI, but how we can actually leverage it to grow our companies in our career. So if you're a stressed out business leader trying to keep up and you're like, how the heck can I when there's new AI advancements every day? Starts here with the live stream podcast, but to take it to the next level, make sure to go to our website at youreverydayai.com. Sign up for the free daily newsletter. We're going to be recapping both today's show and keeping you up to date with all of the latest AI news from today. So let's get into it. ROI is important, right? But how do you evaluate it? Don't get me wrong.

Jordan Wilson [00:02:38]:
There's great resources out there, scientific benchmarks, evaluation sites that look at large language models. But here's the issue. They're just looking at the models themselves. They're looking, mainly if you're using one of these big APIs, but that's not how companies are using AI now. Something I've been thinking about a lot over the last year or so. I've been one of the first people I think shouting and screaming about this AI operating system thing. And now as these AI chatbots are turning into full fledged operating systems. Now people are understanding, oh, maybe we shouldn't just be using models, on the back end via an API connection, and maybe we should be using them in their interface because they're so much more powerful when you have access to all of these different modes and features that aren't available when you're using an API.

Jordan Wilson [00:03:35]:
So now we've seen this influx of large enterprise teams moving away from using models just via the API or now bringing more of their team, but putting them, you know, on a Teams or an enterprise plan for chat, GPT, Gemini, Copilot, Claude, etcetera. But now we're left with this problem again. How do we evaluate this? Well, on today's show, we're gonna go over the common traps that businesses make when trying to evaluate AI models on the front end. We're gonna share the best publicly available large language model evals to at least get you a jump start, And we're going to lay out in detail our seven step plan for evaluating AI models for AI. Live stream audience. Good to see you. What's going on? If you guys have any questions. Good morning, Jason on YouTube.

Jordan Wilson [00:04:25]:
Douglas. Good to see you. If you guys have a question, feel free to drop it. I like when I can get a couple of questions. So if you have any, let's get them in. But first of all, who is this show for? I think this is for teams and organizations that are evaluating using front end AI chatbots as their AI operating system of choice. And I'm gonna be talking a little bit more about this concept of an AI operating system that's finally now picking up steam, even though I've been talking about it for more than a year, and why I think now is a more important time than ever to be talking about this concept of model evaluations on the front end, and how you can create and protect ROI. Because if your team is not properly trained and doesn't understand how large language models work, especially on the front end with all of these new modes that are being added all of the time, you might end up spending more time in AI than you would if you weren't doing it with AI, which is why I think it's important to evaluate these models and to go through this step.

Jordan Wilson [00:05:29]:
Why now? Right. Well, now I think the rest of the world has finally caught up to what I've been saying for a very long time. I think more and more smart organizations are starting to move their day to day business processes inside the front end of a large language model. So if you're brand new here, let me explain the difference between this back end API front end operating system. Right? So I think in early, you know, in 2023 in the earlier days, 2024 larger enterprise organizations, essentially, they fine tune these models from Google OpenAI, Anthropic, etcetera, for their own use. They bought, they built, RAG pipelines and they essentially created versions of these large language models for their, employees to use internally. Right? And those are a little bit easier to evaluate. You know, it's strictly you're working with a model.

Jordan Wilson [00:06:23]:
There's an API, there's benchmarks for certain categories of work, right? There's great benchmarks out there. You know, these models are great for, coding. These models are great for creative writing. These models are, you know, great for multi step research, etcetera. So when you're using an API and you have a little bit more, of a narrow scope, it's a little bit easier to evaluate these models and to measure ROI. But I don't think that's how most people should be using them. Right? I like to think of an actual operating system. In the nineties, you know, most companies or early two thousands had to make a choice on their operating system.

Jordan Wilson [00:07:02]:
Are you going to be a Microsoft Windows organization? Are you going to be Mac OS? Are you going to be Linux? Right? And then from there, you essentially built your processes around those operating systems. Right? When operating systems became, you know, popularized in the nineties, you know, work before computers kind of had to be, reworked and reorganized around what an operating system offered. And I think that's the junction that we're at right now. Smart organizations are understanding maybe using, at least for average, you know, your average knowledge worker, right? Obviously, if, if, if you're a coding shop, you know, if, if, if you're a niche company or a niche department, this is not for you. I'm talking about for the every day knowledge worker, right? People in marketing, HR, sales, right? Not necessarily, you know, people just doing coding tasks because I think that's a little more cut and dry. And now there's this whole concept, especially after OpenAI, earlier this week announced a couple of things at their dev day, but one of the biggest ones was apps. So bringing in entire apps into the chat GPT experience, I'm going to do a show on that later, but essentially what that means, and I think this is, something that the other, Frontier Labs are going to be doing is you're going to be working with your own business context and using entire user interfaces from other websites, all within ChatGPT or Gemini or Claude, etcetera. So we need to start unlearning the old way of working with AI, and we need to relearn how to move our day to day processes inside one of these systems, because that is gonna be the best and fastest way to work.

Jordan Wilson [00:08:54]:
Just like, well, you probably could have figured out in the nineties, MacGyver ing a way to not use one of those three big operating systems. But if you wanted to succeed, you had to, I think you have to make that same choice now. So here's the problem. Even front end AI systems are complicated, confusing, and change too often without notice. Here's here's a good example. So, for the podcast audience, I have something on my screen here. If you ever want the video version, make sure to go to our website at youreverydayai.com. You can watch the video version there.

Jordan Wilson [00:09:28]:
Click on episodes. Right now, this whole, you know, Chad GPT thing, going to GPT five was supposed to make things easier. Right? It didn't, it made things more complicated. And this is why I think it's important to look at maybe our company shouldn't just be using an API and maybe we should be moving our, processes inside of one of these, team modes. Right? So obviously, all the big AI models, you can have a team account, an enterprise account, bring thousands of users. You know, you can share, projects, share GPTs, share chats right through, through most of the, you know, front end AI chatbots. But it's confusing now. So even if you look at chat GBT on my chat GBT screen, let me count them here.

Jordan Wilson [00:10:20]:
I have one, two, three, four, five, six, seven, eight, nine. I have 10 different models that I can choose from. Okay. That's not easy. Right? You have all the different variations of GPT five auto, instant, Thinking Mini, Thinking and Pro, And then you have your legacy models. GPT four zero GPT four one GPT four five zero three zero four minutei confusing. And then even in the thinking mode, you have, I think four or five different layers of thinking. So you, I, I have more than a dozen and many teams have more than a dozen just models they can choose from just in chat GPT.

Jordan Wilson [00:10:58]:
All right. It's not as bad in Gemini, right? There's essentially two, you know, Claude is a little, you know, in between, I think there's about six, six different models, but then you talk about modes and this is where you don't get this. If you're just working, on the backend via an API and these modes are where the magic happen. So using modes like connectors, you know, and all the major, players have their version of connectors. This brings your dynamic data, essentially creating a mini version of retrieval augmented generation or rag pipelines with your company's dynamic data. There's the deep research mode, agent mode, right? I'm going over the modes in chat gbt, image gen mode, web search, canvas, which is so underrated, study and learn, right. Advanced voice mode. So would you rather going back, right? Think of those, those of you that, you know, worked in the, in the nineties, in the early internet days.

Jordan Wilson [00:12:01]:
Imagine if you had a computer with no programs, right? No, no office programs, no terminal, no nothing. That's kind of what I feel, if your company is only using something via an API. And this is why I think the smartest teams have already made the jump to doing a chat GPT business plan, a chat GPT enterprise plan, Google Gemini, team plan, Claude enterprise plan because you need to take advantage of these modes. These modes, right? Difference between models and modes are where work happens. So why then if this is, where work needs to happen and we know AI is so good and so powerful, and we see all these studies, oh, it's so much faster, more efficient, better than humans. Well, why do most companies never find an ROI? Let me give you the reasons. The pilots are always too long. You can't do a year long pilot in AI.

Jordan Wilson [00:13:03]:
That's gonna you're failing before you start. Change management and training are basically non existent in most, even large organizations. I'm constantly shocked. The lack of change management that gets, invested as well as training companies aren't training their people. It's weird. Also not properly measuring pre gen AI human input efforts. How long did these projects take before there was AI? Right? We don't have baselines. We don't have, you know, if, if you want to talk about ROI, what is it? Right? You have to calculate the time that it takes for your people to do a project.

Jordan Wilson [00:13:45]:
You have to calculate any hard costs, right? Software expenditures, whatever. But people didn't measure this before AI. So there's no baseline to compare it to. Also, I think sometimes one lucky run is celebrated when it comes to AI, but variance and reliability are never measured. Generative AI is generative, right? So a lot of times, companies find like one use case and then they just roll that very small use case out to everyone without properly testing it. Right. Generative AI can be kind of like a roll of the dice sometimes. So you, you have to keep that in mind.

Jordan Wilson [00:14:30]:
And then last, but definitely not least one of the common, most common traps of, companies not finding ROI on gen AI is the shiny object AI syndrome, right? Every single week in multiple times a week, there are shiny distractions, right? I just talked about OpenAI's big announcements. We're going to have big announcements this week and we've already had some big announcements from Google. I'm sure we're going to see next week, right? Every, I do this every day. I've been doing this every day for three years. There's a shiny object every single week, multiple times, and companies get too easily distracted by that. So let's talk about evaluations, right? If we're going to talk about how and go over these seven steps to evaluate, what are evals, right? So AI evals are essentially structured quality checks for AI systems. They're important especially for generative AI because you can run the same input and get wildly different outputs. That's how generative AI works, which is why evaluations are extremely important.

Jordan Wilson [00:15:36]:
So using these ensures reliability. So evals verify the AI performs its task consistently and correctly. You're able to catch risk. So evals identify issues like bias and errors before the AI is used kind of by the, by the rest of your company. And also this helps guide improvements, right? To get the most out of gen AI, you can't you have to constantly be iterating, in kind of going in a a cycle of improvement in feedback. And that's what Evals help do. So they provide data driven insights to help your team fix and refine the AI that you're using. And then there are some great publicly available eval sites to get started, right? Because yes, I mentioned chatgbt and they currently have like 12 different models, right? Claude, when you're using it on, and I'm talking frontendchatgbt.com, claud.ai, gemini.google.com.

Jordan Wilson [00:16:35]:
There's, there's multiple models to choose from. But there's good places to get start, to start, looking at evals. Okay. Yes. There are scientific benchmarks, but there's other great resources to at least get started. Right? Because I'm not saying that one model or one mode is gonna be best for every single project, task, or deliverable for your company. You might have to use multiple. Right? You might have, some people, you know, you might have a 100 employees on a chat JPT teams plan, and you might have, you know, 200 on a cloud enterprise plan.

Jordan Wilson [00:17:11]:
So where do you get started? Well, most of these publicly available, AI eval sites, look at different scientific benchmarks as well as user benchmarks. And then depending, you know, like LM arena is a good example. We talk about this a lot on the show, but with LM arena, you can go on there. You essentially have a battle. Okay? You put in one prompt. You get two different responses from a large language model. You choose which one is better. You don't see which model it is.

Jordan Wilson [00:17:43]:
It's a blind, LLM taste test. And then after millions of votes, obviously, you start to see which models are best. And then also they classify what this was about. Was this a creative writing prompt? Was it a factual, test? Was it a a coding task? Right? Was it math? Was it science? So not only do you get scores, these are called Elo scores, but you're also able to classify them, across different, arenas. Right? So if you're looking for the right for your web development team, you can go look at that. If you're a creative, creative writing department, you can go look at that. So different evals classified across different categories. LM arena, I think is one of the best public eval sites.

Jordan Wilson [00:18:26]:
Next live bench. Similarly, they look across, you know, reasoning, coding, agentic coding, mathematics, data, data analysis, all these different things. Then they look at all the different variations of the models as well. Another great one is, epoch AI. They have their AI benchmarking across different categories as well. And then last but not least, this is a newer one. I like this from scale. They're c l l m leader.

Jordan Wilson [00:18:52]:
So a lot of these look at different aspects. Right? Because your team, especially if you're a larger organization, like I said, you might end up using multiple of the big platforms. Right? And I don't think that's a bad thing necessarily. Although it's always best to bring everyone under one roof. Sometimes. Right. Especially like, let's say something like coding, right? That's Claude is very popular, for software development teams. So maybe, you have a team running on that and the rest of your team's on Chachika T or Gemini.

Jordan Wilson [00:19:23]:
Alright. We're gonna get into the seven steps. I had to lay the groundwork first. But before we get into them, very quick word from our sponsors.

Steven Johnson [00:19:33]:
This podcast is supported by Google. Hey, folks. Steven Johnson here, cofounder of NotebookLM. As an author, I've always been obsessed with how software could help organize ideas and make connections. So we built NotebookLM as an AI first tool for anyone trying to make sense of complex information. Upload your documents and NotebookLM instantly becomes your personal expert uncovering insights and helping you brainstorm. Try it at notebooklm.google.com. Alright.

Steven Johnson [00:20:04]:
I had to

Jordan Wilson [00:20:05]:
go through the, the prerequisites of laying the groundwork first. I couldn't just jump in and give you the seven steps, but now I think we're ready to go through them. So before we get started, just a little precursor things to re things to remember. All right. You have to plan your evaluation sprint first. You need to get written buy in from, you know, executive sponsors, whoever that may be, end users, IT security, all that. Right? So here I am crossing my t's, dotting my i's. Right? You need to get a clearance if you have to go take this through legal.

Jordan Wilson [00:20:36]:
Right? What data can we upload? What can we not? Right? This is this is all my like, I gotta get my fine print out of the way because you can't just follow exactly what I say. You gotta get the fine print. Alright? Also, first, like I said, use public benchmarks to start to see which models you should start to evaluate. And then you need to plan a two to four week sprint testing, just one workflow, freeze your model choice and ignore all shiny new objects. All right, here we go. Step one, define your success criteria before you start testing. You need to literally like write a job description for this pilot, right for this test. You need to, explicitly, identify the specific outcome, the constraints, the tools that you're going to be using are allowed, and then any do not do actions.

Jordan Wilson [00:21:23]:
Okay. You need to set clear ground rules because again, these models and modes are always changing. So, you know, if good example, right? A good example, Chat GPT just launched their apps. If you were midway through and you're like, Oh yeah, we could use this app for that. No. You said, this is what we're doing. We're not doing anything else. Okay.

Jordan Wilson [00:21:43]:
You gotta set your dues and your don'ts. Then you need to create a rubric, a grading rubric, you know, one to five, one to 10, whatever works best for you. Defining exactly what makes each eval or test pass or fail. All right. So essentially what you're going to be doing is you're going to be identifying use cases, right? It could be, you know, parsing, certain information from long PDFs and turning it into a table in a chart. Right? That's an easy example. That's that's a use case that probably a lot of companies might do. Right? There's a big industry white paper that comes out every month and your team has historically gone through it manually, pulled out certain insights and put those insights into a chart or a table.

Jordan Wilson [00:22:28]:
Right? You need something measurable that you can define the success criteria before testing something that has a very finite beginning and end that you can measure. Right? And then you need to choose for each of these different use cases, choose three to five simple KPIs that the workflow impacts. Number one, obviously, is human time spent, accuracy, revisions required, value created, etcetera. Right? Because, there's some things that might be gray area. Try not to, with your success criteria and your KPIs, try not to choose anything that's that's, too great. Right. Try to choose as many black and white things as you can. Right.

Jordan Wilson [00:23:08]:
And there's different ways you can grade it. And I'll talk a little bit more about that later, but that's step one, define your success criteria before you start testing. Step two, measure your human baseline first. I don't understand why no one's doing this. This is mind boggling to me. Like literally you have people out there running big organizations, running AI pilots And, you know, they'll tell me about it. They'll be like, okay, you know, we ended up doing this project, you know, here's, here's how long it took, you know, etcetera. And then I'm like, okay, what was it when, pre pre gen AI, pre gen AI? And they're like, oh, well, you know, we didn't benchmark it.

Jordan Wilson [00:23:46]:
It's like, okay, benchmark it now. Do that exact same thing without any AI. Right? I've been talking about this for years. It's very simple. You need to time multiple employees. It depends on how big your team is. Right? If it's a smaller team, you know, if you're running this through a champion team, a lot of times, you know, people identify their champion teams and it's like eight to 10 people. So it depends on the size of your team, but you need to have multiple employees run through the entire process.

Jordan Wilson [00:24:11]:
And a lot of companies aren't willing to do that. Y'all, you need to invest in the same way that you need to invest training and education, learning and development around AI. You need to sometimes have multiple employees do this. And I think so many businesses are hesitant because they look at this as waste. They're like, why would I okay. This is a forty hour project. Right? Why would I have three to five employees do this same project? And if it's just for an internal use case, okay, well, you see all these studies about 60 to 70% time savings when using large language models. Well, do you want that? Well, if so, you gotta get the pre the pre gen AI human benchmark.

Jordan Wilson [00:24:50]:
You mean to calculate the average time error rate, rework minutes, cost per completed task. You need to measure the human input and output. This is your, this is your baseline. And this will ultimately prove whether the AI actually saves time or not without it. You're just guessing. Stop guessing. Step three, you need to build a realistic and controllable test dataset. All right.

Jordan Wilson [00:25:15]:
So whether this is synthetic data related to your company, whether it's publicly available, publicly available data, you know, in your industry, whether it's data about competitors, whatever it is, you need a finite and a realistic data set. You need to gather 20 to forty, twenty to 40 actual work examples. You need to have messiness, right? Don't start with something that's too clean, too structured, too organized, Right? Both for the humans and the AI. You need to really have a real, dataset. Also, I would add six drift cases. You know, rename files, dead links. You need to put actual, traps in there, both for humans and for AI. So again, you know, you're not gonna have the team that's ultimately testing this on the human side, building this, use case.

Jordan Wilson [00:26:06]:
You need to have other people. Right? But you need to build traps because, especially when we talk about agentic models, you can't just hold their hand, right? You have to hope that if they encounter a problem, they will adapt, they will adapt and overcome it just like hopefully a human would. And then you also need to create a pass fail checklist for each case tied directly to your rubric. Step four, you need to configure your workspace like production. Okay. So you needed to set up properly. You need to set up the shared workspace matching a real permissions for the right tools and the models that users will ultimately have. So what do I mean by that? I don't know why so many organizations they're like, oh, I'm not going to pay for AI.

Jordan Wilson [00:26:51]:
I'm going to use the free one. Stop. Right. You need to be using the exact same mode or model that you want to be using on the front end or on, like once this goes, goes live with the rest of your team. Right. Sometimes that might mean paying for the super expensive $200 a month plan. Right? Because there are certain features that are only available on those plans. So it might mean that.

Jordan Wilson [00:27:14]:
All right. Also you need to document the configuration, which model, which tools can be used, which connectors, you know, who has what access level, make sure that the people on the team have access to the same documents. Right? Then you need to verify during the test that the AI uses the expected tools, not workarounds or guessings. This, this is the other thing. So many people don't know prompt engineering basics. So many people don't. All right. This is probably where, even if you follow this steps one through seven, this is probably where you're going to fail.

Jordan Wilson [00:27:51]:
Most people have no clue how to prompt an AI. They don't. And well, why do I know this? Well, people don't even know how to select the right bottle, Right? CEO Sam Altman, recently said around the GPT five launch that only I think 7% of paid users used thinking models, which is an absolutely asinine thing to think about. Right? Why would anyone choose models that are worse? This just shows most humans have no clue what they're doing when they're using large language models. You need to use the right model, the right mode, and the right prompting techniques. You've got to know the basics. All right. Step five.

Jordan Wilson [00:28:40]:
Generative AI is generative. You don't do one offs. You run this multiple times. I'd say you need to run at least this trial three times, both with your humans and on the AI side and demand proof. You need to repeat each test case three times in separate chats. Right? So turn off depending on what, you you know, front end AI chatbot you're using, but you'll you're probably gonna wanna wanna turn off memory. You're, if you're using ChatChippy Tea as an example, or all the, front end AI large language models have something like this. You're gonna wanna turn off memory.

Jordan Wilson [00:29:14]:
You're gonna wanna turn off past chat history, you know, to make sure that it's not pulling from other things. You also need to require working citations, file paths or artifacts for every accepted answer, no exceptions. And then you need to calculate the reliability score. Right? So when you are creating this rubric on the front end, you need to know how it's going to be scored on the back end because again, y'all generative AI is generative. You need to be very detailed in how you're going to score this thing, and then you need to run it through multiple times with humans and with AI systems, because sometimes generative AI might get it right. Sometimes it might get it wrong. Step six, you need to calculate the real ROI with objective grading. Right? Depending on the resources you have internally, I would have the grading, be from someone that doesn't know which ones are human and which ones are AI.

Jordan Wilson [00:30:10]:
Right? So this might also require you having a certain output to where human graders wouldn't be able to tell which one came from a human and which one came from an AI. So what that means, right, until, until the models get perfect, this might mean you you have a, one person or a small team go through and make sure that the, formatting, in the layout is the same between the AI created outcomes in the human created outcomes, just to make sure you remove biases from whoever is ultimately grading. Right? So simple example. Right. A lot of times, large language models will cite things, in line. Right? So you don't wanna just copy and paste that over, because someone's gonna know, oh, this is from an LLM. And then they might, depending on their objective and their agenda, they might grade it accordingly, either better or worse. So, you need to have that kind of blind taste test.

Jordan Wilson [00:31:08]:
Right? Whoever is testing or grading on the rubric can't know what came from AI and what came from a large language model. So same thing, right? Maybe it's a page minimum page maximum, whatever it is, you need to make sure that it is consistent and that human graders won't know the difference. So then ultimately you need to use the automated checks for citations and accuracy, and the humans can grade for tone and quality. You need to convert time savings to dollars using a fully loaded rate and then subtract the subscription costs. Right? Whatever you're paying, for these AI systems for net ROI. Right? So do this across the course of a month because that's, you know, you're paying for these are, you're paying for these AI systems monthly and then you need to report at least seven factors. And this is gonna vary, depending on, you know, your industry, your department, your type of work, etcetera. But I think these, seven factors are pretty good to keep a look, to keep a look at.

Jordan Wilson [00:32:06]:
So cost, latency, accuracy, stability, safety, integration, and compliance. Right? Because you don't know, like you don't want an AI going off the rails and doing something against company policies in the same way you wouldn't want a human doing something. Right. Safety, same thing, accuracy, same thing. All right. Stability. Will an AI system do great the first three runs and then go off the rails. The next three need to be measuring these things and reporting on them as well.

Jordan Wilson [00:32:34]:
All right. And then step seven, you need to retest this monthly to track changes. Alright. Here's something most people don't know. Let me just pull this out as an example. Let's say for whatever reason, you're using, GPT five thinking, as your model of choice, and then you're using canvas mode. Okay? Canvas mode gets updated often. Most people, unless you're a dork like me, don't know this.

Jordan Wilson [00:33:05]:
GPT five, it's gotten updated or sorry. GPT five thinking has gotten updated multiple times since it came out since August. Most people don't know. Right? There's actually a thinking slider now. Right. GPT five auto just got updated this week and you probably didn't know it. Right? That's why you have to retest monthly to track changes. Okay.

Jordan Wilson [00:33:26]:
You don't necessarily have to retest, the human input output. Right? Because that in theory is not going to change or maybe you need to, you know, do that at least quarterly or yearly because, you know, it depends on, again, it depends on your use case, but at least on the AI side, you don't just do this once, set it and forget it because it could get marginally better. It could get increasingly worse. Again, it depends. It depends on what these big AI companies are doing under the hood. Alright. You need to retest though immediately anytime a provider ships, model updates. So a lot of times they're doing these under the hood and you might not know unless you're reading the change log, but you probably should be reading the change log.

Jordan Wilson [00:34:11]:
And if there is a noted model change, you need to rerun the test immediately, not doing so is negligence, right? AI changes all the time. And the other thing is you still need human in the loop, right? You still need human in the loop, both on the testing phase. And ultimately once you deploy this and once you get to the point where yes, we've measured ROI in step six, right? We've taken this to production. It doesn't end there. You still need your human in the loop or expert, expertise driven loop, like I like to say, and then you need to retest monthly to track the changes. And also you need to track your trends versus a three month average, and you need to investigate if the accuracy or savings drop significantly. Like I said, when I just took you through this seven step plan. This is for one type of project.

Jordan Wilson [00:35:06]:
Again, let's go back to my easy, simple example. Right? You have a, a team or one individual that spends 40 hours, a month going through this large industry white paper pulling out, you know, synthesizing, personalizing information for your company and turning it into a little bit of a document. Right? Going through this 100 page report, here's everything that applies to us. Here's some spreadsheets. Here's some graphs, etcetera. Right? That's a simple example. A large language model can obviously do that, but what's the three month trend line look like? Is it getting better over time? Is it getting worse over time? Because if it's slowly getting marginally worse month over month, then you probably need to revisit the model. Right? Going back to looking at the publicly available model evals.

Jordan Wilson [00:35:52]:
So you need to also go through this depending on your use case or potential use cases. You might need to run this through multiple models, and sometimes you might want to do this in tandem. You might want to be testing, you know, Gemini 2.5 pro, alongside, you know, GPT five thinking as an example. Right? And you need to see be constantly reevaluating. One thing people are almost confused about, at least those companies that have already seen and measured their ROI on GenAI, right? They're like, okay, well, what do we do now? Right? You have two choices. You can cut people. Well, technically three. But one is you can reduce headcount and you do that either by laying people off or you stop hiring and let churn, and attrition do its thing, which is what, the latter I think is what a lot of bigger companies are doing.

Jordan Wilson [00:36:45]:
They just stopped hiring. And then they're just, you know, eating up these efficiency gains with AI. Or you can do something else. Right? You can, say, hey. We have huge time savings from AI, so let's start investing in new lines of revenue. Well, how do you do that? This. Right? Every single medium sized business and up. I, not all small businesses can do this, but if you have a 100 employees or more, which is the majority of everyone listening to this, this process that I outlined, this takes humans.

Jordan Wilson [00:37:18]:
You need to have AI testing and deployment teams, even if you're a medium sized organization, right? I'm not talking, oh, this is only for fortune, you know, fortune 1,000 companies. No, you need a team of me's. That's what you need. You need a team of people who are constantly paying attention to AI updates, constantly evaluating different models and keeping a very close eye on these use cases. That's how you get successful AI implementation. You have to invest in the people that are pushing you forward. All right. I hope this was helpful.

Jordan Wilson [00:37:54]:
Y'all as always, put together a little guide. All right. So if something in this struck you and you're like, yeah, my team needs to hear this, but we need a little bit more than this podcast. I obviously put together, an extended guide. So, if you're listening on LinkedIn, go ahead and repost this show. I'll send it over to you. It's hot, fresh, like little Caesar's pizza, ready to go. But if you're listening on the podcast, always check the show notes.

Jordan Wilson [00:38:25]:
Okay. In the show notes, there's a link to this LinkedIn post, right? Each of our live stream podcasts goes out on LinkedIn. So if you want access to the guide, go find in your show notes on Spotify or Apple Music, whatever you're listening to, go click that and then go repost this and I'll send that over to you. So I hope today's show was helpful. If so, make sure you go to youreverydayai.com. Sign up for the free daily newsletter. We're gonna be recapping the highlights and main points from today's show and keeping you the smartest person in AI at your company. So thanks for tuning in.

Jordan Wilson [00:38:56]:
Hope to see you back tomorrow and every day for more everyday AI.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI