Ep 703: AI Hallucinations: What they are, why they happen, and the right way to reduce the risk

Resources:

Join the discussion: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


AI Hallucinations: Specific Risks, Tangible Solutions, and Business Value for 2026

The topic of AI "hallucinations" is evolving rapidly—and for business leaders, the implications are both practical and immediate. Once dismissed as a fundamental flaw, hallucinations (confidently incorrect or fabricated AI outputs) now represent a manageable property of advanced AI systems. The latest insights from the Everyday AI podcast episode reveal how improvements in large language models (LLMs) and best-practice workflows are redefining how businesses mitigate risks and harness AI’s capabilities at scale.



Understanding AI Hallucinations: Defining the Real Risk

Hallucinations in AI are not simply random mistakes; they are a product of how LLMs generate output, fundamentally based on pattern recognition and next-token prediction. When the model lacks sufficient context or encounters ambiguous requests, it may create plausible-sounding but false responses. In legal and financial sectors, this has led to high-profile embarrassments, including lawyers cited for filing court documents with fabricated AI-generated citations and consulting contracts laden with phantom sources.

Recent cases, documented by the AI Hallucination Cases Database at HEC Paris, counted 486 legal incidents worldwide involving fabricated content—a consequence of widespread enterprise adoption without sufficient training or workflow safeguards.

Advanced Large Language Model Improvements: Measurable Error Reductions

Current generation LLMs—from OpenAI, Google, and Anthropic—have achieved significant reductions in hallucination rates. Early models like GPT-3.5 would fabricate up to 40% of academic citations, while subsequent versions dropped error rates below 7% for general queries. This jump is due to the models’ enhanced ability to handle longer context windows: in technical tests, latest models maintained 95%+ factual recall even when pulling from hundreds of thousands of tokens, compared to the sharp drop-offs seen just a year prior.

Businesses can now benefit from LLMs that not only think and reason across extended conversations but also offer summarized chain-of-thought outputs, enabling deeper visibility into how conclusions are reached.

AI Hallucination Mitigation: Four-Layer Workflow for Business Reliability

The reduction in hallucination risk hinges less on “better tech” and more on the implementation of a structured, multi-layered workflow:

  1. Model Behavior Instructions: Custom instructions, embedded within enterprise AI tools, can force models to signal uncertainty, provide confidence scores on outputs, and require source attribution for each factual claim. These protocols prevent models from guessing or fabricating when appropriate context is missing.

  2. Retrieval-Augmented Generation (RAG) & Data Grounding: Modern platforms offer one-click connections to internal business data (Microsoft 365, Google Drive, OneDrive, SharePoint), allowing models to “ground” responses in company-specific facts. A Stanford study found RAG-based workflows combined with human feedback reduced hallucinations by up to 96% compared to baseline models.

  3. Verification Workflows & Expert-Driven Review Loops: Instituting a second-pass review—whether through another model or an expert-driven loop—enables organizations to systematically validate outputs before delivering high-stakes recommendations, filings, or customer communications.

  4. Chain-of-Thought Traceability & Agentic Observability: Latest agentic models document their logical process, giving users side-by-side visibility of both factual claims and underlying inferences, allowing for targeted auditing and correction.

Business Implementation Value: Concrete Organizational Impact

The absence of robust training remains the single largest gap in enterprise AI rollouts. As companies distribute AI licenses to thousands of employees, failure to educate on model selection, data grounding, and error-checking perpetuates costly—and sometimes public—hallucinations. In law, finance, and consulting, high-visibility mistakes result in project refunds and reputational loss.

Adopting the outlined workflow delivers immediate business value: drastic reduction in AI fabrication risk, improved customer trust via monitored AI voice agents (flagging abuse, fraud, or false claims in real time), and safer, more effective collaboration between human experts and AI systems.

Conclusion: Managing AI Hallucinations for Long-Term Success

AI hallucinations will not vanish but, with correct workflows, businesses can treat them as a manageable property rather than a disqualifying flaw. Winning organizations are utilizing latest-model capabilities and systematically integrating multi-layer error reduction into every AI-enabled process. By combining precise data grounding, proactive model instruction, verification loops, and agentic auditability, companies are mitigating risk, optimizing AI ROI, and confidently deploying LLMs for high-value tasks.

For those looking to strengthen their AI implementation, the actionable insights demonstrated here are road-tested and accessible—no technical expertise required, just a commitment to best-practice integration and ongoing user education.


Topics Covered in This Episode:

  1. AI Hallucinations Definition and Causes
  2. Large Language Models' Hallucination Mechanisms
  3. Hallucination Types: Fabricated Claims & Sources
  4. Model Improvements Reducing Hallucination Rate
  5. Context Window Impact on AI Accuracy
  6. AI Hallucinations in Legal and Enterprise Settings
  7. Four-Layer Method for Minimizing Hallucinations
  8. Custom Instructions and Retrieval-Augmented Generation
  9. Expert-Driven Verification and Agent Safety Practices




Episode Transcript 



Midroll [00:00:00]:
This is the Everyday AI Show, the everyday podcast where we simplify AI and bring its power to your fingertips. Listen daily for practical advice to boost your career, business, and everyday life.

Jordan Wilson [00:00:18]:
Most AI tools that analyze sales or customer support calls turn those conversations into a text based transcript, but text only transcripts miss most of the value, and Modulate fixes that. Modulate's new Velma voice native AI model and their ELM technology actually understand what's happening on those calls. It picks up the valuable tone, timing, emotion, and intent that all AI transcription tools can't provide. So whether it's for sales, customer support, or voice agents, Modulate's new Velma model helps you capitalize on what text only AI tools miss. Demand more from your AI today with Modulate at modulate.ai. Let's talk about the elephant in the room when it comes to AI, that one word hallucinations, right? When you rely on an AI chatbots to provide you something in it straight up lies, or it tells you a very non truthful version of what you might be looking for. And if you don't keep up with the advancements of AI, you might still think that AI lies all the time. And, well, if you're using the wrong model, if you're not using best practices, hallucinations even in 2026 can still be a huge problem.

Jordan Wilson [00:01:38]:
But here's what not enough people are talking about. With the latest technology in today's thinking models and the ability to very easily ground responses in your company's data without much techno how, essentially get rid of at least the high rate of hallucinations. So that's why we're gonna be diving into this sticky topic on the fifth volume of our start here series, talking about AI hallucinations, what they are, why they happen, and the right ways to reduce the risk. Alright. Let's get into it. I'm excited for today's conversation. I hope you are too. If you're new here, this is our fifth installment of the start here series.

Jordan Wilson [00:02:24]:
This is the essential podcast series to both learn the AI basics and to double down on your AI knowledge. So whether you are a beginner trying to understand where do I start with AI or you're really just trying to deepen your expertise, our start here series is definitely for you. And if you haven't already, please go to starthereseries.com. Here's why. Well, number one, that's gonna give you access, free access to our inner circle community. Alright? And you're gonna get straight into, our start here series, channel there. So you can go back and listen to all of our, start here series, right there in our free community, but you can also get free access to our, prime prompt and polish, prime prompt polish, chat GPT prompt engineering course. Alright.

Jordan Wilson [00:03:18]:
And if you miss our last episode of start here series that was volume four, we talked about the human AI collaboration and best practices for working alongside AI, which leads us right into today's show. Because if you are doing the best practices of human AI collaboration, what we tackled in volume four, then that leads you straight into reducing hallucinations. Because if you are working with AI in the right way versus blindly trusting its outputs, well, you're gonna, be able to fight back against the hallucinations. So here's what we're gonna cover on today's show. First, we're gonna tell you what hallucinate what hallucinations are, and why they are there. Then we're gonna show you how to assess the risk and talk about how models have improved on hallucinations, but they still run into them. And then last but not least, I'm going to suggest to you a kind of four layer method to reduce the errors and show you how you don't really need to be that scared of hallucinations as long as you are being a smart human. Alright.

Jordan Wilson [00:04:21]:
Let's jump into it and start at the beginning. What the heck are hallucinations and why do they still happen? Well, in general, here's as as bluntly as I can put it. Large language models. Right? And go back and listen to, volume one and volume two, if you need a refresher on how large language models are built. But, essentially, they scrape the entirety of the Internet online, offline datasets, and then humans train these models. So when you and I ask a question, right through this process called reinforcement learning with human feedback, these smart people at, you know, OpenAI and Google and Anthropic and all these other labs have trained models that when someone asks about x y z, here's what you should respond. However, for the most part, AI models are very super smart next token prediction. Right? So sometimes that means if something was maybe incorrect on the Internet or if a model is confused about what you're actually asking, it might present a made up answer, and it might do so very, confidently.

Jordan Wilson [00:05:23]:
Because at their core, AI models are trained to be helpful assistance. That is in almost every single system prompt, that an AI model uses. So that's why sometimes they are gonna make things up because they want to be helpful more than anything else. Right? Because they predict the next word from patterns, and that's kinda why they exist. So it's almost like maybe you've heard this, this saying, that hallucinations are a feature, not a bug. Because at the same time, kind of that same methodology of how they're built help large language models be extremely creative, strategic, and even if we're talking about, you know, scientific discovery, drug discovery, mathematics, that's actually how they're able to solve new problems that humans haven't been able to. But I think a lot of people, number one, are using the wrong model. They're not going through the, you know, context engineering one zero one of how you should be working with a model, and they're using right.

Jordan Wilson [00:06:21]:
They're using like I said, they're using the wrong one. That's why hallucinations have been a huge problem, and they can get you and your company in a lot of trouble if you're not doing the best practices. But, here's what's changing. Models, now they can think and they can reason much like humans do, and that has led to a drastic reduction in hallucinations. Right? You're obviously, the average person is using way more inference. Right? They're eating up way more tokens, to go through basic problems, but that's why I think hallucinations are really on the, decline. So the combination of better models that can think and reason like a human do combined with kind of this rag esque, features that front end large language models give you, are really leading to that reduction. But this is how they work.

Jordan Wilson [00:07:09]:
Right? Because people are always confused. Like, hey. If I give a large language model a simple problem, why does it get it wrong? Or, you know, a a a riddle or asking it how many r's there are in the word strawberry. Why does it get it wrong? Why is there this, you know, this really jagged, you you know, almost polarizing output that a large language model can do. It can do something that's absolutely genius, but then it can get something absolutely wrong. Right? So that is well, because they are next token or next word predictions. They are based on patterns trained by humans. They don't verify the truth against reality necessarily, and it's that same capabilities that allow AIs to be extremely creative and to generate code and to write poetry.

Jordan Wilson [00:07:54]:
Well, at the same times, that can lead them to lie to us. Right? And by definition, a hallucination is just a confident answer containing either fabricated claims because the model just generates the text or it can make up sources. Right? Which is a big problem as well. So I'm gonna be talking about, when we go over our kind of four layer, approach, how you can avoid that type of hallucination. So, there's really different categories of hallucinations. Right? One is it just lies. One is it it's not really lying, but it's sounding overly confident and just giving you some generic information. And then other times, it can just make up sources.

Jordan Wilson [00:08:36]:
Right? So it might give you true facts, but it might make up where it got those facts from. Right? And in most cases, like I said, hallucinations, if I'm being honest, in 2026, they're more human error than AI error. And I know a lot of people won't agree with that. That's fine. But if you get me a room full of humans that are using AI models the right way. Right? Get a a room full of humans that have taken our free, you know, prime prompt polish, courses. They're gonna run into very few hallucinations, right, compared to the average user. Because if you actually know how models work, how to feed them the right data, and how to check and verify on the back end being a smart, you know, human user, you're not gonna run into them very much.

Jordan Wilson [00:09:22]:
But let's talk about the early days because it was bad. Right? So go back to the early models, the GPT three of the GPT 3.5, you you know, the early days of, CHAD GPT in late twenty twenty two. Studies showed that GPT five fabricated up to 40% of academic citations. Right? Fast forward to the next version of GPT four that went down to about 29%. So, hallucinations got a little less, but still rampant. Fast forward to today, g p t five two reports a 6.2 error rate on general queries, and OpenAI claims a 30% reduction in errors in g p t five two versus g p t five one. So we're not talking, you know, 30% fewer errors between GPT five two and that version from three years ago. No.

Jordan Wilson [00:10:17]:
We're talking about in a three month period. And that jump is huge. And I'm gonna explain why, that jump has occurred, and I think why, you know, in a year from now, we may be not even talking or talking very much about hallucinations. And the main reason is just the ability for these labs. Right? So I'm gonna give an example here from Owen AI, but I think that, Anthropic in Google and OpenAI have made tremendous strides, with their models in, through a lot of different techniques, which I'm not gonna get into the technical side. They've made AI models that have much more reliable outputs. And one of the reasons, is their ability to handle longer context. Alright.

Jordan Wilson [00:11:08]:
So for our livestream audience here, I have a screenshot from, OpenAI's GPT five two, model release, and I wanna kind of talk about this long context. So there is a test. It's essentially called the four needle test. So what this is, they have it, pull out, and they ask the model questions. And kind of it's it's a needle in the haystack test. And they see over a wide range of a conversation. Right? If you're using, a a model into the hundreds of thousands of tokens. Right? Like, at 256,000 tokens.

Jordan Wilson [00:11:43]:
So a very long conversation because what has happened in the past, a lot of these hallucinations come, when you are hitting a longer point in the context window. Right? So think of the 3PM brain fog. Right? Let's just say you work nine to five. At 3PM, you're probably not as sharp as you were at 09:30AM, right, when that second Nespresso hits and you're like, let's go, and you're firing all cylinders. Right? For the most part, that's how large language models had been, I would say, even late into 2025. But think of that nine to five instead as a context window. Right? Because all large language models have their constraints. They have their, kind of confines that you can't break through, and one of those is the context window.

Jordan Wilson [00:12:30]:
That's how much information a large language model can retain until it starts to forget. And, potentially, when it starts to forget, it will start to hallucinate. So even great models like g b t five one, as I go here on the results of that four needles test, you saw its ability to properly pull out those facts over a large context window. When it started out, right, it was, you know, about 95%. It was very good at GPT five one thinking. But then toward the end of that context window, it dropped down to, like, 45%, a 45%, ability to recall that information. Whereas, g p t five two thinking hardly any decline at all. Right? I believe it was at, like, 95 or 96% at the end of the context.

Jordan Wilson [00:13:23]:
K. So think that very drained human at 3PM, they might not be able to recall facts. Right? But with today's and when I say today's, I'm saying the latest generation, Gemini three Pro, Opus 4.5 from Anthropic, Claude Opus 4.5, and GPD five two thinking from OpenAI. Their ability to recall information across that context window has greatly improved, which is one of the main reasons why that hallucination rate has gone down because now models can think. They can only think and plan ahead and reason and call tools on their own to, provide you more accurate information, but they're able to do that across the entirety of their context window, which is one of the main reasons why hallucinations are going down. So why does all this matter for your business? Alright. Well, quick, quick word from our partners, and then I'm gonna answer that question. Fraudsters used to need fake documents or stolen credentials to scam a business.

Jordan Wilson [00:14:24]:
Now they just need a few clicks to get your CEO to say anything they want. Voice deepfakes rose more than 680% last year. And most fraud detection systems only flag suspicious transactions after the money is already gone. They're not actually listening to the call where the scam is happening. Modulate is. Modulate's new Velma model analyzes live conversations for the signals that give fraudsters away. Stress patterns that don't match the story they're telling, urgency that sounds performed instead of real, voices that are synthetic instead of human, all detected for your business in real time, not in a report days later. Modulate's Velma model was trained on twenty one billion minutes of real audio and is trusted by Fortune 500 companies.

Jordan Wilson [00:15:11]:
It outperforms voice models from leading AI labs, and it's a 100 times more cost effective. So go see Velma catch what your current tools miss at modulate.ai. So here's why this still matters for your business. Well, let me cut it to you straight. One of the biggest gaps right now and one of the most blaring, things that are going wrong with AI implementation across the enterprise is a lack of training and education. Because you have companies that I've talked to personally, there's plenty of case studies out there that are rolling out access, whether you're talking about, you know, Microsoft Copilot three sixty five licenses, you know, ChatGPT Enterprise, Claude enterprise, Gemini business, Gemini enterprise, etcetera. They're rolling out AI access. You know? A lot of this was in 2025 to thousands tens of thousands of employees, but not giving any best practice training or even education, which is why hallucinations are still rampant.

Jordan Wilson [00:16:17]:
People don't even understand, oh, I need to, you know, click this model selector and choose the best model for the job, or I should be having this model, you know, call a certain tool. It should be running Python. I should be uploading files. People know know the basics. Yeah. They're just trying to save as much time as possible with AI, and that's led to a lot of hallucinations with a lot of, high profile and a lot of press around it. Right? So, as an example, there was, the AI hallucination cases database at HEC Paris documented 486 different legal cases worldwide involving fabricated AI content. Yeah.

Jordan Wilson [00:16:55]:
A lot of these, like, hallucination stories, they come at the most embarrassing point, which is in the legal sector. Right? So, there's been over a 128 lawyers that have been cited for filings with hallucinated cases. One of the most popular is the I think it's the Meta, versus Avianca. I don't know if I got the pronunciation of that right. I might hallucinated the pronunciation, but that's where you saw, attorney sanctioned for six fake citations. That was one of the more, infamous cases of AI in legal early on. And even, Deloitte reportedly refunded part of a $300,000 Australian government contract after it was, they reportedly found some AI generated phantom citations in their report. So this isn't just people who are, you know, using AI to write blog posts.

Jordan Wilson [00:17:44]:
Right? A lot of times, some of these hallucinations make their way to the forefront in very high, high value and just extremely visible places. Right? Like consulting companies, like lawyers. We've seen plenty on the financial side as well. So that's why this still matters. Right? Because maybe what I laid out for you and, like, saying, hey. If you use the right model and, look at these, you you know, the needle in the haystack, that's improving the context when well, most people aren't using models the right way, and that's why this is still extremely important for your business. Don't worry. We're gonna lay it out for you here, but I want I want to be very clear just because I am extremely optimistic about hallucinations decreasing with proper human involvement and training.

Jordan Wilson [00:18:34]:
Doesn't mean they're going away. Right? The very nature of what a large language model is, how it works means that hallucinations will be there for a while. Right? Because the models are always going to optimize for what word or what sets of words could be next even if those things are incorrect. Think of how many times that you've been on the Internet researching something, and you're like, this isn't right. That's not right. Right? In the same way, well, large language models take all that information from the web. So let's say you are a domain expert and there's something, you know, that in your field people constantly get wrong or there's some, some pieces of information out there that exist that are circulating that aren't exactly true. Well, in the same way that you might find incorrect data on websites or you might go to a conference or hear colleagues speak about something that you're an expert in and you're like, nope.

Jordan Wilson [00:19:24]:
That's not right. Well, a lot of times, large language models are reflecting inaccuracies, or, you know, half lies or half truths that exist in the real world. Second, the second reason why hallucinations probably are gonna go away is well, when their needed information isn't there by default by default, large language models are gonna fill in the gap. Right? They're gonna do everything they can to be a helpful assistant because that is the default behavior unless you change it and you should, and I'll tell you how. And then third, the chat infer the chat interfaces rewards that fast confident answers. Right? I'm not gonna get too much into, you know, reinforcement learning with human feedback and, kind of scaling laws and inference and all that. But for the most part, chat bots are trained to quickly give you an answer and being token efficient. Right? So if it thinks, oh, I can spit out an answer rather quickly, it may just do that unless you've told it not to.

Jordan Wilson [00:20:22]:
Right? So, they'll, at times, will reward kind of, giving you an answer even if it's not confident versus just saying, I don't know. I will say this model's more of late twenty twenty five because the models from 2026, you'll see by default, again, using the right models, in the right context. They are gonna say more, and I'd love to hear if you've seen this more. I have. Miles that actually say, I don't know or I'm not sure. Right? Even without giving, you you know, anything extra in your prompt, anything in the special instructions, today's models are more likely to say, hey. I'm not I'm not sure, which is a good thing. So here's how to spot and cut down on hallucinations a quick four step plan.

Jordan Wilson [00:21:07]:
Alright? Number one, you need to change the model's behavior. Like I said, I do think that, I think that in the future, the AI the AI labs are gonna find that sweet spot between, instructing models to be helpful assistance, but not too helpful and not fabricating things. But in the short run, you can do that. Right? So whether it's in your prompting sequence when you are working with a large language model or what I would recommend is setting custom instructions. Right? So, there's different kind of places that you can set custom instructions. But, if you are nontechnical, if you're not really sure what that means, that's in there's settings in most of the big providers that you can essentially put in your own rules. And any chat that you use, any response, that you are, using the model, it will always go through and read and usually adhere to the own set of custom instructions that you put in there. So, you know, even putting something simple, like, if you're not certain or the information isn't provided, say, I don't know rather than guessing.

Jordan Wilson [00:22:13]:
Something simple as that. I I have a set of, custom instructions, that I've done a lot of testing with over the years that are much more robust. But even putting something simple like that, like, hey. If you're not a 100% sure, say you're not sure. Right? Or to require every factual claim to include a source or, you know, labeling everything afterwards. Right? Saying, hey. This is what I'm a 100%, you know, telling the model. This is what I'm a 100% sure on.

Jordan Wilson [00:22:40]:
This is what I'm not sure on. You can even do something like, in each, response, having it give you giving you a confidence score. That's another great thing that I that you could do. And then have it to separate facts from inferences. Right? Because here's the thing, especially if you're using it as a strategy, creative partner, brainstorming, not all those things are black and white. Right? Maybe 99% of what you might use a large language model for is in that gray area. Right? It's strategies. Creative.

Jordan Wilson [00:23:10]:
Right? Being creative. So have it separate facts from inferences. You can ask for a table with columns for confirmed fact versus assumption in a structured output. So number one is doing that either in your prompt or in your custom instructions. Number two, my gosh, the fact that we have this and you don't have to pay any extra is wild. Right? Make it retrieve the information. Right? This whole context engineering, something I've been teaching since 2023 before it was a thing. Right? If if you've taken our free prime prop polish course, this is the refine queue, in the priming.

Jordan Wilson [00:23:51]:
So, we call it the fetch in the insights, but essentially, using a smaller simplified version of RAG, retreat block meta generation. Now you have a simplified version of rag available with a couple of clicks by connecting your company's data. Right? So both in, obviously, Microsoft three sixty five Copilot, in, Claude, in Google Gemini, in OpenAI's chat, GPT, they have different ways for you to connect your business data in a few clicks. So you can essentially, ground it in a way you can ground it in a way you kind of can't. But the combination of number one, those kind of custom instructions. And number two, first putting your company's information. If you combine those two things and you say, always check, you know, this a b c document, before you respond. And then if you make sure you check and verify that it did, I mean, those two things right there, amazing.

Jordan Wilson [00:24:47]:
But it doesn't matter if you're, you know, a Microsoft organization. You can connect your, you know, your OneDrive and SharePoint data, in OpenAI's products. Right? If you're using, you you know, Microsoft three sixty five Copilot, like, online, their online version, you can connect your if you use Google, if you use Google Drive. Right? So being able to connect your company's data to a large language model and then instructing it to always, look at those things first is huge, and it makes a big difference. So there was a 2024 Stanford study that found RAG combined with reinforcement learning with human feedback and guardrails achieved ninety six percent hallucination reduction versus the baseline. Right? Just doing those basic things. Right? Like, having a version of your company's data and proper, instructions are going to cut down on hallucinations, just those two things alone. And then steps three and four kind of combined into one is just the verification workflows and agent safety.

Jordan Wilson [00:25:49]:
So it's kinda like three and three b. Alright. So what is that? Well, you have to be able to catch errors before they escape. This is the, the expert driven loops that I always talk about, not the lazy passive human in the loop. I'm talking about the active proactive expert driven loops. So as these models, they think by default. You know, I I say g p t five two. Right? And if you're listening to this, this episode in July, right, maybe it's G p t five three or Gemini 3.5.

Jordan Wilson [00:26:18]:
I don't know. But today's latest models are agentic by nature. They're gonna go through they're gonna think they're gonna decide how much to think. Should I spend five minutes on this answer? Should I spend thirty seconds? Right? But you have control over that. And you can always kinda build, and I always recommend doing something like this. Build a second pass review where that model's only job is checking claims. That's why I also do kind of a mixture of models set up. But I have systems set up where, you know, I maybe have a Google Gemini deep research run, and then I'll take some of those results, and I will verify it.

Jordan Wilson [00:26:55]:
And I have something set up. The model that I like doing this for is g p t five two pro. I will then use g p d five two pro only to go through and verify every single thing that I get from a different model. That doesn't mean I don't trust model a versus model b. Right? And sometimes I'll flip flop it. It means you should always be doing this. Right? Especially for high value, highly visible projects. Right? If you're just trying to see, like, you know, what's the weather next week or, you know, something topical.

Jordan Wilson [00:27:25]:
Right? Like, what's the best you know, what are the three best softwares for tackling this issue? I don't think you necessarily need to go through all those steps. But if this if you are using an AI model as part of a high value workflow, you should definitely be doing this second pass review. And then, again, requiring even that second pass to show the sources. And then more than anything, you need to be tracing and observing how these models are getting to those conclusions. And what that means is, well, you should be using a thinking model and then reading the summarized or the chain of thought. Right? So what that is, most of the models, you know, different models give you different level of, visibility on what they're actually doing under the hood. But just if you were to hand off an important assignment to a brand new employee, hopefully, you wouldn't just hand that off to the client. Right? You would take a step in between and be like, hey, new employee.

Jordan Wilson [00:28:18]:
Tell me how you got to this answer. Right? Let's say they spent an entire week on this big project. I would hope you would take at least an hour or so to sit down with that new employee and be like, okay. How do we get to this conclusion? Walk me through. Right? So the good thing with these agentic models by default, you can see all that. And you, smart human, need to be able to go through and check and see what it did at each and every step. Did it follow your instructions of, you know, here's here's what's factual. Here's what I inferred.

Jordan Wilson [00:28:47]:
Did it make sure to look at the right documents at the right time? You can go through and sequentially check all of those things. And if it didn't do it, then you can course correct. But that is layers three and four, and especially as AI can take action. Right? As we're talking about agentic AI, that's when you really need to treat it like a junior employee. But, no, they can still make things up. And like I said, most of the time, if you are doing the proper context engineering one zero one, if you're using the right model, and then if you're going through kind of layers three and four right here with the second pass in the observability and traceability of going through that chain of thought, hallucinations aren't gonna be the biggest problem for you. Right? You still need to have those expert driven loops because as you're looking at that chain of thought as you're providing data on the front end, you have to do that at an expert level. You can't just say, here's a thousand files.

Jordan Wilson [00:29:42]:
Good luck, agent. No. You need to, you know, point them in the right direction. Same thing checking responses on the back end. You have to know what all of that means. But if you have expert driven loops and if you go through those four steps that I talked about, I think hallucinations are no longer gonna be the big elephant in the room. You're gonna be able to reduce them. Alright.

Jordan Wilson [00:30:03]:
But remember, it's not going away. Hallucinations are a property that you need to manage, not something that you hope will get fixed one days. And it's not that the winners right now are picking the best models. They're just building those four steps of verification in every single, episode or into every single piece of work that they're doing. Alright. Speaking of episodes, I hope this episode was helpful, because the start here series, I want you to know all the details. Right? As someone that's done this everyday AI thing now more than 700 times, I've got to speak with some of the smartest people in the world. I've realized if you're educated, if you're trained, if you understand how these models work, and if you keep up, that's that's the big caveat there.

Jordan Wilson [00:30:56]:
Right? If you keep up, you don't have to be as worried about hallucinations. I'm not saying you can write them off. You shouldn't. But if you're going through the right steps and what we went over in today's episode, you are definitely able to reduce the risk and to get more out of large language models because that's what we're all about here at everyday AI, cutting through the fluff, giving you the facts, and the right information to grow your company and your career. So I hope volume five of the start here series was helpful. I hope you know more about hallucinations now, why they happen, and how you can reduce the risk. So if this was helpful, please go to starthereseries.com. That's gonna give you free access to our inner circle community, the prompt engineering course, as well as a easy spot to go listen and catch up and engage with others who are going through this start here series with you.

Jordan Wilson [00:31:51]:
So thank you for tuning in. I hope to see you back later for more everyday AI. Thanks y'all. The risk with AI voice agents isn't that they sound too robotic for your company to use. The real risk is that they can sound too confident while saying something completely wrong to your perspective clients or customers. Made up refund policies, promises your company never approved, or discounts that don't even exist. You've gotta give your AI voice agents a trust layer with Modulate. Modulate monitors live voice conversations to flag abuse, false claims, fraud, and user emotions for safer, more empathetic responses.

Jordan Wilson [00:32:33]:
For the guardrail layer you need between your AI agents and your customers, you need Modulate at modulate.ai.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI