Ep 800: Celebrating our 800th Episode: 8 AI Truths, 10 Smart AI Moves and 10 Questions You Must ask

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


8 AI Truths and 10 Smart AI Moves: Critical Insights for Business Leaders at the AI Tipping Point

AI’s impact is accelerating at a pace even seasoned observers struggle to fully process. As seen in a comprehensive analysis on the 800th episode of Everyday AI, recent developments and practical lessons are reframing how effective organizations approach artificial intelligence—from model dependencies to workflow redesigns and cost control. This article provides a granular translation of those insights, organizing them into actionable concepts for business owners, executives, and decision-makers seeking sustainable AI advantage.

    1. Single Model Dependence Increases Business Continuity Risk
      Relying on one AI model, whether for operational automation or customer interfacing, exposes organizations to sudden disruptions. Recent regulatory actions, such as abrupt export controls on leading foundation models, have left companies without access overnight. Organizations building workflows that tie into just one engine risk halting mission-critical operations with no fallback, emphasizing the need for modular, multi-model architectures.
    2. Competitive Advantage Has Shifted Layers
      The defensible edge now lies in how companies build and orchestrate AI adoption. Rather than only focusing on which model performs best, it’s the operating, connector, and plugin layers—sometimes called “harnesses”—that matter most. These layers regulate how models are switched, combined, and updated without disrupting service. Technical decision-makers are evaluating multi-model harnesses, open connectors, and tools that can flexibly integrate various AI engines at any time.
    3. Prompt Engineering Is Losing Out to Workflow Automation
      Whereas customized prompts delivered some early wins, the sector now sees value in large language model automations, agents, orchestration, and scheduled tasks. Standardizing skills and automating deliverable generation has rapidly outpaced manual prompting as the primary return lever. Organizations trailing in this shift are spending time inputting prompts when they should be curating outputs for quality and business impact.
    4. Static Artifacts Now Constitute Organizational Debt
      Files passed by email, stale shared drive content, and static business documents are slowly accumulating as friction points. New AI-powered workspaces—like markdown-based live files and dynamic sites—replace the need for endless file versioning and approval cycles. Time wasted reconciling files or clarifying “which version is current” can be recouped by moving toward AI-native, continuously updated artifacts.
    5. Outdated Processes Exacerbate AI-Driven Workload Increases
      Simply layering AI tools atop unchanged workflows leads to unproductive complexity—more outputs, but no measurable ROI. Without redesigning jobs and processes for AI-native execution, teams end up managing old steps alongside new, introducing inefficiencies. Organizations see real benefit only when the introduction of agents removes steps, streamlines meetings, and directly reworks deliverables.
    6. Job Scope Is Redefined by AI Task Duration
      With autonomous AI agents able to sustain uninterrupted multi-hour or even multi-day project work, the scope of traditional jobs shifts. Instead of incremental disruption, the challenge is mapping how work is reviewed, delegated, and managed as longer, more complex AI-run processes become normal. This shift demands careful restructuring of oversight, review, and exception handling roles.
    7. Open Source and Open Weight Models Offer Resilience
      Recent months revealed that open-weight models, especially in specialized domains like code generation, are not just alternatives but essential resilience insurance. These models provide backup routes when primary providers withdraw access or change terms. Though not always production-ready for every use case, they are rapidly maturing and allow workflows to be rerouted to meet compliance, cost, or operational needs.
    8. Agentic AI Moves Risk from Theory to Operations
      The new wave of AI agents brings genuine operational exposure—models not only generate drafts but can execute real actions, from sending emails to manipulating files. Every deployment of an agentic workflow necessitates robust controls: ownership assignment, approval workflows, permission boundaries, logging, and rollback planning.


    Ten AI Moves Accelerating Organizational Impact

    Analysis of sector-leading organizations highlights ten moves shaping business outcomes with AI:

    1. Development of Portable AI Stacks
      Smart teams are preparing for disruption by architecting redundant, portable AI stacks that include primary, backup, and manual fallbacks—well-documented to mitigate sudden model or API outages.
    2. Workflow Mapping Preceding Tool Acquisition
      Rather than chasing every new model update, effective organizations thoroughly map processes and wins before onboarding new tools, producing more value by replicating and scaling proven ROI.
    3. Role-Specific AI Skill Packs
      Enterprise-wide prompt libraries are eclipsed by modular, skill-specific packs—custom markdown instructions aligned with business functions, vetted for rigor and impact.
    4. Automation of Recurring Reports and Briefings
      Routine reporting tasks, from meeting minute collation to executive brief preparations, are being fully automated. This shift reduces human hours spent on low-value aggregation and reallocates capacity to higher-value analysis and oversight.
    5. Transitioning Static Files to Living Work Artifacts
      Organizations are rapidly converting legacy PDFs, presentations, and spreadsheets into living documents—AI-native, continuously updated, and accessible across teams.
    6. Sandbox Testing of Agents Before Production Launch
      All agents undergo controlled “read-only” pilots with rigorous test cycles before being authorized to take action in live environments, materially reducing costly production incidents.
    7. Routing Tasks by Value, Risk, and Cost Efficiency
      Tasks are algorithmically matched with the most appropriate model (not always the “best” or priciest) based on sensitivity, regulatory risk, and value, supported by emerging tools for dynamic model routing.
    8. Measuring and Rewarding Accepted Outputs, Not Activity
      Business impact is assessed on accepted, utilized outputs—not AI activity logs, tokens used, or lines of code generated. Direct impact metrics drive process improvement.
    9. Red Teaming End-to-End Workflows
      Security, bias, and reliability checks target not just foundational models but the entire workflow: inputs, outputs, tool access, and integrations.
    10. Capturing Decision Reasoning for Long-Term Memory
      Decision logs, rationale, and first-person reasoning data are systematically harvested, creating enduring organizational memory—enabling future automation and helping teams understand why decisions were made.


    Ten Critical AI Questions for Leadership Decision-Making

    Operationalizing AI at organizational scale now centers on ten pivotal questions:

    1. What breaks if a core AI provider or model vanishes overnight?
    2. Is the current challenge a task, workflow, or operating model issue?
    3. What proprietary context differentiates company outputs from competitors'?
    4. What levels of access (read, write, modify, send) does each AI workflow have?
    5. Who owns and is accountable for each automated workflow?
    6. What is the defined cost ceiling for live AI operations?
    7. What constitutes an acceptable, business-ready output?
    8. What mechanisms handle and trace errors and inaccuracies?
    9. What evidence and audit records exist for AI-derived actions six months later?
    10. What legacy processes, meetings, and artifacts should be sunsetted if the AI workflow succeeds?


    Business Value: Converting Technical Truths to Tangible Gains

    The operational and strategic points highlighted are immediately actionable. By focusing on modular stacks, process mapping, output measurement, and context differentiation, organizations avoid continuity risk and gain the flexibility to reallocate resources where they drive results. Open-source adoption and agentic oversight strengthen resilience and compliance. Clear accountability, traceability, and well-defined acceptance metrics mean AI ROI is tied to business value, not technical activity.

    For organizations reshaping their architecture, talent, or processes around these specifics, operational security, cost control, and marketplace differentiation are not distant goals—they become visible milestones within a year. Adopting these practices positions a company not just to endure the volatility of the AI landscape, but to thrive amid it.


Topics Covered in This Episode:

  1. Eight Uncomfortable AI Truths for Leaders
  2. AI Model Risk: Single-Model Dependency
  3. Moats Shifting to AI Operating Layers
  4. Workflow Design vs. Prompting in AI
  5. Static Business Artifacts Becoming AI Debt
  6. Workflow Automation and AI Task Reshaping
  7. Open Weight Models as AI Resilience
  8. Agentic AI: Operational vs. Theoretical Risks
  9. Ten Smart AI Team Moves Explained
  10. Portable AI Stack Building Strategies
  11. Workflow Mapping Before AI Tool Adoption
  12. Building Role-Specific AI Skill Packs
  13. Automating Reports and Briefings with AI
  14. Converting Static Files to Living Artifacts
  15. AI Agent Deployment: Read-Only Mode Testing
  16. Routing Work by Model Value and Cost
  17. Measuring AI Output Quality, Not Activity
  18. Red Teaming AI Workflows Beyond Models
  19. Capturing AI Decision Reasoning for Business
  20. Ten Critical AI Questions for Leaders




Episode Transcript 


Jordan Wilson [00:00:15]:
Well, we made it to 800 episodes, and guess what? We're still here, and so is artificial intelligence. Turns out this AI thing really matters. So much so we're still doing the everyday AI podcast three and a half years after we started. But here's the other realities. No. There is no AI bubble. Yes. AI is smarter than all of us, and, no.

Jordan Wilson [00:00:43]:
AI has not taken 50% of what color jobs, regardless of what some CEO that's trying to drum up IPO hype tries to sell you. At everyday AI, I've spent the last 799 episodes trying to break down complex theories into digestible daily takeaways with my random rumblings sprinkled in, obviously. So for today's eight hundredth episode, we're going to continue the no nonsense trend as we give you eight uncomfortable AI truths, 10 moves smart leaders are making, and 10 questions every leader must ask. See, it's eight times 10 times 10. I'm not gonna give you 800 random AI facts or something like that, like I used to do when we were at episode 100. So and also sometimes these 100 ish episodes, right, 500, 600, 700, sometimes they go kinda long. I'm gonna try my hardest to make this one a shorter, value packed episode, so let's get straight into it. And if you are new here, well, welcome to everyday AI.

Jordan Wilson [00:01:46]:
My name is Jordan Wilson. We do this thing, well, every day. It's your daily unedited, unscripted livestream podcast and free daily newsletter to help everyday business leaders not just keep up with what's happening in the world of AI, but how we can make sense of it to grow our companies and careers. So starts here, but make sure to go to our website at youreverydayai.com. We're gonna be recapping the highlights from this show as well as all of the other AI news that's important that you need to know to get ahead. Alright. Let's get into it. And if you did miss our seventh hundredth episode, we're taking a kind of similar, take.

Jordan Wilson [00:02:20]:
But if you listen to this episode and you're like, wow. This was really helpful. I think you'll like our seven hundredth as well. Although some things might be a little bit out of date now even though it's only a 100 episodes ago, that's how quickly AI moves, but you might wanna go listen to it. So we went over seven ways AI is reshaping how we work, 10 AI workflows that actually deliver ROI, and 10 AI skills every professional needs in 2026. But let's start off this round with our eight uncomfortable AI truths. So number one, single model AI strategies are now continuity risks. And, I mean, the timing of this is obvious.

Jordan Wilson [00:02:59]:
Right? We saw Infropic roll out a very capable, and much hyped model in their fable five, which is of the Mythos five family. Right? It's Mythos five without guardrails. And I saw a lot of chatter online, you know, companies saying like, oh, you know, we're goodbye, ChatGPT. Goodbye Gemini. Goodbye Copilot. I'm trying to move all of their operations inside this new model because it is a step change or I could say it was or kind of is because it's currently no one in the world can use it. So that's why this is a single model strategy is a risk because we saw The US exports control action on Fable five. So for companies that rushed to try to move all of their operations into Fable five from anthropic, number one, okay.

Jordan Wilson [00:03:48]:
Scroogie McScrooge duck writing, blank checks. How can you afford that? It's a very expensive model. But number two, I think that shows where, everything is headed right now because you have to be thinking, across multiple models. And I don't know if this, kind of US government export, action could be a common thing. You know, the only thing with the new president's new executive order is a thirty day review period. So, this kind of export control action was not part of that initial, you know, executive order, but we've seen that anthropic and the US government have kind of had some ongoing beef for the last couple of months. So, this isn't shade at anthropic. This is just the reality of if you can trust your entire organization's AI strategy on a single model that the federal government literally at 5PM on a Friday can click, nope, and then the entire world loses access.

Jordan Wilson [00:04:44]:
Because if one model breaks your workflow I mean, procurement misses business continuity planning. Right? I I did a training for a company that was using Claude models, and they were trying to set up all of these complex workflows. And one thing I talked about them after the session I did is what's your backup plan? You have to be building these things modularly because if you spend so much time, you you know, especially weaving some core part of your business's operations into a single model, You could be screwed come Friday at 05:01PM when said model is done and gone, whether that's last for a couple of days, a couple of weeks, or indefinitely. Alright. So that is the uncomfortable AI truth number one. Truth number two, the moat has moved from models to operating layers, and that kind of goes with uncomfortable truth number one. And we actually did have a start here series, on this one. I believe it was trying to do my math.

Jordan Wilson [00:05:40]:
I think it was yesterday. Right? When we talked about AI super apps kinda being the, the new wave or the new trend on where work happens. But, I mean, codecs, you can't overlook that with now 5,000,000 weekly users. So, yes, in codex at least, you are still using GPT models, but you are using models, multiple of them in the same conversation with a unified memory, being able to use a a a browser, a file browser, an Internet browser, being able to control your computer, being able to, use your browser, all of those things. That's why I think the models are becoming less and less important as our, capabilities that we get through large language models become more and more agentic. Right? As they start using browsers and using computers and using terminals and using files in the same way that humans do, I do think that there is maybe less of an overall emphasis away from a single model. And that's why I do think when we talk about moats, when we're you know, which I think for, you know, business leaders that are working on the front end AI strategy before you, you know, say, hey. We're for this department, we're gonna go kick off a a one year, program with this big AI company.

Jordan Wilson [00:06:51]:
You have to understand what their moat is. Right? Like, are they developing a harness that is gonna be able to fully use any of these tools? Is it maybe better to look at a more open harness, and then you can plug and play different frontier models, at will. So I think that's a really important thing. You know, talking about things like, plugins, that bundle all of these different skills and apps together, from codecs, I think, is another important thing. Also, as we talk about context, connectors, all of those things, All of these things on the outside, the second layer, if you think of the model as the base layer, all of the things surrounding them, I think are gonna become more and more important. I'll probably do a maybe my final start here series. And if you've missed that, you know, make sure you go to starthereseries.com. Go listen to that.

Jordan Wilson [00:07:37]:
It's a, you know, an ongoing series. I think we're gonna cap it at about 30 episodes because that's probably enough. But I think it's been great for, yeah, you you know, beginners to AI leaders to tackle all of these ongoing changes. But I do think that there's maybe gonna become a certain cape a a certain capability gap, not where humans are, you know, not using the full capabilities of models, but actually where we well, for most of our day to day work, at least as it is today, our work will not require 95 of what these models are capable of. You know, we kind of saw this with Microsoft, looking at DeepSeek as a potential, you know, maybe cost alternative, to running the GPT and clawed models. And at first, I'm like, well, that doesn't make sense, at least now. But it might make sense in, six to nine months because I think at least until the work that knowledge workers are required to do changes drastically. Right? For the most part, I think in six to nine months, you know, some of the more basic I'm not saying, you you know, deep seek.

Jordan Wilson [00:08:38]:
Right? I'm not saying that model, but I do think some of the more basic models or even a company's second or third most powerful model, maybe for a period of a couple of years might be sufficient enough to do the majority of all this work, which is why I think, you know, yes. When we were talking about models capabilities from, you know, 2024 and 2025, the model really mattered. But I think it matters less and less today than it did. Truth number three, prompting is losing to workflow design. Alright. And if you haven't noticed this, again, this goes along with the push toward the super app and the the harnessing and all of those things. But, you know, even Google. Right? Google is pushing agents, scheduled tasks, sub agents, orchestration.

Jordan Wilson [00:09:22]:
Right? If if if you are still spending a majority of your time or your team is still spending the majority of your time prompting large language models versus reading and acting on their outputs, you are behind. Let me just say that very bluntly. Right? If you are still going in and typing things and, you know, your team has a prompt library, it's not like that anymore. Right? We have skills. That's a universal language that all large language models speak. For the most part, they all have a scheduled task or automations. So for the most part, humans should be working more on the front end, you know, making sure that, your context stays up to date, making sure that, you know, you are making sure whether you're working with markdown files, skills across your organization, making sure those stay up to date with your most up to date context. But then from there, you should really be spending.

Jordan Wilson [00:10:10]:
Right? Humans should be spending more of their time taste making. Right? If you have a project that you would normally go in and prompt a large language model for, right, and that project happens every Monday afternoon, instead of you prompting every single day at Monday afternoon, well, you should be looking at the deliverables that five different agents delivered and saying which one is best. And then Tuesday, after that deliverable is due, you pour in with new updated contacts and make sure that, hey. Next next Monday, you know, those five, you know, reports that five different agents are gonna take, a a a swing at, they actually have more better context and directions. Truth number four, static business artifacts are becoming organizational debt. And I I didn't say have become on this one. Okay. I'm saying they are becoming.

Jordan Wilson [00:10:58]:
Let me give you the quickest example. Again, did in a recent, episode on this, actually, when we talked about codex sites, so make sure to go back and listen to that one. But as the harness has become more popular, right, as things like codecs, cursor, right, clogged desktop, I think, has some catching up to do. We'll see, you know, if Google continues to invest in the anti gravity two point o app or not. Right? But as these harnesses become, number one, more capable, but more team friendly, I think the concept of, you know, emailing files back and forth or even dropping, you you you know, static files in a shared drive. I think it's pretty soon gonna become archaic. And I think our artifacts, first, need to become AI native. Yes.

Jordan Wilson [00:11:48]:
We need to figure out what that means for history and versioning and all of those things. But, you know, I think codec sites is maybe the first iteration. Maybe it's not the final or the best. Right? But I think that's a good example of the difference between how a static business artifact can actually become organizational debt. And I didn't think it at first, but I thought back throughout the the the course of my career. How much time that I spent either communicating about different file versions, trying to find old versions, you know, corrupted files. Right? All these things just tackling, communicating, finding, organizing around different files, sharing them. You know, you're waiting for approval on all these files because, oh, turns out they had v four and they actually needed v four final.

Jordan Wilson [00:12:36]:
Right? All these things. I think we have to look at what does an AI native deliverable look like. And maybe it's a codec site. Maybe it's a, you know, clawed live artifact. I don't know. But that is truth number four is we have to start thinking about AI native deliverables. Truth number five, AI can create more work if leaders don't remove old work. Alright? Let me tell you what that means.

Jordan Wilson [00:13:01]:
AI often just adds more and more. Right? If you're looking at text walls of text, that doesn't necessarily mean that your organization is winning back time. Right? When we talk about ROI, if all you're doing is continually experimenting with different AI with different large language models, or if your employees if your team members are just banking their save time, which I would argue the majority of companies that saw, you know, AI gains, in theory, but not on the books in 2024 and 2025. Yeah. That's just your, you know, employees having AI do all the work for them, doing the same amount of work that they were doing pre AI and then saying, well, look. Right? But on the flip side of that, AI, if you don't properly, hold it in, can actually create way more work if you don't redesign workflows to be AI native. So if your job description hasn't changed, if your SOPs haven't changed, especially since '20 late twenty twenty four, early twenty twenty five when models became more agentic, you had to do so because all you're technically doing, is you are just doing a disservice actually now that I think about it a little bit more. Right? If you are still producing artifacts for clients, industry reports, whatever the the novel work is that you're creating business value within your organization, If you're still doing it the old way, let's just say the twenty twenty two way with twenty twenty six tools, that's bad.

Jordan Wilson [00:14:33]:
Right? Yes. You are still creating more work, but one of the reasons you're creating more work is because you are creating all of these processes that in theory are not gonna be transferable once your company, your organization, or your artifacts, becomes AI native or with whatever your whether it's for a client, an internal team member, etcetera. Right? What about when their demands change? And I think the consumer demand will begin to change. Right? We're getting over with all this AI slot. But I think that demands are going to change. So if you are just trying to use as much AI as possible on old processes, old job descriptions, companies and and organizations that are set up in a pre AI way, I think ultimately, you are just setting yourself up for failure even if it seems like success in the short term. So real ROI starts when the AI actually goes in there and removes steps, removes meetings, and reworks what you're actually working on. Truth six, longer AI task completion will reshape jobs.

Jordan Wilson [00:15:40]:
Alright. So good example here, Infropic says autonomous task length has doubled roughly every four months. And as an example, Anthropic said that Claude authored over 80% of Anthropic's merge code in 2026. So what does this mean? I'm not sure yet. Right? This is one of those things that still kinda keeps me up at night. And I think we it's it's gonna take even people on the edge of AI. It's gonna take them a while to figure out what does the future job look like. Right? Not like, what happens when we have, you know, embodied AI and human you know, human AI robots doing all these things even before that.

Jordan Wilson [00:16:20]:
Right? What happens when autonomous AI, harnesses are doing all of our normal deliverables that we were doing, three years ago. What happens then? But I think that managers must start to redesign work around delegation, supervision, review, and exceptions. Right? We have to start redesigning old work processes because as these models can work, whether we're talking about in loops, in managed and controlled loops, or just working truly autonomously toward a goal. Right, as they start doing things that would normally take many, many hours or many, many days. Right? We don't know how jobs are going to look. And I'm not talking about job disruption. I'm not talking about job creation. I'm talking about the jobs that will still be here, which I think is the majority of jobs.

Jordan Wilson [00:17:11]:
How are they gonna look differently? But as you have these, models now that can run and do these long tasks, and they can work for eight, ten, twelve, fifteen, eighteen hours on a single type of project that would normally take a team. What does that mean for the future of jobs? I don't know yet, but it will reshape jobs. Truth seven, open weight models are becoming resilience insurance. Right? So I already talked as an example about the recent Microsoft news that said that they're looking at DeepSeq as a potential rollback rollback or different offering. Right? So they're looking at fine tuning a version of DeepSeq and maybe using it internally, offering it up, to customers within Copilot. We'll see. Right now, it's just reporting that's still shaking itself out. But the reality is open source and open weight models will soon become at least a resiliency insurance.

Jordan Wilson [00:18:09]:
Right? This is something I've even started to experiment with my with on my own. I will say this, though. Open weight models, they're great. If you think that you can slap, right, a general open weight model and, you you know, have it run, you know, locally on your machine, like, oh, why would I ever, you know, use Fable five, or why would I ever use g b d five six? I can have an open weight model running on my computer. No. That makes you look foolish, FYI. I'm just saying that. In that case, open weight models will still, you know, open open weight or open source models will still be multiple years behind in terms of what a consumer can go and run on hardware.

Jordan Wilson [00:18:53]:
So unless you're spending $30,000 on hardware, it is much more economically feasible, and just responsible for you to be paying $200 a month than just use the Frontier model. Right? But open weight models are becoming a great backup plan for specific use cases. So as an example, and this one's kind of recent as well, GLM five two max is now technically the number one model in the world for front end coding because the only model that's ahead of it, Fable five, is not technically available. So it is the number one available model. Think of that, an open source model. Alright. Again, you can't really run this locally unless you have a cluster of supercomputers. But, again, thinking about how this shapes the future of work and how this impacts decisions that your company makes.

Jordan Wilson [00:19:44]:
I mean, you have to look at what Microsoft is doing there. If Microsoft is thinking that, hey, we might be able to whether it's internally or for Copilot customers. If we put enough human effort into this, we could fine tune an open source model, and run it at production for even half of internal or external purposes. That's big. And I think that business leaders need to start looking at that. For most companies, it's not there yet, but you have to understand that open fallbacks matter because vendor changes are gonna happen. Policy prices access, I I think through the rest of 2026, we'll see what shakes out between the anthropic and White House, but you have to start thinking of these things. And you need to start routing workloads by sensitivity, volume, compliance, cost, and quality threshold.

Jordan Wilson [00:20:30]:
Truth number eight, agents make AI risk operational, not theoretical. Alright. And I talked about this a little bit more in our 2026 AI prediction and road map series, but as the default as the default models become more agentic, they're getting read write actions. Let's even take out the fact of running these, you know, desktop super apps that can access every single file in your computer. They can read, write, use the terminal, all those things. Even look inside something like ChatGPT. Even look inside something like, I don't know, Grok, perplexity computer. Even the, quote unquote, nontechnical AI chatbots are getting more agentic power.

Jordan Wilson [00:21:16]:
Right? Very small example. ChatGPT can send emails. Alright. Doesn't sound like a big thing, but ChatGPT has always had the ability to draft emails and to read your emails and to pull in all this context. But just as of a week ago, it can actually send an email. Do you know how much risk there is when you give any AI model the ability to send an email? Right? There's great promise, but there's also a great downside, if you don't have the right expert driven loops. Because every agent needs an owner, limits, approval, logs, and rollback plan. Alright.

Jordan Wilson [00:21:50]:
So those are our eight uncomfortable truths. Now let's get into the 10 AI moves that smart AI teams are making, and, yeah, we're gonna pick up the pace here. Alright. Move number one is they build portable AI stacks before disruption hits. So I'm talking about primary backup, cheap, private, and manual fallback packs paths. You these all have to be documented. Right? What happens when AI works? You have to be able to answer that first. And then what happens when that AI that works fails? Alright.

Jordan Wilson [00:22:21]:
So smart teams are making those moves already. People always are are so quick to say, once you find a win and you're like, oh my gosh. This is gonna save us 40% of staffing costs. So we're just gonna go ahead and cut 10% of staff. Call it a win. Right? No. Use those humans and start building these portable AI stacks because disruption will hit whether it's internal, external, from a policy standpoint, whatever. You need to be, kind of creating these redundancies.

Jordan Wilson [00:22:51]:
Move number two, they map workflows before buying more tools. Alright? I think so many organizations that I talk to, they find one success, and then they roll that success out across different teams. Right? Which is great. That's a great way to go and and do things. But then what happens is then they just say, okay. Then what's the next good tool that we could do this same process? Right? Instead of replicating that workflow mapping across the entire organization. Instead. Right? And I get it.

Jordan Wilson [00:23:24]:
It's hard because companies always wanna be chasing the shiniest, latest, greatest AI tool and model. Let me let me be very direct here. Even if you I'd say, if your organization unless you have a team of a 100 people whose only job it is is AI deployment in your organization, which is no one, they have no outside responsibilities or deliverables, Unless you have a 100 people that that's their only job, you and your company will get much more much more value if you literally stop. Don't try g p t five six. Don't try fable five one or whenever it comes out. Don't try Gemini three five pro. You will find more business value if you literally stop trying the next best model. And instead, look at your next or your last big win and then use the findings from that and map that out to everything within your organization.

Jordan Wilson [00:24:22]:
That's very hard to do. So instead, normally, you just go through this concept of we're just gonna map, or we're just gonna try to replicate our big win. And, oh, by the time we, you know, rolled out that one big win across all organizations, there's a new model. So we're just gonna go chase that. It's hard not to, but you're gonna get much more, ROI, by just mapping entire workflows before you just jump on to the next big model. Move three. Smart teams are building role specific AI packs. Alright.

Jordan Wilson [00:24:56]:
Luckily, Anthropic was really ahead of the game with these skills. And surprisingly, it did take OpenAI a little bit longer than I thought it would, to kind of, jump, I won't say, on the bandwagon, but to follow suit. Skills are extremely powerful. Right? This kind of takes away, a lot of the, grunt work of prompting. Right? So skills, I mean, you could make the argument skills are just a bunch of text prompts stacked together. Right? And they sometimes do some technical things behind the scenes. Right? But this is the basics of, well, prompt engineering. Right? The skills are essentially this, prepackaged markdown files that tell a large language model in a very specific way what to do, what not to do, but around different skill sets.

Jordan Wilson [00:25:44]:
So as an example, there's great skill packs for different type of marketing rules. Right? So if you're a marketer, you should probably at least start with a marketing skill even before you go in and prompt something or customize a workflow to work for how you need it. But this is huge. And, you know, skills are obviously transferable across different systems. Codecs just rolled out a bunch of, kind of, you know, plugins, I guess, is what they're saying because their version is more, using skills plus apps, which is great. Right? But every organization needs to be experimenting with skills, and you need to find a good skill that's already prebuilt for your organization. There's great open, open resources for skills because you can plug and play them into anywhere. Right.

Jordan Wilson [00:26:34]:
You need to start with a skill that is highly rated, vet it across your organization, test it in a sandbox way, make sure it works, customize it, and go on to the next. Alright. Move number four, smart teams are automating recurring reports and briefings. My gosh. If you are still having 20 people in this boring three hour meeting every single week and someone's still typing up to dos because no one has access to the AI note taker, oh, or can't use it for this meeting, and then you're preparing a deck for the meeting and a deck after the meeting. My gosh. No. Stop.

Jordan Wilson [00:27:12]:
You have to start automating all of those things. Because number one, the majority of the people in those meetings don't wanna be in there anyways. Right? The the majority of people, reading a a 30 page slide deck, they only care about one little slide anyways. Right? That's the reality. I think you have to start pulling in all of these things that are, spending that you are spending too much human capital on in a non AI native way, and you have to start blowing up those old processes. Alright. Move number five. Already kind of talked about this, but it's worth talking about again.

Jordan Wilson [00:27:48]:
Converting static files into living artifacts. So sorry. Spreadsheets, PDF, PowerPoints, all these things. You have to and and a good one too is skill skill libraries. Right? Your company, your organization needs to have living AI native documents. Old static files are going to slowly die. Probably not by this year. Right? It'll probably still be another year or two before they actually die.

Jordan Wilson [00:28:18]:
And until we see the rise. Right? Even look in the AI community, how popular things like markdown files are now. Right? We are going to get whether it's something like codec sites. I don't know what it is, but you have to start converting your static files right now into living artifacts. Move number six. Smart teams are testing agents in read only mode before giving them control. My gosh. I can't tell you the amount of times I've seen this.

Jordan Wilson [00:28:47]:
Right? Whether it's, you know, hearing about it on on, you know, conversations I don't have here on on, on the show, reading about it in news articles. The process of sandboxing an agent and putting an agent into production is pivotal. Right? I think a lot of teams think, okay. Well, let's sandbox the agent. Let's test it against the guardrails. Make sure it's safe. Make sure it doesn't drift. You know, they go through.

Jordan Wilson [00:29:19]:
They check their boxes. Right? That makes the compliance team happy. And They're like, look. We're we're super responsible. Right? And then they do this, you know, this is usually those, those companies that I think are a little too AI trigger happy. And then all of a sudden, they start setting these agents out. And well, they're like, oh, don't worry. Bill from IT.

Jordan Wilson [00:29:38]:
He's our human in the loop. You know what? I know I talk about this a lot, but the world would be a lot better place if no one ever came up with this human in the loop thing because it's absolutely garbage. Because in I I tell you, 90% of the time, the human that's in the loop, number one, is the wrong person. Number two, they're not driving or steering. They just look in there when it's it's it's more of, they're a crash scene investigator instead of an FAA flight operator. Right? Most of the time, a human in the loop is only, quote, unquote, involved when something goes wrong. So you have to test agents in read only mode before giving them control. Because as the models driving these agents become more sophisticated, so too does your process of the human or humans you have overseeing them.

Jordan Wilson [00:30:40]:
AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead. But you don't have ten hours a day to understand it all. That's what I do for you. But after 700 plus episodes of Everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in.

Jordan Wilson [00:31:24]:
Number one, go scroll back to the first one in episode six ninety one. Number two, tap the link in your show notes at any time for the start here series, or you can just go to starthereseries.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same. The start here series will slow down the pace of AI so you can get ahead. Move seven smart AI teams are making. They route work by value, risk, and cost. So we had an episode recently in the start here series talking about going from token maxing to token efficiency, and this part's huge. You should probably not, if I'm being honest. You should probably not be using Fable five for 99% of your tasks.

Jordan Wilson [00:32:16]:
You shouldn't be using GPT five five x high for 99% of your tasks. Right? It's crazy the amount of people that are just using the best model for everything. Blankets. That's not the way to do it. Right? I think most companies take these, I won't say their habits, but maybe best practices. Because for the most part, in 2023 through 2025, we were all living in this subsidized, AI nirvana, right, where it's like, oh my gosh. I can just use the best model for anything regardless. Alright? But times are changing.

Jordan Wilson [00:32:53]:
Times are changing. Token maxing has led to token efficiency. Right? We saw all these stories, all these huge organizations, you know, hit their AI spend limit or their token budget within three or four months, and now they have to change. And the way you change that is by routing the type of work to the proper model or models that it needs. That's why I think that model routers, mixture of model setups, I think there's gonna be some great AI startups. I've already seen a few launch in the last few months that do exactly this. Right? So you are not wasting, as an example, a billion tokens a day. It's not crazy for you know? I very easily go through 500,000,000 to a billion tokens a day sometimes.

Jordan Wilson [00:33:44]:
Me. Small business owner. Right? Luckily, I'm still heavily relying on the subsidized, you know, OpenAI $200, ChatGPT Proclaim. But subsidies are all going to eventually fall or diminish, so you have to start making these calls appropriately. And, you know, similarly, you do have these now very capable models like GLM five two, the Kimmy models. Right? We're not saying use them for every task, but instead of using, you know, one model for 50 different tasks within your organization, you might be using, you know, five different models for eight different tasks or something like that or 15 different models for three different tasks each. So you have to start routing them by value, risk, and cost now because if not, you are going to be in for a rude awakening. Because as an example, right, if anthropic's like, yeah.

Jordan Wilson [00:34:44]:
You know, here's enough Fable five. We'll see if the June 22 date gets extended or not. But, you know, it's it's very likely and and anthropic might go down this path of, you know, releasing these models just to all subscribers. People are gonna be reworking their workflows, then they get pulled, and you have to start paying API costs. So, you know, $200 becomes 14,000 give or take. So you can't keep doing that. You have to work smarter. Move number eight, I'm seeing smart teams make.

Jordan Wilson [00:35:16]:
Measure accepted output, not AI activity. Right? Lines of code is not return on investment. Right? Tokens burned doesn't mean a thing. Are your deliverables getting better? Are they moving the needle? Are your blog posts bringing in more human visitors? Right? All of these things that I think we rushed to start using AI for. Oh my gosh. Look at claw design so great now. I can go get out five x the number of proposals. Okay.

Jordan Wilson [00:35:50]:
Number one, did you? And number two, are you bringing in five x the amount of business, or are you just forcing AI slop down people's throats who are already being avalanched upon with just, you know, more and more content. Just because we can create more and more text, more and more photos, more and more videos, more and more audio, doesn't mean that we should be. I think human taste really matters here. So you have to look at accepted outputs. Right? Which of those five you you know, to pull from my random example, random rambling from earlier, you know, when my agent came back and gave me five versions of a certain report, which one was number one accepted by the expert driving that agent loop? And number two, which one ultimately won over my peers, won over the internal stakeholders, won over the external stakeholders as well? Which one got us more business? So you have to start measuring those things, not activity. Activity is not movement. Move number nine, they red team the workflow, not just the model. Okay? Yes.

Jordan Wilson [00:36:57]:
Large enterprises need to have red teams. You absolutely have to. Right? We saw this with the, you know, even though we don't have the full story yet, but we saw this with the, the Fable five rollouts. Right? Essentially, at least according to the White House and the federal government, you know, they said, if Robich shift a model that wasn't safe. So what happens if you put that into production and for whatever reason, one of your customers, one of your clients who maybe relies on your services or your outputs that you use somehow stumbles upon a piece of code that wasn't safe. And you're like, well, well, this model did it. Well, okay. Yes.

Jordan Wilson [00:37:35]:
We assume that labs like Anthropic and Microsoft and OpenAI and Google, have the best red team red teamers in the world. Right? It's it's it's those, those engineers that go in and make sure a new model is safe, and they try to break it internally, before releasing it. That's kinda like what red team is. That's red teaming for dummies one zero one. Right? But you have to have an internal team that does that as well, not just the model, but the harness. So in the same way, you know, if if if you say, hey. We're gonna roll codex out to 10,000 people. We're gonna roll roll out Claude Code.

Jordan Wilson [00:38:11]:
We're gonna roll out Cursor. Right? I think before, when you would roll these things out, it was usually developers using them to write code. But as we start producing nontechnical work documents as that becomes the norm, how are you red teaming that internally? That's what you have to do. Agent risk includes the inputs, the tools, permissions, actions, and the integrations. So you t your teams must test prompt injection, what happens when you use bad data, tool misuse, and also runaway costs. And move 10, smart teams are capturing decision reasoning in systems they own. Alright. Models are rented, but company reasoning history can become durable memory.

Jordan Wilson [00:38:52]:
I've been talking about this for years. One of the most important things your companies can do is to start collecting first company reasoning data. Right? FCRD. I've used the acronym once or twice before. So the simplest way to think about this, the old school transformer models needed structured data. Right? It gobbles up everything on the Internet. Models that reason. Models that think.

Jordan Wilson [00:39:22]:
That's today's models. Yes. They're still good. Yes. They still need structured data. Good data. Your company's good data. It also means how your company thinks and how your company makes decisions.

Jordan Wilson [00:39:35]:
Because the smartest thinking model in the world is only gonna be as helpful as the nuanced information that you give it. What do I mean by that? In the same way that you would teach and invest in a human over years to understand why you made this decision. Hey. Here's why I rejected this proposal, and here's why I crossed off this line item. Those things aren't always found in a spreadsheet. That's not always structured data. Right? These are the things that smart teams are capturing, putting them in documents and training models to work through. Alright.

Jordan Wilson [00:40:15]:
So now let's get to 10 questions every leader must ask. Alright? Question number one, what breaks if our primary model disappears tomorrow? Yeah. Maybe a little recency bias in putting this all together, but I think the whole situation with the anthropic fable five and the US government really sheds a much needed light on this AI race that has become a frenzy. What happens if that model is gone? You have to be able to identify the workflows, the customers, employees, and systems that are dependent on that one model. And what happens if it goes away? Alright? I'm not gonna answer these questions. These questions are for you and your team. Play this episode for them. Share this episode with them.

Jordan Wilson [00:41:06]:
You need to be talking about these. Question two. Are we solving a task workflow or operating model problem? Because tasks need prompts, workflows need systems, and operating models need redesign. So if you just have an AI summarizing meetings, that's a lot different from updating projects and owners. Because you need to be not just solving the task, but the bigger picture as well. Question three. What proprietary context makes our AI better? And this goes back to that first company reasoning data. Generic models just produce produce generic outputs.

Jordan Wilson [00:41:49]:
And it's the same thing even though these models are becoming smarter and smarter. If all you're doing is feeding your data to a world's smartest model in the correct way, that's not gonna help. I mean, it's gonna help a little bit, but that's not gonna be what ultimately separates you Because your competitors are gonna be doing the exact same thing. They're gonna be putting their best structured data with the best reasoning models or the best harness. Right? If competitors can match your outputs instantly, then the model was never your moat. Question number four, what can AI read, write, change, or send? You need to know all of those things. And there are very few yeah. Let me say this honestly.

Jordan Wilson [00:42:35]:
There are very few AI leaders out there that are in charge of AI rollout across their organization that can accurately answer these questions. And it's not their fault. It's because this space moves too quickly. So unless you have a team of people reading every single change log, right, or unless your whole team listens and reads the you know, if you listen to our newsletter every day, if you or sorry. If you listen to the podcast every day, read the newsletter, you're probably ahead of most teams. But still, even at that point, you have to know every single model, every single harness, every single connector, which one has read write access, which one requires approval, what permissions, right, which permissions are needed, you know, when using which AI super app, what can do what. You have to know those things. Question five, who owns every AI workflow? What one person? Asian crash is going to happen at your organization, in your industry, your competitors, etcetera.

Jordan Wilson [00:43:35]:
You have to know now. Number one, you have to map the workflows. But number two, you have to know who owns each workflow. And that person has to be able to answer those questions like question four. What can AI read, write, change, or send? And you have to know who owns that workflow. Question number six. What is the cost ceiling? The ceiling is important. We saw some stories.

Jordan Wilson [00:43:58]:
We don't know if they're actually true or not. There's one story that came out that one company, quote, unquote, accidentally ran up their monthly AI bill to $500,000,000 because they didn't put a spend cap on Claude. Was that true? I don't know. But what is the cost ceiling? You have to have those measures in place, not just how do we monitor what the cost is, but what cost should you be paying? Right? Should you be paying, I don't know, a million dollars a month to summarize emails? Maybe. I don't know. If if if your organization if you have hundreds of lawyers reading very important emails, maybe. Right? But you have to know what the cost ceiling is because agentic AI can spend money invisibly while appearing to be productive. Right? That's why you need to look at things like artificial analysis is cost per intelligence.

Jordan Wilson [00:44:52]:
Right? Because a lot of times, people just look at a single benchmark and be like, oh my gosh. This Fable five thing is great. Well, it what if it costs three times as much, as something like GPT five five? If you get the same exact output, but it costs three times as much. What's your cost ceiling? Question seven. What counts as a useful output? My gosh. I've seen people just really redefine and twist what a certain end artifact should be because an AI created it. Generated content means nothing until the business accepts it, uses it, and you can prove an ROI on it. So you need to say, what counts as a useful output? Where is your AI spend going? And is it producing an actual useful output? Question eight, what happens when the AI is wrong? K? Not when it goes out, but what happens when it's wrong? Do you have the right expert driven loop that can identify when an AI output is wrong? Hallucinations, yeah, they are not really a problem anymore, but not everyone is using the right model.

Jordan Wilson [00:46:00]:
So they obviously still exist. If you do the context engineering thing correctly, if you have expert driven loops, if you're using the right model for the right purpose and you've gone through the guard railing and the scoping and all those, let me just say it. Hallucinations are extremely rare if you do all those things, but most companies do not do those things. So you have to ask the question, what happens when the AI is wrong? Question number nine, What evidence could we show six months laters? Six months later. Yeah. Leaders, you need records of models, data, prompts, actions, approvals, and changes. Some, right, at some enterprise levels, like, if you're using JetGPT Enterprise, Cloud Enterprise, right, Copilot, they do have some useful metrics, I guess. Right? I like Microsoft's approach with their intra ID, and and a lot of the other companies are starting to adopt something similar as well.

Jordan Wilson [00:46:54]:
But every single piece of value that is created via an augmented relationship between a human and AI, It needs to be observable. It needs to be traceable. It needs to be auditable. Alright? You need to be able to go back six months later and say, oh my gosh. We ended up winning this huge RFP. Right? And we had this, you know, RFP agent cooking it all up. Let's go back. How do we do that? Who set that up? What models were we using? What was the workflow design process? And why did we get wildly different results from all these other RFPs when we were using the same process? You have to not only be able to go back and say, hey.

Jordan Wilson [00:47:37]:
This worked. Let's go back and, make sure that we apply some of our successes across the board. We have to also say, why did it work in some scenarios and not others? And then the last question, question 10, what should we stop doing if this works? And this is the big one. This is the thing about shifting human agency. And maybe I'll this is good to end this episode, episode 800 with a bigger question. Something maybe thought provoking because it's something I'm thinking about all the time. What happens in six months? What happens in nine months? What happens in one year? When these agents are even more and more capable. Go back and look at where we were a year ago and look at where we are now.

Jordan Wilson [00:48:20]:
It's scary. There is no wall. It is a ramp going straight up. If capabilities continue, if the harnesses continue, if these agents that run on your local machine twenty four seven and can produce artifacts in the same way that humans can and that that they are economically as valuable or more valuable than expert humans if all these things continue? What should we stop doing when this continues to work? Should we stop spending time as humans on the things that we've spent the majority of our career on? Should we stop? Let's just say let's just say you love creating PowerPoint presentations. Right? Let's just use that as a small example. You love getting the team in the room and, you you you know, brainstorming and, you you know, pulling in all this doc all these documents and creating a banger presentation, and you present it to the client, and they love it. What's that's your job. Right? If that's your job and you love it, but it's shown that an agent is better, what do you stop doing? I don't have the answer, but maybe on episode 900, we will.

Jordan Wilson [00:49:30]:
But that's a wrap for today celebrating our eight hundredth episode, eight uncomfortable truths, 10 moves smart teams are making, and 10 questions every AI leader must ask. I hope this was helpful. You know what? I'll say one thing that is helpful for me hearing from all of you. So if you didn't make it this far fifty minutes into this episode, reach out. Let me know what would you like to see in episode eight zero one. What would you like to see down the road? I know we normally, on Wednesdays, do our AI working Wednesday, so thanks for, you you know, letting us throw a little curveball in the rotation. We'll be back next Wednesday with our normal working Wednesday series. But thank you.

Jordan Wilson [00:50:12]:
Whether you've been around for one episode or all 800, I hope we can keep going to 900 a thousand, and beyond. So if this was helpful, do me a favor. Tell someone about it. I can only make it to 901,000 and more by giving you unbiased straight AI info when you tell people. So, please, number one, subscribe to the podcast if you haven't already on Spotify or Apple Podcasts, then make sure you go to your everydayai.com. Sign up for the free daily newsletter. Do me a favor. Drop me a line today.

Jordan Wilson [00:50:43]:
Drop me an email. I would love to hear, you know, whether it's if something's been helpful along the way, how long have you been listening? What would you like to see us do differently in the next 800 episodes? I work for you. So thank you for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI