Ep 718: Agent Risk, Security, and AI Sprawl in 2026: Why AI That Acts Changes Everything (Start Here Series Vol 9)

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


AI Agent Risk, Security, and Sprawl: Actionable Insights for the 2026 Business Landscape

In early 2026, business leaders face a fundamentally different landscape shaped by the emergence of advanced AI agents. The latest episode of the "Everyday AI" podcast delivers a rigorous walkthrough of how agentic AI transforms risk, security, and operational complexity—backed by specific real-world scenarios, practical frameworks, and concrete actions for organizations navigating this new era.

AI Agent Risk: From Chatbot Errors to Autonomous Actions

Between 2022 and 2024, AI risk was limited to misinformation, accidental data leaks, and occasional hallucinations from text-based systems. The transcript highlights how this risk profile expanded dramatically in mid-2025, when AI models gained the ability to act autonomously, connect with organizational systems, and execute actions—from modifying files to scheduling meetings—often without clear visibility from human overseers.

Today, AI risk is not hypothetical. Agentic models can move faster than any employee, create sub-agents autonomously, and spread within organizational data structures much like a digital virus. The capacity for silent, large-scale action is no longer just a future concern; it is reality. Business value and risk have never been so closely intertwined.

Agent Security: Practical Surfaces and Vulnerabilities

The transcript establishes a clear mental framework for agent risk, breaking it down into three actionable surfaces:

  1. Input Layer: Untrusted content—including prompt injections and hidden instructions—can trigger unintended agent actions. While outright malicious attacks require technical acumen, inadvertent copy-paste errors in everyday workflows open up similar vulnerabilities.

  2. Tool Layer: Every connector and permission expands the potential blast radius. The transition from stationary brains to proactive models with tool access means agents can now run code on machines, touch sensitive infrastructure, and cascade actions across systems.

  3. Action Layer: The biggest change is from outputs (text, reports) to actions (data modification, emails, purchases). Silent, unintended workflows can proliferate, leaving organizations exposed to invisible risks.

AI Sprawl: Dark AI and Observable Challenges

Understanding agent sprawl is essential for mitigating risk. The transcript categorizes three distinct forms:

  • Shadow AI: Unapproved or unknown use of AI tools by employees (e.g., personal ChatGPT accounts for work).

  • Agent Sprawl: Known but unmanageable proliferation of agents—approved tools whose full operational footprint cannot be tracked.

  • Dark Agent Sprawl: Unobserved agents, including malicious ones seeded intentionally (future malware, ransomware analogues), working inside company infrastructure.

Unchecked, these phenomena multiply risk and force companies to address identity, permissions, and ongoing policy compliance at a deeper level than traditional IT.

Enterprise Response: Strategies from Leading AI Labs

Major AI companies are actively responding to these new threats:

  • OpenAI: Implementing human approval workflows for agent actions, using command centers like Codex to monitor decisions.

  • Anthropic: Focusing on defense against prompt injection, ensuring browser agents operate in isolated virtual environments, and adopting domain allow-listing for safety.

  • Google: Project Mariner uses virtual machines for browser agent isolation, minimizing privilege escalation.

  • Microsoft: Copilot Studio scales governance, utilizing monitoring, logging, and identity frameworks tailored for enterprise contexts.

Leading labs recognize that rapid development must be balanced with robust security architecture. However, most organizations currently lack visibility and centralized control over their agent footprint, making sprawl difficult to detect until damage is done.

Actionable Monday Morning Playbook: Governing AI Agent Risk

Key insights from the transcript translate directly into pragmatic playbooks:

  • Bounded Autonomy: Begin with suggestion and proposal phases, progressing to human-approved limited execution. Avoid full autonomy until monitoring, traceability, and governance are established.

  • Least Privilege Principle: Start agents with read-only access. Restrict write permissions to narrowly defined, observable tasks.

  • Mandatory Human Approvals: Require explicit signoff for irreversible actions such as deletions, permission changes, and purchases to contain risk.

  • Decision Trace Logging: Establish tools to log agent tool calls, trace decision paths, and flag abnormal action patterns. This is foundational for post-hoc auditing and future agent ops teams (the transcript predicts agent operations will mimic DevOps structures in 2026).

Anticipated Trends: Supply Chain and Marketplace Risks

The transcript pinpoints several trends to monitor throughout 2026:

  • Browser agents will become major risk surfaces as more workflows migrate online.

  • Open source agent pilots carry heightened risk, especially as plug-ins and skill marketplaces expand. Malicious code has already been observed in popular tools.

  • Point-and-execute patterns (agents directed to URLs without human review) are spreading. This accelerates exposure if sites are compromised.

  • Production mindset should replace “side experiment” mentality. Winners will treat agents as central infrastructure with enterprise-grade controls and compliance.

Conclusion: Opportunity Hinges on Understanding Agent Risks

Organizations seeking to harness advanced AI agents must invest as much in risk management and governance as in skill acquisition and operational deployment. Specific, layered security practices, visibility into agent actions, and deliberate scaling are now table stakes for sustaining business value amid rapid technological change.

Business leaders are urged to move beyond optimism—balancing excitement for new capabilities with practical, strategic realism. Only through a detailed, hands-on approach can the full potential of agentic AI be realized without incurring unmanageable risk.


Topics Covered in This Episode:

  1. Evolution of AI Agent Risk (2022-2026)
  2. Agentic AI Security Threats Explained
  3. Types of Agent Sprawl & Dark AI
  4. Key Differences: Chatbot vs Agent Risk
  5. Major AI Labs' Risk Response Strategies
  6. Agent Actions, Outputs, and Enterprise Risk
  7. Monday Morning Playbook for AI Security
  8. Agent Skill Marketplace & Supply Chain Risk




Episode Transcript 



Jordan Wilson [00:00:15]:
There's always been AI risk. But in the early days of large language models and chatbots, that risk was like Bill getting something wrong in the blog post or Deborah putting up a hallucinated stat in the onboarding guide that was littered with m dashes and delves. But AI risk today is a legit different ball game than risk was three and a half years ago. I mean, heck, AI risk today is unrecognizable from what it was three and a half months ago. And that's not an exaggeration because after hearing for, like, five years that we're six months away from real AI agents, well, it finally happened. And it was actually this perfect storm of multiple events that led to an unexpected business scenarios that the corporate world now faces today. You either get on board with AI agents quickly or get left behind, but do it too quickly and you could go under. The risk, the security, and the sprawl are real.

Jordan Wilson [00:01:17]:
And so we're gonna tackle it all today on everyday AI, the start here series edition. Alright. Well, welcome to everyday AI. My name is Jordan Wilson, and, if you're new here, this thing's for you. It's your daily livestream podcast and free daily newsletter helping everyday business leaders like you and me keep up with all the news, like a thousand new agents a day. What do we do? What do we try? Well, tune in. I tell you and, help you make the right decisions to grow your company and your career. So after 700 plus episodes, I realized I couldn't answer the most common question that people had for me.

Jordan Wilson [00:01:50]:
Like, Jordan, you have a lot of episodes. Where do I start? Well, that's why I created the start here series. So the start here series is the essential podcast series to both learn the AI basics and to double down on your knowledge. So this is volume nine, and you can go listen to all of the, episodes in our start here series if you just go to starthereseries.com. So that will give you, free access to our inner circle community, and it'll put you right in the start here series space. So you can go, listen to all of them, watch them, read about them all in one place, interact with others who are doing the same. Alright. So, one other thing that you need to do before we get started my gosh, y'all.

Jordan Wilson [00:02:36]:
You you have to go listen to these. Episode seven twelve and seven thirteen. That is our 2026 AI prediction and road map series. That's like a culmination of a thousand hours of work over the years, or over the past year, to give you guys the blueprint for 2026. So make sure you go listen to those. Alright. If this sounds kinda like our last episode in the start here series, not exactly. So make sure you go listen to that one if you didn't already.

Jordan Wilson [00:03:03]:
So this is more of just the state of AI agents, where they are, what they are, how we should use them, should we use them. So make sure you go listen to, volume eight from yesterday. That's episode seven seventeen. But today we're here to talk about the other side, the risk, the security and the sprawl that's changing everything. So here's what we're going to be covering on today's show and well, why it urgently matters because AI models didn't just get smarter. They got hands. Right? The risk model changed when AI moved from generating tax like it was three and a half years ago to now it's taking real actions. And a lot of times, actions we're not aware of, and that's the scary part.

Jordan Wilson [00:03:47]:
And an agent connected to your email and calendar or, you know, your company's data can act fast, confident, and wrong in the same way that AI models hallucinated three years ago. Well, they can still hallucinate now. And this isn't a future concern. This is happening right now. So on today's show, we're gonna go over a simple mental model for why agent risk is fundamentally different from chatbot risk. We're gonna talk about what OpenAI, Google, Anthropic, and Microsoft are actually building to address this risk, and I'm gonna give you a practical Monday morning playbook that you can start using this week to address the risk security and sprawl. Sound good? Yeah. Sounds good to me.

Jordan Wilson [00:04:29]:
I'm I'm excited, and I wrote all this. Alright. So let's get a quick little catch up here. So probably if you're listening to the show, AI is not new to you. Right? But, let me just give everyone the the the briefer here. Right? So, essentially, from 2022 to mid twenty twenty three, early twenty twenty four, large language models were largely text systems. Right? With the the risk was just limited to misinformation and data leaks and, you know, using hallucinations and looking foolish. But that started to change, I'd say, in the mid twenty twenty five.

Jordan Wilson [00:05:06]:
So that's when the model started getting well exponentially more capable and agentic by default. That's the thing that people don't understand. Right. I've I've actually had two great conversations with, two really smart minds, kind of the head of, agents at CloudFlare and then, the head of Microsoft Research. And I learned a lot by talking to them both, you know, before and after the show as well. But one thing that kind of came true as well, no one really knows what constitutes an agent and what doesn't. But I think that we can all agree that even today's large language models, they're agentic by nature. Right? They well, a lot of us are giving them access to all of our data.

Jordan Wilson [00:05:48]:
And not in all cases do they have right access, but in many cases, they do. So when these models can think and act on their own and, you you know, spit up a virtual environment and a terminal and access your computer. Right? I keep saying this. Even right now, I have multiple agents running on my computer. Right? I have codex and Claude code going right now. I'll probably spin up, anti gravity, later. I always have agents running, and they have access, and I don't necessarily know what they're doing, you know, in between. I always go back and look when they're done.

Jordan Wilson [00:06:22]:
But that's what's really changed, in mid, you know, 2025. But here we are in early twenty twenty six, and this is where, now we're gonna start talking about and actually putting into practice all those buzzwords, right, that people just, you know, started chatting about in, like, 2024 to, you know, sound super smart, like, oh, governance and audit logs and, you know, isolation, protocols. Right? Like, all those well, okay. Well, good thing that everyone was talking about it, because, you know, we were trying to get our our buzzword bingo. Well, now we actually need it. Right? Now it's a time to talk about, ethics and govern governance and guardrails when it comes to, agentic capabilities. I like to put it like this. Keep it very simple.

Jordan Wilson [00:07:17]:
2022 and before, I'd say that AI was a dumb, stationary brain, but it was a brain. Right? So everyone was blown away. Like, oh my gosh. It can think. Well, it was dumb. It was stationary. Didn't move. Couldn't really think ahead.

Jordan Wilson [00:07:31]:
Right? And in 2023, well, it became a dumb stationary brain with tools. Right? That's when, you know, the early version of, Chad GPT plus, you know, GPT four. It had tools. Right? It could go on the Internet as an example. So that's when it, really started to open up what it could do or at least the information that it could access. In 2024, well, it was still a stationary brain with tools, but it went from a dumb brain to a smart brain. I would say in 2024 was the first time that we actually had smart models, because at the end of 2024, that's when we got reasoning models. And then in 2025, I think we still had smart brains with tools, but the difference now instead of it being a stationary brain, it was a proactive brain in 2025.

Jordan Wilson [00:08:17]:
It could go out at least especially at the end of 2025, the second and third quarter. It can make moves on its own. Right? Proactively could schedule things and or just, you know, they could start acting over long periods. And then what brings it to 2026 is, well, now we have that smart proactive brain with tools and arms. Right? So tools are cool. But when you have arms, you can actually use them in a real way. And I think that's where, you know, agents now have teeth when they are autonomous, proactive, and smart. So here's why agent risk, I think, feels different.

Jordan Wilson [00:09:03]:
Because, you know, like I said, with the chatbots, there was always risk, but was it really? I mean, yeah. Worst thing you can do is you get in trouble putting out something hallucinated, and you look foolish. Right? Does your company go under? Probably not. Right? Do you expose every single dark seat? Right? It's not that bad. I mean, it is, but it's not. But that the the the new agent layer is just a whole new type of risk that we're not even ready for. If I'm being honest, if you listen to the show, I say this a lot. I'm like, this year is going to be scary.

Jordan Wilson [00:09:37]:
People are not ready. Businesses are not ready. And I mean that. Right? That kind of business predicament that I talked about in the opening of the show there, that is real. If you don't run to use AI agents this year, you toast. You are literally toast. I don't care if you're a small business or a $20,000,000,000 revenue business. You're toast.

Jordan Wilson [00:10:05]:
Right? But if you sprint too quickly, you can go under because you could make devastating mistakes that would be highly improbable to recover from because AI agents can do things much worse than even the worst human. Right? Have you ever heard stories or maybe you've experienced this. Right? Kind of a rogue employee. Happens too often. Right? Someone's someone's bitter maybe about getting fired or not getting a promotion, and they do something absolutely crazy. Right? Maybe expose all the company secrets or, you know, release some files. I don't know. Okay.

Jordan Wilson [00:10:42]:
That's a human. That's a single human. And you can see that human. You have eyes on that human. Right? Bill in IT is watching that human. Agents are different. Agents move a 100, a thousand times faster than that one person, but you can't always see agents. And guess what? That one disgruntled employee, that's one.

Jordan Wilson [00:11:04]:
This isn't a video game. You can't respond 10 times agents can. Agents can spawn sub agents like that. And those sub agents can spawn like that. Right? So think of in the way, like a virus might spread across your computer or across the human body, right, and replicates and duplicates. It's the same thing with AI agents. A rogue employee can't do that, and that's why this risk is very different. It is very real.

Jordan Wilson [00:11:31]:
So I think that we spend too much time, thinking about the positives in the optimistic side of AI agents, which is great. Right? Yes. Oh, now all of a sudden I have, you know, 320, AI agents, you know, doing all my work for me around the clock. Oh, that's cool. Right? But what about the risk, the security, and the sprawl? And more teams are experimenting now than ever before because no one's got it all figured out. And that means more sprawl, and more exposure, especially if you don't have guardrails up. So here's essentially the three surfaces where agent risk actually lives. Number one is the input.

Jordan Wilson [00:12:12]:
That's kind of that untrusted content that can contain hidden instructions, agents, you know, treat like real commands. Alright? That I don't think is going to be as big of a deal. Right? Your inputs, don't get me wrong. Things can go wrong. Right? People are blindly copying and pasting, so you can copy and paste prompt injections, but prompt injections are the big thing, and we're gonna talk about that here in a second. But still inputs, they can be poisoned. Right? There's there's things that you might not know. In in the same way, a rogue human can create a lot of rogue AI agents that can create a lot of risk and a lot of sprawl that's gonna be uncontrollable.

Jordan Wilson [00:12:48]:
So, you know, inputs, I think you really only have to worry about it. You know, if someone really doesn't know what they're doing or if someone is trying to be malicious, which, again, those things are gonna come up. But think of that one bad employee, what they can now do if they know agentic AI. Right? Yeah. Talk about malware or spy spyware. Right? Ransomware that we you know, these companies that have paid, you know, millions of dollars, hundreds, tens of millions of dollars for for ransomware. It's gonna be way worse with agents. So the second layer is tools, and I think this is where it starts to get a little dicey in terms of, well, the capabilities are wild.

Jordan Wilson [00:13:23]:
Right? So every permission and connector that you add expands that blast radius essentially when something goes wrong. In the same way that I walked you guys through the, oh, it's a dumb stationary brain, but, hey, once that brain gets tools, okay, now it can start doing a lot of things, and the same thing on the agent gone wrong side. Right? If it was a dumb stationary, agent with no tools and no arms, it's like, alright. Well, have fun, buddy. Right? You're in a glass case of emotion just shaken up. When they have tools, that's where things go wrong. Right? When they have, access to your terminal, right, to your computer terminal, That's where things can go wrong when they can run code on a machine. That's where things can go very wrong.

Jordan Wilson [00:14:04]:
And then last but not least, and this is the big one. This is actions. And this is the biggest thing that, if I had to boil down the biggest change in risk and why it matters now more than ever, it's outputs to actions. Right? What we had to worry about a couple of years ago from AI in terms of risk was the output. Now we have to worry about the actions, but it's actions at scale and actions that we might not even necessarily be able to see. Right? Silent unintended workflows that you may not be tracing tracing. And really what this comes down to, well, it's this combination of increased capabilities from agents and moving in the shadows, and that's an enterprise nightmare. That is literally the formula for an enterprise nightmare.

Jordan Wilson [00:14:50]:
So some stats here for you. Right? So right now, 57% of employees, at least admit to using personal AI accounts for work. It's way more than that. Let's be honest. That's just the number that admit. A third admit to inputting sensitive data into unapproved tools. Right? So your shadow AI use case there. And here's the thing.

Jordan Wilson [00:15:11]:
You can't govern what you can't see. And most organizations can't see their agent footprint at scale. This is what I call the three types of dark AI. Alright? For the most part, you're not gonna find, you you know, the three dark types of or the three types of dark AI online is something that I've kind of, categorized them in. But I think it's really helpful, to get a better glimpse and a better understanding of the categories of risk. So number one, we all know this, shadow AI. Right? That's just essentially unapproved or unknown AI use. Right? If you've been using, you know, CHAD GPT on your personal computer, because Copilot is the AI that's approved, but you wanna use CHAD GPT and you copy and paste things over as an example.

Jordan Wilson [00:16:00]:
Right? That's shadow AI, but that's been around. Everyone knows that. Right? What you've maybe heard of, maybe not, it's the next kind of tier, and that's agent sprawl. But agent sprawl is known. Right? So that's essentially when you have approved agents, but you're not sure how to wrangle them or observe them all. You're like, oh, well, yeah. Bill gave us that, you know, agent to help with finances, but we're not really sure what it's doing. Right? We think it's doing good.

Jordan Wilson [00:16:32]:
Right? I checked the outputs, but I don't really know how it's getting there. That's that's that's the beginning of agent sprawl. But the thing is agent sprawl goes quickly. In the same way, a snowball at the top of the mountain might come rumbling down a thousand times the size. That is where we get into then dark agent sprawl. Alright? It sounds like a like a screen name for, for aim back in the back in the nineties. You guys remember that? AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead.

Jordan Wilson [00:17:16]:
But you don't have ten hours a day to understand it all. That's what I do for you. But after seven hundred plus episodes of Everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode six ninety one. Number two, tap the link in your show notes at any time for the start here series, or you can just go to starthereseries.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same.

Jordan Wilson [00:18:09]:
The Start Here series will slow down the pace of AI so you can get ahead. Right. Aim, you know, instant messenger. Right? Dark agent sprawl, twenty twenty six. That's gonna be my, my username on the, everyday AI inner circle community. But, essentially, that's where there's well, dark agent sprawl can be a couple of things because agent sprawl, it's a problem. But like I said, for the most part, that's a known risk. And you're like, yeah.

Jordan Wilson [00:18:40]:
We have these agents going all all over the place. We we have no guardrails. We have no, traceability, observability. Right? It's but but you know of it. Dark agents for all is agents you don't know about. Those are unapproved agents that are working in or on your company, and you can't observe them because you don't know about them. So this is well, this can be shadow AI, you know, gone agentic. Right? So people plugging in agents that aren't approved and company doesn't know, but there's also another kind.

Jordan Wilson [00:19:12]:
Right? And I think we're gonna see a lot of this, maybe not in 2026, but in 2027. That's going to be the equivalent of malware or spyware, but agents. Right? People sending, and seeding agents out specifically in the same way that they would for spyware, malware, ransomware to make money, right, to to extract, you know, value from businesses, and that's what's gonna happen. That's the next version of this. That's not necessarily dark agent sprawl, but there's two sides. Right? Dark agent sprawl can start out innocent enough. Right? Oh my gosh. You know, our company is not approving, you know, any of my AI agents.

Jordan Wilson [00:19:53]:
Well, okay. Whatever. I'm gonna go ahead and, you know, unleash, you know, Claude Code and get, you you know, 50 instances of Claude Code going up, and they're all gonna spin their agents. Well, one person might know, but the rest of the company is in the dark. And those agents could replicate, duplicate, and you can't observe them. But the other thing is, well, bad actors, dark agents, sprawl, that's the thing too. So why is this all happening now? All these risk and security concerns and the sprawl, why why now? Because like I've said, we've been six months away for five years. Right? We're six months away from great agents.

Jordan Wilson [00:20:28]:
But then it's all it almost seems like we didn't get a six month warning. They just popped up. Right? Like, December 2025. Right? Like, everyone was winding down for, you know, the the holidays, you know, especially here in The US. And then you go back to work in January, you're like, what the freak happened? Right? You're like it's like, woah. Woah. Wait. The the agents are actually here now.

Jordan Wilson [00:20:48]:
Where's the six month warning? Why don't we get get word of this in in June or July? But there's three reasons why. And it is literally the perfect storm, not just the ingredients happening, but at the exact right or wrong time. Alright. So number one, the number one element, that was mixed in the bowl and it's exploding is the reasoning threshold. Right? So improved reasonings from models like, you know, g p t five two from OpenAI, Gemini three one, their new model, very impressive. And, you know, Opus SONNET four six from INPROPIC. These are all built to be agent native. That's how they build them now.

Jordan Wilson [00:21:32]:
Right? They're not building models that are great at, you know, reading, writing, and comprehension first. No. What they're focusing on is is the harness and and the tool use. That's what's first. Right? Even, you know, Google's, you you know, their new model yesterday, everything that they highlighted was tool use. Right? And and and how to, you know, improve tool use in the scaffolding there has, you know, really helped them, improve their outputs. But that's that's the thing. These models, their reasoning ability is legit through the roof.

Jordan Wilson [00:22:04]:
They can plan steps ahead, self correct errors, and move beyond, reactive behavior into proactive. And that's kind of reduced agent reliability from around 50% to 90%, and that's perfect storm right there. Right? 50% coin flip. You're not doing that. Right? Would you hire an employee that has a 50% chance to fail right away? Probably not, but 90%. Okay. That's an a employee. Step two, computer use improvements.

Jordan Wilson [00:22:32]:
This seems small. This is big. Okay? The models have to be insanely smart. Right? They have to be able to think reason. They have to have genius level IQ, which is what today's models do. Right? If you'll get offline IQ tests, they're scoring in the genius level. Smarter than 99.9% of people. Right? But the computer use is important.

Jordan Wilson [00:22:50]:
Doesn't matter if it can't go and use a computer, but because at least for now, that's how businesses generate value. I I think eventually, value will be just generated agentically. Right? The web has gone agentic. Right? Google and Microsoft, announced support for essentially an MCP version of the web, where websites can talk to each other, agents can talk to each other. So we'll see how, you know, business value is extracted in the future because a lot of times it's been through the the website stack, the software stack, but, you know, who knows what'll be in the future. But right now, agents can use computers better than humans. Right? So, computer use means, these models can use a mouse computer. They can, click.

Jordan Wilson [00:23:30]:
They can use APIs. They can talk to each other. But the big thing is the computer use capability gap has essentially been solved. Right? So Claude Sonnet 4.6, which came out, well, earlier this week. Alright? So if you're listening to this in six months, right, it came out in, mid February. So imagine in six months, these these, computer use models are gonna be, insanely good. But it's the first, the first one that scored better than humans. Right? So 72.5, success rate on the OS world benchmark, and that's surpassing human performance for the first time and nearly quintupling the scores of twenty twenty four.

Jordan Wilson [00:24:09]:
That's the thing. And I think that's one of the reasons why, you know, even though the, maybe the foundation was there for AI agents, you know, a year, year and a half ago, they couldn't use the computer. It was so slow. Right? It was, you know, very, very, very slow computer vision. Right? Every time, if you wanted to click something, it would take three minutes. Okay? Now models can go as fast as humans or at least, you know, not just clicking and navigating interfaces, but just, a range of computer use tasks. And then the third thing that has created this perfect agentic store, and well, it's the context window and memory. Right? So being able to work on task for a long period of time, right, I think my, my record, so far, I think I got to about 10, overnight the other night on codex, which was really fun to do.

Jordan Wilson [00:25:01]:
Right? But it kept its memory persistent the whole time. Right? It didn't forget what I told it when I went to wake up. You know, when I woke up, yeah, I spent time reading through the chain of thought, and I'm like, oh, sweet. Not only did it remember anything, but I'm going through them, you know, tracing it, you know, making sure it kind of stayed within the confines of the instructions I gave it, and it did. But, you know, that's another reason why it's happening. And, y'all, let me tell you. If you haven't in the last, like, three, four weeks, if you haven't went out and used ChatGPT or, you know, OpenAI's codex, if you haven't used, you know, Claude Code or Claude Co work from Anthropic, if you haven't used, Anti Gravity, from Google, I'm not saying this to, like, be that guy. You you you're gonna get left behind.

Jordan Wilson [00:25:44]:
Right? And if you're listening to this show, I don't want you left behind. I'm actually thinking alright. I don't know this yet. It might might kind of, get a little bit of a vibe working focus, you you you know, here in the, first or second quarter. So, make sure if you're not already in our, inner circle community, get in there. Yeah. I'm gonna start put putting some stuff in there. Just keep keep, keep your eyes open for that.

Jordan Wilson [00:26:09]:
Alright. But let's keep going a little bit here as we, get gonna be wrapping up here in a minute. Let's talk a little bit about how the biggest AI companies are actually responding right now because the risk is well known because it has literally been an explosion, and it is going to get worse. So I wanna talk about a little bit about kind of the, big picture focus of the big four labs. So right now, OpenAI is taking more of a human approval, approach. Right? So Codecs is kind of their command center, to review agent decisions. Anthropic is really going on the defense against prompt injection, for browser agents, you know, with, making sure, that virtual machines are isolated and have minimal privileges, and going with domain allow list versus black list. Right? So everyone has a little bit of a different approach here.

Jordan Wilson [00:26:56]:
You know, Google as an example there, Project Mariner, browser agents run-in virtual machines as a safety, to isolate them from from anything else they can kinda get their hands on. And then Microsoft. Right? There's millions of ways that Microsoft is doing it and millions of ways that all these companies are doing it. I just kinda picked out, you know, different aspects to illustrate their different approach. You know, Microsoft as an example with Mike, their Copilot Studio, I mean, the governance is everywhere. Right? They're the sentinel monitoring, the, the purview logging, you know, the agent, intra ID. Right? So Microsoft, it is very enterprise. Google is as well.

Jordan Wilson [00:27:30]:
I was just giving an example of, you know, how their project barrier kind of runs in an isolated virtual machine. Because, essentially, I think from the people I've talked to, at least, three of those four companies, they know risk is part of it, right? Like there's like any lab, right? They call them frontier, you know, or AI labs for a reason. You run experiments and part of it is, well, you know, that things need to go wrong. Right? So these companies are intentionally trying to get all the risk and all the security nightmares and all the sprawl they can. Right? So then when they set a new model out in the wild when they're done with the red teaming, you know, they hope that they have a good idea. So the labs are working on this, but here's the reality. I think there's a pressure to allocate more resources to the development of models versus research and security. Right? I don't have that on, you know, privileged information.

Jordan Wilson [00:28:37]:
That's just from, you know, talking to a lot of people and reading a lot and just the, I think the reality of the world that we live in. And that's why a I sprawl is going to be hard to miss for most business leaders, because everything's unclear. Right? The tool like, do you know? Think of the the the AI model you use most. Do you know the tools it has? Right? Even if you think of, you know, Chad GPT, Gemini, Claude, do you know what tools the base models have? Probably not. You'd be surprised. Right? Like, I'm one of those weird guys that that reads, change logs, in model cards. Most people have no clue. Right? And and this is much worse than classic shadow IT because agents can act across systems, not just, you know, within the confine confines of, folders and files, which are very structured.

Jordan Wilson [00:29:32]:
The whole kind of, it's kinda like how people talk about hallucinations. Right? Oh, you know, they're not a bug. They're a feature. Right? What makes agents great is also the built in nightmare. So if you wanna have the good, you have to recognize the bad that they can do. Right? And that's because that's how they're made. They are made to go build their own path. They are made to blaze their own trail.

Jordan Wilson [00:30:00]:
Right? So by their very definition, agents aren't necessarily always good at staying within guardrails. Right? Because if they think that the destination that has been given to them, at times, right, depends on the model, the setup, all that. Right? But it's very common for an agent to hop over guardrails in order to accomplish a task, not knowing that that's a bad thing. Right? And every new skill that you add that that you give to, an agent or a model, every connector, every workflow, all that does is it expands the attack surface for everyone. And that's another realization that I think most people are going to struggle to grasp, right? Because you think that you're expanding the amount of work that you can get done. And are you? Sure. But I think you almost have to think of it from, like, I don't know, like an old school, like, military perspective. Right? It's your land.

Jordan Wilson [00:30:59]:
Right? And, hey, we're gonna go discover new land. We're gonna go out and, you know, win new deals. Right? So we're gonna expand our land. I don't know if any of you have ever played, like, you know, risk. I used to. I wish I still had time. Right? It's one of those games, you know, it's like, oh, okay. This is good to make sure my brain still works.

Jordan Wilson [00:31:16]:
But it's kinda like that. Right? Your agents, without you knowing though, are acquiring new territories. And that's great. And you think, oh my gosh. Look at these new skills and capabilities. Right? It's like, oh, it's like having a thousand new employees. Okay. Great.

Jordan Wilson [00:31:30]:
But your surface area for risk, attacks, and sprawl is also multiplying at the same time as your capability, but you're not focused on that. You're focused on, oh my gosh. Now all of a sudden I have a designer. Now all of a sudden I have a data analyst. Okay. Well, what's that data analyst doing with your data? Are you going to check? Right? Is it, using an in open source? Right. This is one thing I was kind of blown away by the other day. Gave gave, you know, wasn't anything bad.

Jordan Wilson [00:32:04]:
Right? Gave a model a very hard task. This is one of those that overnight, I came back and I realized, oh, it downloaded another large language model locally to do the task. Right? It did it locally. Right? So, kept all the data, kind of on on prem, on on my local machine. But what happens if in the future, other AI models are just gonna go use other AI models, but on the web? Right? They can. What if they do that without you telling? Or what if they go use a, a site that's unsecure? And and that and this is why I think, you know, OpenClaw, as amazing as it is, this is why it's also scary. Right? Because, you know, these, you you know, autonomous open source agents, the open source, if I'm being honest, and I'm sure, you know, I'm gonna catch some flack from this from some OG open source people. Right? OG open source, I'm not talking about you.

Jordan Wilson [00:33:01]:
Today's open source, it's taken a weird turn over the past, three to six months. Right? It's almost become this, you know, crypto infused wild wild web point for west. Right? It's not like traditional open source anymore. So when we talk about open source AI agents, right, so many of of what's so much of what's out there is risk. Right? It is, and it's unknown, and it can be scary if you don't know exactly what you're getting into. And I think this is why small agent pilots are becoming huge risk. Right? Because when you build your own agents without centralized visibility or governance, that's where you get in trouble. So many people when they're using their own agents, it's for them.

Jordan Wilson [00:33:43]:
Right? But it can be go going and doing work for the whole department, the whole company, or accessing data from the whole company. But if one person has eyes on it, and if there's not central, organization, that's where you run into something. So here's the playbook. Right? If you're feeling a little scared, you know what? Maybe that was my intent. I got you all pumped up and, you know, excited with yesterday's show talking about, you know, the new capabilities of agents, but you gotta get it in check. Right? You have to balance that that optimism with a little bit of realism, but here is the Monday morning playbook for how to deal with the risk, the attacks, and the sprawl. So like what we talked about yesterday, touched on it briefly, but you need to start with bounded autonomy. Right? When you think about autonomous agents, it's not, you know, zero to a 100.

Jordan Wilson [00:34:35]:
It's bounded autonomy. What that means is you start with suggesting, then propose, then approve, then limited, execution. Right? So many people go to full execution first. It's baby steps. Right? A human being doesn't go from womb to sprint. Right? You go from womb to, you know, being, being on your stomach, to sitting, to crawling, to walking, to right. Your agents have to go the same way. Even if you know, out of the box that the agent can sprint, you're not prepared for that agent to sprint.

Jordan Wilson [00:35:13]:
Part of it is for the humans. Right? We think it's more for, understanding the the agent's abilities. It's not. Right? This is to make sure that you don't have a bunch of lazy human in the loop, and instead, you have proactive expert driven loops. So start with least privilege by default. Read only first. Write access only for narrow defined task. Right? Especially in enterprise environments, you don't wanna send out a bunch of read only general agents.

Jordan Wilson [00:35:42]:
Start with read only. First, observe, improve your loop. Then you go to, limiting, limited execution, and then you can finally get it to, writing for narrow tasks. And then last, in your Monday morning playbook is you need to require human approvals for irreversible actions, like sends, deletes, purchases, and permission changes. These are all things agents can do. Right? Agents can literally do anything. Agents, right. And this is going to get even harder to manage when we start to see true agentic commerce.

Jordan Wilson [00:36:23]:
Right? And I'm not just saying, oh, you know, my agent is gonna go buy something, you know, from, the Amazon agent. That's not what I'm saying. Right? There's going to be agent, bartering. There's going to be agent, you know, liaisons. Right? So, you have to work on human approvals now. Right? It's almost like you have to understand that agents are capable of, you know, a 10 and you have to give them a a two because we as humans need to first learn to adapt our behavior, before just giving in to agents. And then you need to build governance before you scale. And I think that's, like, just what I was getting at.

Jordan Wilson [00:37:08]:
We get all excited for what we know it can do. And then we're like, wait. We gotta do this boring stuff first? Absolutely. Because if you don't do the boring stuff first, if you don't go through the sitting on your tummy, right, sitting up, crawling, walking, then you can't expect to understand, the path of the sprinting agents. And every agent run needs a decision trace that you can expect after the fact. So you have to have be able to log all your tool calls, capture your decision traces, and monitor for abnormal action patterns. This is one of the reasons why I said, I think we're gonna see in 2026, it's gonna be extremely common to have agent op teams. Right? In the same reason or sorry.

Jordan Wilson [00:37:51]:
In the same vein, you know, dev op teams are very common. Right? We're gonna have agent op teams for this exact reason. So here's what to watch for for the rest of 2026, when it comes to agent risk, security, and sprawl. Browser agents are gonna become a mainstream risk surface. I think we know that. I think we're gonna see a major open claw type. I'm not saying open claw themself, but I think we're gonna see a major, open source agent crash. I talked about that in the 2026 AI prediction and roadmap series.

Jordan Wilson [00:38:20]:
I actually think it's gonna come from right? There's this new trend, like, trend over the last, like, week or two where everyone is just sending pointing their agents to a, a website. Right? When they want their agent to do something, they're not even saying, alright. Let me, let me, the human, understand this. Let me write it out. Let me apply this to my business logic. Let me apply this with my guard. Nope. They're just saying pointing their agent.

Jordan Wilson [00:38:45]:
Hey, agent. Here's this video. Go watch it and do it. Hey, agent. Here's this website with a whole bunch of cool stuff. Just go do it. You know, there's this whole concept now going around of, you know, point your agent to a URL. Okay.

Jordan Wilson [00:38:59]:
Have fun with that. Don't do that. Right? Don't. Right? Especially if that agent has its own dedicated machine. Right? And I'm saying, if you're tinkering around, right, it's it's it's your personal computer. You're the business owner. If you wanna take that risk, that's fine. Right? Especially if you're you're you're tinkering around, that's fine.

Jordan Wilson [00:39:17]:
But don't do that in an enterprise. And, unfortunately, people are doing that, and that's a bad thing. Because guess what? A lot of those sites, quote unquote, oh, you know, someone puts up a huge tutorial. Oh, oh, okay. They're putting it up on their personal website. You know how easy it is to to hack, to fish any of these websites? Okay. You see let me give you an example. Something goes viral on Twitter.

Jordan Wilson [00:39:41]:
It's some, you know, literally, some 20 year old kid, you know, put a great, you you you know, guide up on the website. Cool. And, you know, everyone's just like, hey, agent. Go read that. Go read that. And then you have millions of agents going to just read this, you know, this nice kid, put up a cool guide. Okay. Guess what's gonna happen? Someone's gonna see that.

Jordan Wilson [00:39:59]:
They're gonna hack that site, and they're gonna inject malicious code in there that no one's gonna see. That's what's gonna happen. Alright. Next, the agent skill marketplace is going to expand, the supply chain risk. So you inherit what you plug in, and we've already seen that with, some skills and plug ins for agents. Some of the top ones were found to have, you know, malicious code in them. You need to, we're gonna see identity and permissions becoming a board level compliance requirement, not a nice to have. And last but not least, the winners are gonna treat agents like production software, not side experiments.

Jordan Wilson [00:40:32]:
Alright. That's it. That's a wrap. Volume nine of the start here series. I hope this helpful. I hope this is helpful, and I hope you understand a little bit more the risk security and the sprawl, that you need to be aware of and how this having AI that acts, having AI, right? That has a smart brain. It's proactive. It has tools and it has arms.

Jordan Wilson [00:41:00]:
That changes everything, right? This perfect storm that has happened in the last thirty days. We've seen, I've said this before, we've seen more both on the opportunity and the risk side. We've seen more in the last thirty days than we've seen in the last, probably two years. All right. So I wanted to take a moment for the start here series to address this. And I hope this helps you not only understand the risk, but if you want to take advantage of the opportunities, you can't do it without knowing the risks. So now you do. All right.

Jordan Wilson [00:41:36]:
So I hope this was helpful. So if so, there's already eight other, in our start here series that you should go check out. So like I said, please go to starthereseries.com. That is gonna give you free access to our inner circle community. Once you sign up, you're gonna be, just straight up loving it, hopefully, in our start here series space and connecting, with now more than a thousand people, in our, inner circle community. So thank you for tuning in. Hope to see you back for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI