Episode Categories:
Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
Managing the AI Capability Gap: Why AI Outpaces Most Companies and How to Catch Up
As artificial intelligence accelerates past the point of matching and exceeding human experts in most knowledge work tasks, businesses face an immediate and measurable challenge: the AI capability gap. Recent events and data illustrate that the essential bottleneck is not technology, but organizational readiness. Here's a focused examination of what drives this gap, how leading firms are addressing it, and the exact operational benchmarks businesses need to adopt to catch up.
AI Capability: Benchmark Data Shows AI Already Outperforms Human Experts
Frontier AI models, when used correctly, now match or exceed professional output in most knowledge work assignments. The GDP Val benchmark—run by OpenAI—evaluates AI performance on real business deliverables across 44 high GDP occupations. In the most recent assessment, GPT-4.5 matched or exceeded industry professionals in 83% of these blind, head-to-head comparisons, a dramatic rise from 47% just five months prior. This leap demonstrates that AI capability is advancing on a timeline measured in months, not years, and consistently outpaces human domain experts in complex workflow outputs such as research, pitch decks, spreadsheets, and other economically valuable deliverables.
AI Adoption Study: Meaningful Business Value Remains Rare Despite Tool Proliferation
Despite widespread claims of AI adoption—88% of organizations cite using AI—the majority are not seeing measurable business value. Only 6% report significant profits or operational improvements, according to a recent McKinsey study. The breakdown is primarily organizational: most firms limit applications to superficial tasks instead of deploying AI as a core workflow engine for defined, strategic processes. This pattern appears across industries, regardless of company size.
AI Capability Gap: Quantified by Task Automation and Workplace Application
Anthropic's labor data study provides granular evidence of the AI capability gap by mapping 20,000 public work tasks across 800 occupations to millions of real-world AI interactions. While current AI models can, in theory, automate up to 94% of tasks in computing and mathematics, observed usage is only 33%. In domains like management, business, finance, and legal, theoretical automation coverage exceeds 80%, yet observed application is under 20%. This quantifiable gap persists even in roles generally perceived as the lowest-hanging fruit for automation, such as software engineering and coding.
Moreover, workers most exposed to AI tend to earn 47% more and are four times likelier to have graduate degrees, defying earlier assumptions that AI would primarily affect junior or manual roles. Physical and manual occupations, for now, are least exposed.
Causes of the AI Capability Gap: Five Core Drivers
AI Model Capabilities: The functional scope of top models is no longer easily tracked. Recent months have yielded advances—which were previously considered science fiction—making it hard for even technical professionals to keep up.
Business Leader Knowledge: Most decision-makers cannot accurately assess current AI capabilities, partly due to the overwhelming pace of technical updates.
AI Literacy: Foundational understanding is lagging. Many users gravitate toward basic feature usage while skipping essential groundwork needed for strategic application.
Human Skill Set: Applying AI requires both foundational literacy and the ability to reimagine workflow processes—a combination that eludes standard upskilling efforts.
Tool Access: Not every role has immediate access to cutting-edge models or clear policies on usage, compounding the adoption challenge.
Recursive Self-Improvement: The Engine Behind the Fast-Paced Change
A pivotal shift in 2025 saw AI labs incorporating recursive self-improvement into their largest models. These systems now write and optimize their own code, generate new features, and build solutions autonomously, dramatically compressing the technology development cycle. The leading labs—OpenAI, Anthropic, Google—have acknowledged that flagship models are being continually improved internally by themselves, not just by human engineers. This technical leap means model capabilities now accelerate far faster than traditional enterprise learning or change management cycles.
Business Process Redesign: What Top Performers Are Doing Differently
Top-performing companies are now allocating over 20% of their digital budgets specifically to rebuilding knowledge workflows for AI integration. They are not simply layering AI on existing routines but are rebuilding workflows, separating tasks into risk tiers and aligning human oversight accordingly. These companies measure operational reliability—specifically, the percentage of AI-assisted workflow steps accepted without rework or incident—and use this as the baseline for scaling automation.
This approach eschews annual or quarterly planning cycles; instead, these organizations track, iterate, and retrain monthly to keep pace with new model capabilities. By continuously measuring how many workflow steps can be automated with the latest models and grouping tasks by risk level, they are able to manage risk and operationalize at speed.
The Key Metric: Track AI-Assisted Workflow Acceptance Rates
The most actionable starting point for managing the AI capability gap is to measure, at least monthly, the percentage of AI-assisted workflow steps that are accepted without rework or incident. This metric should be split by risk level: low-risk drafting tasks versus high-stakes financial or legal outputs. Among organizations that have rescaled knowledge processes, anywhere from 30% to 60% of work can be automated in a low or medium-risk tier, depending on function and sector.
Organizations that fail to adopt this pace and operational insight risk falling behind rapidly as AI-adaptive competitors capture productivity and profit advantages not possible with outdated processes.
Closing the AI Capability Gap: A Call for Continuous, Measurable Change
AI capability has entered a phase where quarterly or annual updates are insufficient to remain competitive. The gap between what AI can do and what most businesses apply it for is rapidly expanding. The most successful organizations are those that audit their workflows with the most advanced models, invest heavily in redesign, and relentlessly track real-world automation rates.
The imperative is clear: Inaction or complacency in facing the AI capability gap today will result in missed profit opportunities and eventual competitive stagnation. Only those who adapt their organizations in step with technical advances will secure sustained business value.
Topics Covered in This Episode:
- AI Capability Gap Definition & Urgency
- Frontier AI vs. Human Expert Benchmarks
- Anthropic AI Knowledge Worker Usage Study
- Top 6% AI Company Adoption Strategies
- Five Causes of the AI Capability Gap
- Recursive Self-Improvement in AI Models
- AI Adoption vs. Organizational Workflow Design
- Closing the AI Capability Gap with Metrics
- AI Automation in Professional Knowledge Work
- Managing AI Risk Tiers & Process Redesign
Episode Transcript
Jordan Wilson [00:00:16]:
There's been a lot of talk lately about these AI models that are scary good. And even more chatter about how many of the brightest minds in AI and all the tech AI CEOs feel that we've already reached artificial general intelligence and are racing toward superintelligence. But what does that matter to your your company? Because right now, those are kind of just theoretical issues. It's kinda like when you see, an email on a Friday and you look at it and you're like, yeah. That's a Monday problem. But do you know what's scarier than some private AI model that's scary good or artificial superintelligence? The AI capability gap. That's the scariest of all. Why? Because it's very real.
Jordan Wilson [00:01:02]:
And unlike that Friday afternoon email that you're maybe going to tackle on Monday, the AI capability gap is a today problem. And if you don't address it head on, it's going to slowly stall your company's growth. And if your company isn't growing, the AI capability gap will undoubtedly stop your company dead in its tracks. And this is not an exaggeration, and the timing here is specifically urgent. So we're gonna hit rewind today and explain the capability gap in AI and why AI is racing ahead of your company. And we're gonna show you how to slow it down a bit and catch it. Alright. You ready? This is part of our start here series on everyday AI.
Jordan Wilson [00:01:51]:
Let's get into it. So if you are new here, well, let's just get to the big picture. AI capability right now is outpacing business adoption, and it has changed recently. And the pace is too fast for any of us. Right now, Frontier AI. And when we talk about actual capabilities, this is not an exaggeration. Frontier AI models when used correctly, and that's the big asterisk here, they now match or exceed human professionals on most defined knowledge work tasks. And there's a massive gap that persists between what AI could, in theory, automate today and what organizations are actually using it for, which is often just topical.
Jordan Wilson [00:02:34]:
Right. There's a McKinsey study that said only about 6% of organizations are generating meaningful profits despite widespread AI adoption, and the bottleneck is not actually the AI models. You could have made that argument maybe mid twenty twenty five, but not anymore in 2026. It's actually about organizational adoption, workflow design, and training and education. So here's what we're gonna tackle on today's show. So stick with me for the next I'm gonna try to make this one twenty five minutes. We'll see. Alright.
Jordan Wilson [00:03:08]:
Stick with me for the next twenty two ish minutes, and you're gonna learn the benchmark proving that AI outperforms human experts, and it's not even close. You're gonna know more about Anthropic's new ish study revealing most knowledge workers barely touch AI's true automation potential. You're gonna learn why the AI capability gap only started to show itself over the past few months. It's actually kind of new. And what the top six percent of AI performing companies actually do that everyone else completely ignores. Alright. Welcome to Everyday AI, and this is our start here series. This is the essential podcast series to both learn the AI basics and for AI experts to double down.
Jordan Wilson [00:03:48]:
Alright? I started this thing because after 750 plus episodes, everyone always said, Jordan, great podcast. Right? When they found it, they're like, where do I start? I have no clue. And I'm like, I also have no clue. So That's why I started the start here series. It's best if you are brand new here, start an order. Alright. These are faster, you know, podcasts, usually about twenty five to thirty five minutes. But if you listen on two x, they're even half that.
Jordan Wilson [00:04:11]:
Right? But this is a great way to listen from one to now volume 19. And then also make sure to go to starthereseries.com. That's going to give you free access to our exclusive inner circle community. You can't find access right now. Any A gap report card. Alright. So, little bit more on that at the end. But trust me, you are going to want to repost today's episode.
Jordan Wilson [00:04:55]:
I'm just saying. Alright. And if you miss our last start here series in volume 18, and this was episode seven fifty, we went over the vibe coding boom, why vibe coding isn't going away and how it's both good and bad. So make sure to go check that one out. But now let's talk about how you can actually manage the AI capability gap and why AI is more than ready and companies are not. Alright. And I'm not the only one talking about this. I've been talking about this now since, I think, late twenty twenty five, but a lot of super smart people are starting to dive in deep on the AI capability gap.
Jordan Wilson [00:05:32]:
I really like what Jack Clark said. He is, the anthropic, cofounder of Anthropic, and here's what he said in a post a couple of weeks ago. He said most of AI progress has this flavor. If you have a bit of intellectual curiosity and some time, you can very quickly shock yourself with how amazingly capable modern AI systems are. But you need to have that magic combination of time and curiosity, and otherwise, you're going to consume AI like most people do. As a passive viewer of some unremarkable synthetic slop content, or at best, just asking your LLM of choice how to roast a turkey and keep it moist, or Tony box lights spinning but not playing music, What do I do? And all the amazing advancements are mostly hidden from you. Alright. So that's what Jack Clark said in a, kind of a viral, post that he had a couple of weeks ago, on x.
Jordan Wilson [00:06:30]:
And I think this is very telling. And the two things that he talked about, I think are very true. Alright. To start, well, first to even realize the AI capability gap, let alone close it. You have to be extremely curious. And the reason why I say that is because if you work how you've been working for the last few decades, you're not going to discover or your or your team, your organization is not truly gonna discover that capability gap because AI is not meant for the average knowledge worker. Right? That's why I talk about all the time, how I hate absolutely hate the concept of upskilling because AI is not something that you sprinkle on top of a seasoned knowledge worker. Right? Like myself, I've been working full time for twenty ish years.
Jordan Wilson [00:07:18]:
Right? I can't just sprinkle AI on the top. I have to I have to unlearn. Right? So you do have to have the time and you have to be curious on, hey. What does my role look like if I completely started over? And if I kind of forgot everything that got me to the point I am in my career, and that's what you have to do. And you also have to have a lot of time, and you have to devote a lot of time to it. Olivia Moore from a sixteen z, talked about she said, OpenAI dropped a state of enterprise reports across a million customers. This was a couple of months ago, but her response, she said, the gulf between AI power users and everyone else is wide. The ninety fifth percentile user spends 6¢, six times more messages than the median with coding, writing, and analysis showing the biggest gaps.
Jordan Wilson [00:08:00]:
So, yeah, those people that are power users, well, because they actually understand the capability gap, they're using AI all the time. This is why myself and I'm not saying this is, like, some weird flex. I'm saying this because I'm part of this, you know, power user group. When I say that I use AI for ten to twelve hours a day, right, that's how much I'm working. You know, it's not like I'm working sixteen hours a day. I am only working in AI every single day. No matter where I am, I'm using AI from beginning to end in literally every single step in between, and it has completely reshaped my workflow. Kevin Roose, New York Times columnist said this, and I really like how he put this.
Jordan Wilson [00:08:42]:
He said, I follow AI adoption pretty closely, and I have never seen such a yawning inside outside gap. People in SF, San Francisco are putting multi agent cloud storms in charge of their lives, consulting chatbots before every decision, wire heading to a degree only sci fi writers writers dared to imagine. People everywhere are still trying to get approval to use Copilot in Teams if they're using AI at all. It's possible that the early adopter bubble I'm in has always been this intense, but there seems to be a cultural kick takeoff happening in addition to the technical one, not ideal. And I'll say from my experience, I think there's five main reasons that the AI capability gap has really been exposed in 2026. And number one, the models of what they can actually do. Right? I'm just gonna roll through these things really quickly. So number one, what AI models can actually do.
Jordan Wilson [00:09:40]:
Number two, what business leaders think AI can do. Number three, AI literacy. Number four, human skill set. And number five, AI access. Now let me break those five things down a little bit more in-depth. So number one, what AI models can actually do. A year ago right? So I don't know when you're listening to this start here series episode. You know, could be in April 2026.
Jordan Wilson [00:10:02]:
You might be listening to it at the 2026. I don't know. Right. But a year ago, in the 2025, you could, for the most part, understand what AI models are capable of. Today, you absolutely can't. And it is literally getting to the point of science fiction. Right? All of these math problems and science problems that have been plaguing researchers for decades are now being solved. I look at them and I have no clue.
Jordan Wilson [00:10:27]:
Right? I I read the papers and I'm like, okay. I don't have any clue. I can follow and understand it a year ago. What AI models can actually do today is mind boggling. Number two, what business leaders think AI can do. I get to talk to a lot of smart people. AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead.
Jordan Wilson [00:11:00]:
But you don't have ten hours a day to understand it all. That's what I do for you. But after seven hundred plus episodes of Everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode six ninety one. Number two, tap the link in your show notes at any time for the start here series, or you can just go to starthereseries.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same.
Jordan Wilson [00:11:53]:
The Start Here series will slow down the pace of AI so you can get ahead. It's re really right. I've had hundreds of guests on this show. I'll say maybe a handful that I've talked to and I'm like, yes, this business leader fully understands what AI can do for the most part. And that's not their fault necessarily because most people in their role, they have to be an expert at one application. Right? You can't be in like, very few business leaders know what AI can do. I feel fairly confident that I'm in that category, but only because this is all I do. If I had another job, right, and I was just an AI champion at a company, you can't.
Jordan Wilson [00:12:38]:
Like, you literally can't. You used to be able to. You can't anymore because literally every single day, right, whether you're talking about, Google, Microsoft, OpenAI, Anthropic, now Meta's back in the conversation, perplexity, etcetera. Right? You literally can't have an actual job where you have to, produce something for a company and still understand what AI can do. Not anymore. You used to be able to juggle that. Now unless you're literally working twenty plus hours a day, you can't do it. Alright.
Jordan Wilson [00:13:05]:
Number three, AI literacy. And this kind of goes, you you know, hand in hand with, you know, number two, what business leaders understanding. So there's a difference between, understanding AI's capabilities and then AI literacy because you have to actually be able to speak the the language foundationally, which again, most people skip over that. Right? They go straight to the bells and the whistles. They go straight to clicking the button, not knowing what happens, you know, from a to x. They only wanna get to y in the result that brings on z. Number four is human skill set. Alright.
Jordan Wilson [00:13:41]:
This is different than literacy. You have to have the foundational understanding, but then you also have to have these skills. Right? I know a lot about basketball, but my basketball skill set, not great anymore. Right? I think I peaked in eighth grade. But you have to have that human skill set. And then last but not least is AI access. So that's access to the tools, to the technology in your role. Okay.
Jordan Wilson [00:14:04]:
So what can AI models actually do? Right? And I do wanna spend a little bit more time on that because that will better explain this gap that I think started manageable. And now it's very hard to manage that capability gap because it has turned into a frigging canyon. So the reason being is because AI models are better than expert humans. All right. I've talked about this, this evaluation a couple of times, but the more I follow these AI benchmarks and all of these other things, right, there's there's dozens of them. I think there's about five that matter. Right? And one of them, and I think probably the most important one, is GDP val. So this is a benchmark from OpenAI.
Jordan Wilson [00:14:48]:
And the reason why I think it's the most valuable benchmark to look at is because it's about creating business value. So more or less, this benchmark evaluates AI models against real professional deliverables across 44 different high GDP occupations. So these are things that you and I may do. You know, going in and doing research on a fast moving market segment and then creating, an artifact like a, you know, a a pitch deck or creating a spreadsheet, something like that. These are real world domain specific tasks where both a human and an AI model from start to finish, completed a task and then submitted their final version to a panel of experts who are experts in that field. Alright? And expert judges then compared the unlabeled AI and human outputs in a blind head to head pairwise comparison. And what we've seen now is, as an example, the best model for actually creating front to back economically valuable outputs like a human would is the OpenAI's GPT five four model, and it matches or exceeds industry professionals in 83% of these evaluations. Right? I remember when some of these first GDP valve numbers came out and I'm like, oh, that's that's pretty impressive.
Jordan Wilson [00:16:08]:
Right? When we were in the 40%, but now it's it's undeniable. And I do assume that by the end of the year, that number is gonna be like in the mid nineties. Right? My, like, one of my predictions is it was gonna get to 80% and we're already at 80% is just a huge jump. That's why I I keep talking about the AI capabilities are just running too fast for most business leaders to understand and keep up. And it's nearly doubled. That GP Bell score has nearly doubled in five months. So when I talk about the last, you know, since essentially late twenty twenty five, I'm gonna talk about why, but it has been we've seen more developments in the past five months than we have in the five years prior, at least when it comes to large language models, and it's not even close. Right? So the in October, the best AI model scored on that was a 47%.
Jordan Wilson [00:17:04]:
So it hasn't quite doubled, but it's nearly doubled in about five or six months. And right now AI handles most defined professional tasks yet most organizations barely use models in this way. It's like, did you know, right, as an example, your, you know, GPT and Claude models can do this. They can create artifacts by default. You don't have to have anything special turned on. It can just create spreadsheets, presentations, word docs, etcetera, all personalized based on your company's info. Alright. Next, I wanna talk very briefly about the anthropic labor data study that came out, about two months ago.
Jordan Wilson [00:17:46]:
I did go over this in more depth. So if you're interested in this, go listen to episode seven thirty. Here's essentially, similarly, whereas the GDP valve benchmark looks at actual economic output, this anthropic labor data study looks a little bit more of the capability gap. So that's why I think these two, kind of this benchmark, OpenAI benchmark in the anthropic labor database labor data study in tandem really tell a powerful story of where we're at and why this capability gap even exists. So in this study, anthropic mapped over 20,000 work tasks. Right? So they actually used, some public, some public jobs data from the federal government. So they mapped over 20,000 work tasks across 800 occupations to millions of real AI conversations with their Claude chatbot. These were anonymized, but essentially, they said, okay.
Jordan Wilson [00:18:45]:
What are people using Claude for? And they're matching millions of these chats to 20,000 work tasks across these 800 occupations, from this federal data. And what they found in theory was that computer and math roles could solve 94% of, sorry. Today's most capable AI models could automate up to 94% of computer and math tasks. Right? But they were only seeing 33% usage. Right? Because people assume, oh, computer and math. Yeah. Everyone's using AI for that because everyone knows that's the lowest hanging fruit, right? Anything in coding, software development, etcetera. Yet even in the area where people assumed, oh, yeah.
Jordan Wilson [00:19:34]:
Literally everyone, if you're working in anything computer, right, software engineering, anything dev, anything with math, of course, you're using AI. And they said, well, actually not. Right? And even office admin, business, financial, and legal roles all revealed the same deep, kind of adoption shortfall. Right? So, for our livestream audience, you can always see the video version of this on our website at youreverydayai.com. So I have the theoretical capability and observe usage by occupational category, kind of map that anthropic put together, very fascinating. But all this shows is in certain categories, right, such as management, business and finance, computer and math, you know, architecture and engineering, legal is another high one, arts and media. Right? Kate, like, AI's capabilities when mapped to literal, the actual, work tasks that the federal government uses to define these roles. I mean, most of them have coverage in the 80 to 90% yet the observed or or sorry, the, that's the theoretical coverage.
Jordan Wilson [00:20:43]:
So in theory, today's most powerful models could do anywhere from 80 to 90% of the actual work. Right? Computer and math was the high one where they had an observed coverage of a little more than 30%, but these other areas are not even 20% right there for the most part, you know, in the twenties, some of them below 20. So you have this huge gap where in most of these general cases, right? Like management, you know, in theory, today's AI models, if you understand the capabilities, they can automate 85 plus percent, yet you don't even have a 20% coverage. That is a crazy gap, right? Where essentially you have a magic wand that if you know how to use the magic wand and you know the magic words and you say, poof, work be gone, poof, the work is gone, but people don't know because it's impossible to keep up with. And some other things that they found, on this study, which is actually the workers that are most impacted, they said that workers most exposed to AI right now earn 47% more in hold graduate degrees at nearly a four times rate. So, essentially, I think people early on assume that AI would displace or would, potentially have the capabilities to displace workers who were maybe a little more junior or, you know, weren't as high up the kind of quote unquote knowledge work totem pole, so to speak. And what they found was the exact opposite. Right? It is those, people who have those higher degrees, and then conversely, people with physical and manual occupations register the lowest exposure according to Anthropic's study.
Jordan Wilson [00:22:23]:
So how did we get here so quickly? Right? How do we get as an example from GDP Val where literally five and a half months ago, it was even right. It was a little less. It was about a 47% tie or win rate against humans. Right. It was a, it was a coin flip On if the off the shelf model that you can pay $20 a month for was as good as a, you know, a human with a decade of experience who specializes in that to now humans aren't gonna be able to compete in a couple of months, right? By blind benchmarks, almost every single time within probably six months, humans are always going to prefer the AI model. How do we get here so quickly? Right. Two years ago, the AI models were not good at all. Right.
Jordan Wilson [00:23:14]:
They're they're actually pretty bad. Right. Especially, obviously, we have today's comparison to draw back on. But one of the biggest reasons I think is recursive self improvement. Stick with me here if you're not super technical. Alright. So recursive self improvement or RSI, it's a concept within artificial general intelligence where essentially an AI system improves its own source code, architecture, or training data, leading to a more capable model, which then improves itself further, creating a self reinforcing loop. Alright.
Jordan Wilson [00:23:41]:
So why am I talking about this kind of strange niche concept, called recursive self improvement on something about managing the AI capability gap? Well, because at the 2025, the AI companies started to either directly admit or to, kind of allude to the fact that they were all now using recursive self improvements, on their models. Right? So now their quote, unquote big models were good enough that it could start writing its own code. It could start improving itself. Right? I think, you know, anthropic with their cloud code, probably one of the most consequential, products of the AI life cycle outside of, ChattGPT. You could make the argument that Claude code is one of the most important outside of ChattGPT. Right? The lead at Anthropix Claude code says, yeah. We don't write code anymore. You know, Claude codes Claude code.
Jordan Wilson [00:24:36]:
Right? And we've seen the same, inferred by different researchers at in sorry, at OpenAI as well. You know, Google has been a little less direct, but they've still alluded to the fact that a lot of their models today, not just the smaller versions that are distilled from the biggest versions, but the biggest versions are being improved by themselves. Right? And these new big scary models, right, like anthropics, mythos, or, you know, open AI's, whatever it's called, you know, spud or glacier. Right? If you're listening to this in six months, these names these code names don't matter anymore. But one one of the reasons that maybe the general public isn't getting their hands on them is, well, they're too compute intensive, but they are using these models internally to improve the best consumer models and to also ship new products. And that's why Anthropic as an example, their ship rate in in February and March was straight up off the charts. Right? And then we found out later, well, one of the reasons was because they were using this mythos model to put out, a lot of this a lot of these new features that we're all using now. And and and this is why.
Jordan Wilson [00:25:43]:
Right? Because a year ago, it would have taken a team even using AI. It would have taken teams way longer. But now that you have this kind of recursive self improvements or models improving themselves and building new features that, use those models. Right? This is why it's hard now to keep up. So this started, like I said, in 2025 and ever since large language funnel capabilities have far outpaced enterprise training and learning and development. You know, even the companies that wanted to do it right, I think they could they could do it, you know, in, you know, quarter two, quarter three of last year. But now you can't. Right? Unless you have an entire segment of people, you know, maybe 5% of your total workforce that literally has no deliverables.
Jordan Wilson [00:26:33]:
And all they do right? I've said this all along. Your team needs a bunch of mes where all they do all day is they just scope different AI models. They play with AI releases. They sandbox thing. They're not building anyone for anything. They're just building solutions for what they think the company needs, and then they're training those people. But companies don't have that. Right? And that's why, at least right now, it is nearly impossible to keep up.
Jordan Wilson [00:27:00]:
And most companies, though, claim AI adoption, but they can't actually prove real results, and it's getting even harder and harder. Right? That McKinsey study that I talked about in, 2025 said that 88% of organizations are using AI, but only 6% are generating meaningful business profits. And right now, I think leadership confidence is running far ahead of what frontline practitioners report actually seeing on the ground. That's the thing. Right? So, a lot of times, the people who are in charge of AI, at certain companies now because it's getting easier to build, I think maybe their hands are on keyboard or their eyes are on monitoring agents a little bit more, and they are getting removed from what the frontline practitioners are actually experiencing. So not only is it pretty hard, to manage that gap just from a technical perspective, but from a change management and a people management perspective, it's getting even harder. And I think that's why some AI capability gaps are closing, but others are still remaining stuck. So as an example, right, AI is helping AI enabled humans close some gaps, but many still exist and are getting worse.
Jordan Wilson [00:28:13]:
So as an example, coding performance. Right? This in in in a lot of different benchmarks, this surge from single digits to 90% on structured benchmarks in three years in terms of what the AI itself was capable to do. And that helps, obviously, the humans that use this as part of their daily daily workflow close a big chunk of that gap individually. Right? But you now have this, you know, PhD level, models with, you know, science and standardized math capabilities that are rapidly approaching the ceiling of the current tests. So you aren't even necessarily able to know by the benchmarks what the capability gap is because these benchmarks are becoming saturated. So we even the AI community doesn't even fully understand these models capabilities because the benchmarks, right, a lot of them have been stuck in the 90 to 95 percentile. You know, a lot of them are at 98, 99. So I think the benchmarks themselves are getting saturated.
Jordan Wilson [00:29:08]:
So even the people building the models aren't even fully aware or understand what these models are actually capable of. So let's talk about how you can actually manage this gap and start to close it. Alright. Right now, the top performers, right, the top companies are investing more than 20% of their digital budgets to fundamentally rework their processes. All right. Let me repeat that 20% of their digital budgets. So whatever your digital budget budget is, I'm guessing most companies are probably at 1%. Maybe, maybe 5%, very few, only the top performers are investing that 20% of their digital budget to rework their day to day knowledge processes.
Jordan Wilson [00:29:54]:
Right? Because you need to learn to separate workflows into risk tiers with verification and human approval matched to each level. In organizations that wait for AI to become ready, they're gonna find that their competitors have already captured the advantage. Because in the same way, right, just to draw a little parallel here. What anthropic, and I keep saying they've won twenty twenty six so far. Well, it's because they were first. Right? They were first to, at least internally, close that capability gap and put it to work. Right? And that's why they've been able to outpace their competitors. I don't know how long it'll last.
Jordan Wilson [00:30:35]:
We'll see, because I think the, the other labs have closed, that gap. And I think we'll start to see that soon, but think of that within your own organization. Right? The first to close the gap is going to be able to accelerate at a pace that we haven't seen before. Right? That's why you see, unfortunately, a lot of these big companies, you know, Block as an example, cutting 40% of their workforce, and they seem fairly confident that they're going to be able to actually grow revenue, because of the way they they've completely reworked their organization. Actually, Jack Dorsey had a very fascinating essay that we shared about in our in our newsletter about how they're essentially flipping the work pyramid on its head. So here's what I want to have you focus on. One metric. Okay.
Jordan Wilson [00:31:24]:
I I want to make this very digestible for you, and hopefully actionable as we close out today's show. I want you to track the percentage of AI assisted workflow steps right now that are accepted without any rework or incident. And here's why. Because you probably most organizations, they kind of get their AI plan or their AI training for the year. Right. Or some companies, unfortunately, it's kind of like one time, Maybe the companies that are really investing heavily, you might get it once a quarter. But when was the last time that you used the most capable thinking models? I'm talking about GPT five four pro. I'm talking about Opus four six with extended reasoning.
Jordan Wilson [00:32:07]:
I'm talking about Gemini three one pro with, a higher thinking budget. Right? When was the last time that you used those and tracked the percentage of your workflows that get accepted without rework or incident? And then I want you to break that metric down by risk tier. Right? So the low risk, you know, drafting versus the high stakes legal or financial outputs. So organizations that measure and understand that operational reliability for the high percentage of AI assisted workflows that can be accepted that are lower risk. You will be surprised that depending on what your team does, depending on what your company does, depending on what sector you're in, you'd be surprised without too much investment aside from just reverse engineering your current day to day processes, you'd be surprised to say that about 30 to maybe 60% of a lot of the work that many of us do. Right? I'm not talking about, you know, people in specialized role. I'm saying if your organization has a thousand employees, you know, if you look at that work collectively, you you'd be surprised. I would say 30 to 60%, if you rescope everything, would fall in that can essentially be automated without rework or incidents and is in a lower risk or a medium risk tier.
Jordan Wilson [00:33:30]:
And if you're not measuring that on an ongoing basis, you can't scale. But that's step one to managing the AI capability gap. And you have to be doing this right. At least monthly. Again, a year ago, you could get away with quarterly. The way models, right? If you're using a model from last quarter, good luck. Right? That's showing up to a, to an F1 race in a bicycle. Good luck.
Jordan Wilson [00:33:59]:
You're going to get smoked. You don't see it a chance. So you can no longer take these year long pilots. These quarter long plans. You have to be agile to actually understand number one, the capability gap, but to begin to manage it. Alright. I hope this was helpful in our start here series. Like I said, make sure to repost today's episode on LinkedIn.
Jordan Wilson [00:34:19]:
Here's why. We put together the AI capability gap report card to help you know where your organization stands when it comes to number one, understanding this gap, and number two, how you can actually tackle it. So in this report card, it's a great guide, for you and your team to go through it together to understand the latest capabilities of all the models and how they break down for different types of work. Alright. So if you didn't know, yes. If you're listening on the podcast, this is actually live streamed on LinkedIn. So in the podcast show notes, we always put a link to today's, LinkedIn, show. So just go click repost, and I will send that capability gap report card your way.
Jordan Wilson [00:35:06]:
Alright. Thank you for tuning in. If you haven't already, please go to starthereseries.com. That's gonna give you free access to the inner circle community, and you can go listen to all of the start here series in order in the playlist that we have, and you can go read and listen to all of the shows there. So thank you for tuning in. We hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.
