Ep 804: Open Source Surge? Does GLM-5.2 Make Open Source an Enterprise Priority? (Start Here Series Vol 29)

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


Open Source AI Models: Evaluating GLM 5.2 for Enterprise-Grade Deployment

A significant convergence is redefining the path forward for large-scale generative AI adoption. Three concrete developments are now shifting enterprise priorities: open source models have substantially narrowed the gap with proprietary AI; cost pressures are shifting organizations away from indiscriminate AI usage; and major technology companies are actively exploring open source alternatives to existing commercial solutions. This intersection, highlighted by the release of GLM 5.2, serves as a moment of critical evaluation for enterprise leaders tasked with managing the cost, performance, and strategic risk of AI investments.


AI Performance Benchmarking: Open Source Model Competitiveness

Robust comparison across established benchmarks now shows that open source models such as GLM 5.2 have credibly matched or even exceeded some closed alternatives in specific domains. GLM 5.2, a 744-billion-parameter MIT-licensed mixture-of-experts model with a one-million-token context window, targets advanced coding and autonomous agent workflows. Notably, leading AI analysis indices rank GLM 5.2 just behind the latest Claude and OpenAI models—surpassing any Google offering as of the most recent updates. This marks the first time in recent memory that an open model ranks within the top three for overall utility, changing the selection calculus for enterprise applications that demand reliable and sophisticated language intelligence.

Enterprise Cost Management: The Token Efficiency Imperative

Rising operational costs have forced large enterprises to confront an era where AI usage is no longer measured by token volume but by token efficiency. Recent research and industry reporting document how strategies such as “token maxing”—rewarding employees for generating AI activity—have resulted in wasteful spending. Specific cases show single engineers utilizing hundreds of billions of tokens in a month, and organizational lapses leading to accidental outlays of half a billion dollars on a single model. A pivotal study from Stanford demonstrates that agent-based workflows can consume up to 1,000 times more tokens than legacy chatbot queries, confirming the growing gap between desired AI outcomes and sustainable resource allocation.

Security and Privacy: Open Source Risks in Production Environments

The integration of open source AI models into production carries well-documented privacy and compliance concerns. Chinese models such as GLM 5.2 are developed using methods that sometimes involve distilling U.S. frontier models, raising questions about trust, intentionality, and potential for geopolitical influence on output. Furthermore, the opaque nature of certain open source weights complicates direct validation and auditability, demanding a proactive approach to data governance, output scrutiny, and compliance reviews before production deployment.

Cloud Infrastructure and Local Deployment: Realistic Options for Enterprises

Despite public perception, running a model of GLM 5.2’s caliber is not feasible for average users or small businesses—hardware requirements start at no less than $15,000 per machine and deliver suboptimal performance in local environments. The viable deployment model for such technology is through robust cloud infrastructure or via enterprise-managed APIs, frequently leveraging established hosting providers like AWS or Hugging Face. For most organizations, the cost savings becomes significant only at high scale. GLM 5.2 offers a fraction of the operating cost compared to closed alternatives like Opus 4.8 or GPT-5.5, particularly for companies with their own GPU infrastructure and high sustained API volume.

Workflow Redesign: Balancing Capability Gap and Workflow Overshoot

A fundamental challenge in enterprise AI is not just matching model performance to business requirements, but also redesigning workflows to maximize ROI. Two interconnected barriers persist: the model capability gap—where organizations underutilize AI’s full potential; and what’s now identified as “autonomous workflow overshoot.” The latter describes how most enterprise workflows are not yet designed to absorb the continuous, multi-turn, agent-based automation these models can enable. For example, standalone agents capable of complex, unsupervised operations outpace current enterprise process controls, requiring strategic roadmap adjustments in job roles, deliverable management, and compliance oversight.

Scenario Planning: When Open Source AI Becomes an Enterprise Fit

Open source AI models are now realistic priorities for three specific enterprise scenarios:

  1. Organizations with Existing Compute Resources and High API Costs: Large companies—typically Fortune 100 scale—can ingest open models like GLM 5.2, fine-tune internally, and provide controlled employee access, offsetting large commercial subscription costs.

  2. Non-Agentic Task Automation at Scale: Enterprises can segment workloads, identifying routine, non-agentic tasks (e.g., summarization, basic content generation) where open source models offer cost-effective alternatives to closed, high-end models, especially via API deployment.

  3. Futureproofing for Task-Specific Model Adoption: As model distillation and specialization advance, state-of-the-art open models tailored to specific tasks (front-end coding, text summarization, PDF parsing) will proliferate, enabling enterprises to match model choice directly to business need without incurring excessive cost from generalized systems.

Conclusion: Open Source as Foundation for Enterprise AI Strategy

GLM 5.2 does not immediately bring open source AI to every desktop or small business, but it may represent the tipping point where enterprise AI roadmaps incorporate open models as an informed choice, not an afterthought. Cost considerations, benchmarking parity, risk mitigation, and workflow redesign are coalescing into a strategic imperative. Early adopters will not simply save on compute—they will set the blueprint for future task-optimized AI that balances sophistication with operational discipline.


Topics Covered in This Episode:

  1. Open Source AI's "ChatGPT Moment"
  2. GLM 5.2 Model Benchmarks & Performance
  3. Enterprise Adoption Drivers for Open AI
  4. Microsoft Evaluating DeepSeek for Copilot
  5. Token Maxing to Token Efficiency Shift
  6. GLM 5.2 Infrastructure vs. Consumer Use
  7. Autonomous Workflow Overshoot Explained
  8. Capability Gap and Workflow Challenges
  9. Enterprise Scenarios for Open Source Models
  10. Future of Task-Specific SOTA AI Models




Episode Transcript 


Jordan Wilson [00:00:16]:
There's three important things happening right now that make me think open source AI might be having its Chad GPT moment. Number one, the models are actually pretty good with a recent splash from z AI's GLM 5.2 leading the way. Number two, the era of token maxing is over as companies cut AI spend. And number three, one of the biggest and most influential companies in the world is looking at open source as a viable option. Granted, this doesn't mean that you'll have a frontier level model operating twenty four seven on your computer. That's not how any of this works. But for large enterprises, they will and do have that option today. But even if you're not a Fortune 100 company with GPUs to spare, you too are gonna have to start paying very close attention to open source models in 2026.

Jordan Wilson [00:01:08]:
Yes. The Chinese companies are distilling from American labs, and there's privacy considerations, but that doesn't change the fact that US tech companies are using these models in production as AI costs are starting to skyrocket. So will models like GLM 5.2 thrust open models onto the streets of mainstream AI America, or will this just be another drop in the bucket until the next wave of US lab models make the current open contenders look archaic in comparison? Well, let's find out on today's edition of Everyday AI as part of our start here series. Alright. If you're new here, welcome, but let's talk about the big picture here. Open source AI has nearly caught the top of proprietary models. So I think most even people who are bullish on open models would have admitted that for the most part, open source or open weight models are about six months behind. And I'd say now that gap is maybe only two months, two or three months, which is pretty incredible to see.

Jordan Wilson [00:02:10]:
And then a lot of benchmarks, which we're gonna look at, open sources kind of caught the closed proprietary models. So they are now finally credible enough and powerful enough for serious enterprise evaluation. Also, Microsoft is reportedly looking at DeepSeek as it looks to, lower its cost in Copilot co work. So that's huge. And the model pushing all of this, I think, right now is z AI's GLM 5.2. I know that's a mouthful, but but we're gonna look at some of the charts that show that this is now a big picture model. This is a big shake up, and you have to be paying attention to GL five two and what comes after this. And as companies now shift from token maxing to token efficiency, open models may finally be having their chat g b t moment.

Jordan Wilson [00:03:04]:
So on today's show, here's what you're gonna learn. You're gonna learn why Microsoft reportedly looked at DeepSeek for lower cost copilot agents. You're gonna see why GLM five two is an enterprise infrastructure play, not exactly a business laptop AI. You're gonna know why autonomous workflow overshoot blocks, adoption more than model quality. And I'm gonna tell you about that secret issue that I think people aren't paying attention to when it comes to, close proprietary and open source AI that autonomous workflow overshoot. Alright. Let's get into it. Welcome to the start here series.

Jordan Wilson [00:03:42]:
This is the everyday AI essential podcast series to both learn the AI basics and to double down on your knowledge. So if that's what you're trying to do, sweet. Me too. So, this is an ongoing series. We're actually on volume, like, 29 now. So make sure you go to starthereseries.com. That's gonna give you free access to our exclusive inner circle community. And there, there's a playlist that has all of the start here series, all in one spot on a Spotify playlist, all of the newsletters, everything all in one spot, and you can go connect with other business leaders that are trying to do the same.

Jordan Wilson [00:04:18]:
If you missed our last start here series episode, I think it was actually a really important one. We talked about AI super apps and why every company is racing to create one and what they are. That was volume 28 or episode seven ninety nine. And today, let's talk about the open source surge. So let's quickly recap these three different things happening all at once. So number one, Chinese open source models have kind of closed the gap. They haven't closed the gap completely, but they've closed the gap to its, like, teeny. Alright.

Jordan Wilson [00:04:52]:
So, obviously, I'm not gonna talk a whole lot on model distillation in this episode. But if you don't know what that is, essentially, Chinese companies steal more or less from American companies. Right? We've seen the US government is working, you know, on this with the top labs. But, Google, OpenAI, and Anthropic have all but said, yes, the Chinese companies are stealing all of our work in making open source models. So we're gonna look past that. And the reason why we're actually looking past that is, well, Microsoft. Right? Microsoft being one of, anthropic and OpenAI's biggest investors is reportedly looking at DeepSeek as a viable alternative to using the anthropic in OpenAI models. And that's one of the reasons why I think it's now finally time for enterprises to take a serious look.

Jordan Wilson [00:05:45]:
So, yes, there's obviously a lot of privacy data considerations when it comes to using open source models. Alright? Because you don't always know the weights. So, you know, you might be getting an output and blindly copying and pasting that knowing there might be a geopolitical reason you're getting a certain answer. So there's obviously a lot of considerations to take into account. But I think the fact that we're seeing, number one, the models are good enough. Number two, token maxing is going away. It's no longer about, oh, you know, everyone go burn 5,000,000,000 tokens so you can go climb up the internal company token leaderboard. That's over.

Jordan Wilson [00:06:20]:
Companies are cutting AI spend. And number three, Microsoft looking at open source models like DeepSeek as a viable alternative. So let's talk a little bit more about this new model that is catching everyone's attention. This is from z AI. I believe they're formally called Zifu AI, but they are called z AI, based in China. And GLM 5.2 is a 744,000,000,000 parameter MIT license open weight mixture of experts model. So it was developed by z e z AI, and it has a 1,000,000 token context window, and it is designed specifically for complex long horizon coding in autonomous agent workflows. And it targets coding tool use long context and agentic work.

Jordan Wilson [00:07:10]:
And here's the reason why we're talking about this y'all. It is incredibly good. Okay? I've used it a little bit. I've been impressed. I haven't had as much time as other people, and I'm gonna be reading some thoughts from other, kind of leaders in the AI space. But, when you look at artificial, the artificial analysis intelligence index, which was actually just updated to 4.1, a little detail there. But, essentially, this takes, about a dozen or so, widely used and widely respected benchmarks, puts them all together, and gives all of these models a score. So this is a good way to look at all the different kinds of benchmarks and to know about how smart or good a model is.

Jordan Wilson [00:07:55]:
So, right now, obviously, we don't have Claude Fable. Right? That could change at any minute. But, the best model right now, it is one a and one b, at least at what's generally commercially available. Claude Obis four point o 4.8 max and GPT 5.5, x high from OpenAI. And then not too far behind, now you have a GLM 5.2, which surprisingly scores higher than any Google model. Again, that could also change because, you know, Google did say in June they're coming out with their new model. So, you know, that could change later today or later this week or next week. But regardless, this is the first time that, you know, in recent memory, at least, you know, since the AI race became more than just OpenAI, this is the first time that an open source or open weights, company has cracked a top three company.

Jordan Wilson [00:08:57]:
Right? So it's now Anthropic, OpenAI, and z AI. Yeah. So coming in, ahead of Google, coming in ahead of Grok, coming in ahead of, you know, DeepSeek, Meta, the, Kwan models, all of these other, you know, even the other Chinese open source models, incredibly, incredibly well benchmarked. So a lot of accusations out there that maybe they're a little bench maxed or over fitting, you know, to make sure that they score really well, on these benchmarks. But I I mean, let me just read, for you all, some reaction from some respective people in the AI community. So this one from, Chris Som, who's a partner at Active Capital, Capital, who said, I'm starting to feel like GLM five two might be better than Opus. I was spending $300 a day on Claude, switched to GLM, switched to GLM, spent $3.82 today, and it found and fixed a bug, a Claude bug from yesterday. I honestly can't tell which is better anymore.

Jordan Wilson [00:10:06]:
Alright. Jeremy Howard, very well known name. He was the founding president at Kaggle. So he said, wow. Z a i's GLM 5.2 is a marvel. It is at least as good as Opus 4.8 in GPT five five. It is super fast, inexpensive, and not too verbose. It responds with nuance in judgment and handles long context very well.

Jordan Wilson [00:10:28]:
I've never experienced an open weight model like this before. And then, Guillermo Roche, who is the Vercel CEO. So, yeah, these are big big names. Right? You know Vercel. They're one of the leaders, in the AI space. So he tweeted out, genuinely impressed, almost almost shocked at how good GLM five two by z AI is at coding. This changes things. And then last but definitely not least, another prominent name in the AI space.

Jordan Wilson [00:11:00]:
This is, Matt Velaso, who's a former VP at Google DeepMind, also, VP at Meta and worked at Microsoft as well. So he said, all day using GLM five two didn't miss much. Alright. So saying compared to using other models. So all day using g l m five two didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be the same, and then talking about how we need to get some serious hardware. Alright.

Jordan Wilson [00:11:29]:
So now let's talk about this because, when people think open source models, what they default to is, well, oh, that means I can download this thing on my computer and I can run it twenty four seven and I can plug it, plug it into Open Claw or Hermes agent. All of a sudden, I have a, you know, a model that's the level of, you know, g p t five five and Opus four eight running for me twenty four seven, and I don't have to pay any API bills. I don't have any subscriptions, etcetera. No. Not at all. Not at all. Alright. No matter what anyone tells you, because you're gonna need at least $1,520,000 dollars in hardware, to get anywhere near that type of performance.

Jordan Wilson [00:12:11]:
Right? And I'm just saying the outputs. So if you're just comparing outputs to outputs, yeah, you're gonna need a machine that's at least $15,000, probably a little bit more, and it is going to be slow as a snail. So it is not economically feasible or economically reasonable, right, to, you know, run a model like GLM five two locally. You will also be running a quant version, so you'll be running like a two bit quant. So it's gonna be a dumbed down, very slow version. So, yeah, good luck doing that. But, you know, there's, people that argue out there like, oh, yeah. I'm doing this.

Jordan Wilson [00:12:46]:
So unless you have some very rare special reason, you you know, that you have to be able to run things locally. Right? Because that's the other, obviously, aside from costs, you know, once, models do become more powerful and smaller, and and efficient enough to run on consumer hardware, obviously, the privacy. Right? Being able to run things locally and not having to send things to the cloud. But that's not where you're at with GLM five two. Right? There's maybe, I don't know, a couple people in the world that have a powerful enough, machine to actually do that. But it's for most people, this is not going to replace your $200 subscription to Claude or Chad g p t. But this is a legitimate option for large enterprises who have access to compute. Alright.

Jordan Wilson [00:13:35]:
But small teams skip this. Alright? I I I'm not saying skip GLM five two because obviously, you can still use it on the API side. You know, there's plenty of hosts from, you know, AWS to Hugging Face, ever everyone in between that you can actually go in and run this model, using their servers, their GPUs. Right? But this is not something you need to overlook. Alright? And what do I mean by that? GLM five two might not be a a a model that you're going to run locally. It may not be even something that you run on the API even though it's a fraction, of the cost. Right? That's the other, big thing. It depends on what level you're using.

Jordan Wilson [00:14:17]:
Right? But if you're comparing, the highest output of GLM five two to the highest, you know, GBD five five or Opus four eight or Fable five, whenever it comes back, it is a fraction of the cost to run it on the API side. So, yes, there still is, you know, if you're building modularly, which is smart to do, and can, you you know, swap out a set of API keys or if you're using something like open router that makes it even easier to do that, sure. You know, maybe you're running this model in the cloud. But this is not something, you know, I think most people think, oh, open weights, you know, I can, you know, download it, fine tune it. You're not gonna be doing this on a consumer piece of hardware. But so that's a factor one. Factor one is, well, the model is actually good enough. Why this might mean, open source is having its chatty b team moment.

Jordan Wilson [00:15:08]:
Number two, the Microsoft move. So according to Axios, and this is just not even a week old. We covered this on our Monday show. So Axios reported that Microsoft is actually looking at lower cost open source alternatives to anthropic in, GPT models. So right now, both in Copilot, co work, but specifically Copilot, you know, they mainly have always relied on GPT models from OpenAI. Over the past year, they've started to integrate and, offer other, models specifically from, anthropics Claude models as well. But co, Copilot Cowork is obviously an agentic offering. This is based on Anthropic's very popular co work technology.

Jordan Wilson [00:15:56]:
It's kinda Microsoft's version. Microsoft is obviously a big, financial backer, in both Anthropic and OpenAI, and they also, well, serve those models. So they make money, on the cloud when their customers use these models. So the the the fact that and I do have to, you know, break this down, and I won't go too much into it. We just talked about it on yesterday's show. But the fact that Microsoft is potentially looking at DeepSeek of all companies. Right? Because all The US companies have called out DeepSeek by name for distilling their models. So, I don't know if I'm if I'm in leadership at OpenAI and Anthropic, not super happy with my my biggest financial backer, by potentially even though this is just a report, the report could be all wrong.

Jordan Wilson [00:16:48]:
Maybe it's true, maybe it's not. But the fact that this report is out there from a reliable source in Axios that Microsoft, one of the biggest companies in the world and one of the most trusted and respected companies when it comes to AI deployments. Right? The fact that they're looking at a company like DeepSeek to potentially be an option under the hood for running Copilot Cowork to make it more affordable to customers is big. Right? Because now that Copilot Cowork is generally available, that's new as well. Before, it was just in beta. But now they're charging usage. So they're having to, instead of subsidizing this like so many custom, companies are, right, they're charging end users for usage. And, well, what they're gonna see very far is Copilot, co work.

Jordan Wilson [00:17:37]:
No one's gonna use it if they have to pay the APIs for, you know, Opus 4.8, which is one of the most, aside from Fable, from Anthropic, it's Opus 4.8 is one of the most, you know, expensive models out there. You know, depending on the task, it's actually twice, more than twice as expensive as GBD 5.5. So pretty big here. But the reported goal was just cheaper agents inside of Azure security protections. And that pressure exists because of the agentic work behaves just like it's it's nonstop. Right? I think when we go back to the chatbot era, we didn't have to worry as much about, you you know, token usage. But now when these agents run-in loops, right, and it's it's easier to set off loops now, in in codex and clawed code. It's easier, to, you you know, put into this goal mode and, you you know, have an agent work for not just hours, but days.

Jordan Wilson [00:18:37]:
Right? So now all of a sudden, we have to start thinking about things like token efficiency and, you know, the cost of compute. Alright? And that leads us to well factor number three, and that's the, the shift that we've gone. Well, now companies are not just saying let's burn tokens. They're saying let's save tokens. So there's been a lot of recent reporting over the last week or two. You know, it's New York Times, article from this week that read tech workers maxed out their AI use. Now they're trying to minimize it. Business insider story titled Silicon Valley's AI token craze is facing a reality check, and then a TechCrunch article that said the token bill comes due inside the industry scramble to manage AI's runaway costs.

Jordan Wilson [00:19:23]:
And I did cover this, in-depth on episode seven eighty nine also in the start here series when we talked about token maxing, the shift from token maxing to token efficiency. But in short, you had Meta in other companies that had these internal leaderboards where they were reportedly just rewarding employees for using tokens. And I think it started maybe with good intentions. Right? Because AI leaders thought, well, hey. If our people are using AI a lot, that means that the business is going to grow, and we're gonna be saving time. But that backfired because people were just burning tokens. You you know, a reported example is one engineer, used 281,000,000,000 in a month just to climb the rankings, right, which is, pretty pretty crazy number. You you know, you had reports that employees were just running agentic loops intentionally to keep running even if they're, you know, doing personal projects.

Jordan Wilson [00:20:23]:
Right? But they were just burning tokens just to burn tokens because, well, there was a thought for a brief period of time, I'll say, from the end of twenty twenty five to early twenty twenty, '6. Well, people looked at employees all, like, for a brief period of time. Right? They saw token usage as, like, a KPI, as a a a key thing to evaluate employees on. Right? Like, hey. If you're burning tokens, thumbs up. You're doing a great job in our book. And then there was the report that a company reportedly spent $500,000,000 on Claude on accident in one month because they forgot to set limits. And that shows you just the amount of tokens.

Jordan Wilson [00:21:05]:
Right? So a Stanford study found that agentic task can use up to a thousand times more tokens than a single chat. Alright? Because chatbots answer once, agents can call tools endlessly in a loop. Right? I actually, for fun, I have my, Claude code ultra code going, on a simple task just to see how much usage it's gonna burn when it shouldn't burn anything. Right? But I'm looking at the chain of thought, and I'm seeing, Opus 4.8 going in some silly loops just burning tokens, you you know, like, I don't know, burning marshmallows on accident. But this is going to a broader capability gap that companies are already struggling to close. And that capability gap is half the problem, and I'm gonna go through this one quickly because I have covered this one in-depth before on episodes seven thirty five and episodes seven fifty five. So if you want to know more about this, I did cover it in both of those episodes. But, great study from Anthropic that talked a little bit about the capability gap.

Jordan Wilson [00:22:14]:
And this does get to, the kind of closed source versus open source and GLM five two stick with me here. But in that study, essentially, anthropic looked at anonymized chats and they hit a ceiling for these different, categories of work. And they said, here is the ceiling of a model's capabilities. And according to all these anonymized chats, here's what people are actually using it for. Right? So one of the best categories that had the best usage was only 33%, and that was computer math tasks. But mini tasks was not even 10%. So, you know, you could have the best, as an example, business strategist, and people were only using 10%. Right? And that's the baseline kind of model capability gap, and that's the first bottleneck.

Jordan Wilson [00:23:03]:
But there's kind of a new this is kind of my, my my secret term I teased in the beginning. I've been trying to put a name and a face on this. So, I'm trying this out. Maybe I'll rename it down the line. Right? There's always these, these common, common concepts that come up over and over. And it takes me, you know, sometimes a a month or two to put, you you know, to put a put a label on it. So you're not gonna see this anywhere. This is something I'm trying on for size here.

Jordan Wilson [00:23:33]:
But I think aside from the model capability gap, I think the bigger problem maybe is autonomous workflow overshoot. Alright? And I think that's the next gap that you need to prepare for. And that might you might see stick with me here. You might see how that might actually make some of your day to day tasks, Looking at a model like GLM five two, even though it's text only, it's not multimodal, it still might make it a feasible option in the future. So let me talk about this concept of autonomous workflow overshoot. Alright. So models right now, right, models by default, g b d five five, Opus four eight, Gemini three five three five Flash. They carry your company's context.

Jordan Wilson [00:24:22]:
They can plan. They can act. They can call tools. They can spin sub agents. Right? So, essentially, I ask companies, what would you do if every single employee had at least one or many twenty four seven agents? AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead. But you don't have ten hours a day to understand it all. That's what I do for you.

Jordan Wilson [00:25:05]:
But after 700 plus episodes of Everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode six ninety one. Number two, tap the link in your show notes at any time for the start here series, or you can just go to starthereseries.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same. The Start Here series will slow down the pace of AI so you can get ahead. And it's not a rhetorical question because those capabilities are here now.

Jordan Wilson [00:26:07]:
Right? I've had both Claude agents and Codex agents run for more than twenty four hours. That's the autonomous workflow overshoot because the capability gap is humans are not fully understanding the model's capabilities, or using the models to their fullest extent. Right? This is a human not taking advantage of what's there. But this is different. This is overshoot. Right? So, essentially, I'll argue that 9999% of bleeding edge capabilities go unused. Alright. And that's from a combination of the capability gap, but also workflow overshoots.

Jordan Wilson [00:26:53]:
So what is autonomous workflow overshoots? That's essentially that the capabilities of the agentic models are uncharted in your standard enterprise workflows. Because I will say that most workflows still need that human handoff. Right? If if if if I told you at your company, hey. Everyone has a twenty four seven agent that can carry your company's context. It can plan. It can act. It can research. It can call tools.

Jordan Wilson [00:27:27]:
It can create PowerPoints, Excel, anything, websites. Right? Hardly anyone. Hardly any business leader would be prepared to implement that. That is the overshoot. Right? Today's models are more than most companies can handle, not more than most humans can take advantage of the capabilities. That's two different things. Right? They can't handle this. That is the workflow overshoot.

Jordan Wilson [00:27:58]:
They these models are capable of just too much. Because I think that even AI forward companies right now need multiple quarters a year or even more to completely rebuild job descriptions, their inputs and outputs, approvals, and even the type of work that you do. Right? These things all have to change. This is also this concept of, yes, there is gonna be all these new AI. There there's gonna be all these new jobs that AI is gonna create. I ultimately think, AI will take away more traditional full time, roles than it will create. But AI is obviously going to create millions of jobs that we just don't know what they look like Because it's gonna take businesses a while to understand where all these autonomous agents capabilities are headed so we can start creating, both the context that the agents need, the expert driven, the expert driven loops that they need to do these workflows right, but also the lines of business, the streams of revenue that go along with those things. Those things are all gonna change, and that is the autonomous workflow overshoot.

Jordan Wilson [00:29:05]:
So thank you for sticking with me because now I can tie this together and answer the question well, should open source be an enterprise priority? Alright. And I'll say yes in three different scenarios. So let me lay those out for you. Number one scenario, if your company has access to compute and they also have a high API bill, right, they're shifting away, you know, from token maxing to efficiency. I think a general purpose model like GLM five two makes sense. Okay? Again, this is a very small sliver of companies. Right? This is essentially your fortune 100 companies. Not everyone has, you know, a a rack of GPUs sitting on the shelf and can, you know, if they want to, you know, roll out the type of compute that their enterprise needs.

Jordan Wilson [00:29:56]:
Like I like I'm saying, GLM five two, you can't download this on your every employee's laptop. It's not feasible. Doesn't make sense. Not gonna work. But if you have the compute, you can. Right? Yeah. And and what's crazy is I've I've talked to and I've worked with plenty of companies that fit into this category. I've seen their, you know, server racks and all the, in in Nvidia chips blazing and all the cooling, you know, water going underneath it.

Jordan Wilson [00:30:27]:
It's all above my head, but there's companies out there. Yeah. They can, you know, find they can download GLM five two. It's open weights. They can fine tune it, and they can, you know, make a a a portal, or a way that their employees can access this model. And in theory, they can use it twenty four seven around the clock. Right? Obviously, there's still bandwidth and other issues that you have. It's not as easy as, you know, click download and, you know, click deploy.

Jordan Wilson [00:30:51]:
But there are companies, number one. Well, maybe because of that, what I said, the autonomous workflow overshoot, I'm literally thinking of one company in particular that's spending millions of dollars a year on anthropic as an example. They could probably do this because I would say less than 1%, of people using, this, their current AI system that they're paying multiple 7 figures for are less than 1% are using it to its full capabilities. So in most cases, the 99% of people, a GLM five two would probably be enough. Aside from the fact that it's not multimodal, that that that is a huge downside. Right? But for the most part, number one, the companies that should be using it are those that have access to compute. And for the most part, they're not needing twenty four seven or can't take advantage of twenty four seven autonomous coding agents. Number two, those four non agentic tasks.

Jordan Wilson [00:31:53]:
Alright. And this is just a a different chunk, a different segment of work. So this is the, non agentic tasks that can and should be chunked for future open models. Alright. So like I said, so few people right now, their workflows require something like a fable five. Yeah. It's fun to go in there and, you know, let me make this three j s world view game and, oh, let me go code up this webs right? Yes. There's obviously people on in in in software development and dev roles that need that twenty four seven.

Jordan Wilson [00:32:29]:
But most people even using this, you know, you don't need Fable to write better emails. Right? So I think when you start chunking your large enterprise companies, start chunking your non agentic tasks, your non frontier tasks. I think that that's a big group of people that can start looking at an open source model like this, maybe using it via the API. And then last but not least, it's preparing for the future because I I hope that intelligence just, you know, I hope we truly do get the intelligence too cheap to meter, promise at some point soon. But there's a reality that as the capabilities increase, right, at least for now, for the most part, just well, that Stanford study showed that, well, agents just can burn a thousand times more tokens than a simple chatbot query. Right? And that is one downside of GLM five two. It is it is token inefficient. It burns through a lot of tokens to get that level of intelligence, although the level of intelligence is extremely high.

Jordan Wilson [00:33:38]:
The highest we've literally ever seen for an open model. But I do think that we're gonna see smaller, task SOTA models in early twenty twenty seven. Let me tell you what I mean by that. You you know, task is SOTA. So, you know, task specific state of the art models. Right? So right now, when we talk about models, they're just general large language models. Right? They're one model and people use it for everything. I've been a huge advocate and believer in the future.

Jordan Wilson [00:34:08]:
Right? There's gonna be a mixture of models technology. Thank you, certain companies that made my crazy 2023 prediction a reality in 2026. I was only a couple years too early. Right? But we've seen that, you know, OpenRouter just had their fusion technology, you know, perplexity with their computer. Right? You put a prompt out, and it will route route it to whichever model it thinks is best. It might put it through multiple models. Right? But I think we're gonna see that on a small, small language model platform here in 2027. Probably not this year, but I I essentially, what's gonna happen.

Jordan Wilson [00:34:45]:
Right? The model distillation is both, you know, when you I think most people think about model distillation. They think, oh, you know, Chinese models distilling a big frontier trillion parameter model, right, which is what we're seeing here with a lot of the, you know, the deep seeks and the quens. Right? All these accusations flying around. But what about well, just when, you know, frontier companies distill their own models. Right? The legal and, you know, teacher student model scenario. Or, you know, we're obviously gonna see a lot of these, Chinese companies come out with, small language models or task specific models, but I think those are gonna be state of the art. Right? So as an example, I think that whether it's through, you know, quote, unquote, not exactly legal distillation, or intentional legal distillation, I think we're gonna see state of the art open models for things like, you know, specific tasks, you know, front end coding, web search, text summarization, PDF parsing, you know, copywriting. I think we're eventually gonna see dozens after that probably, hundreds of open models that are state of the art at one task.

Jordan Wilson [00:36:00]:
Because when you can create a model around one certain task, it can be smaller, it can be better, it can be more token efficient, which in in turn, makes it cost efficient. Right? Which is what people are wanting. So those are, I think, the three scenarios where enterprises should be considering open source models, whether it's GLM five two or something else. So to quickly recap, number one, if your company has access to compute, in a high API bill. Number two, if you're an enterprise company that can start chunking all of these different, AI tasks into agentic versus non agentic, and maybe you keep your current agentic options that are maybe a little bit more expensive, and then you take your non agentic options too maybe on the API side, something like a g l m five two. And then last but not least, it's more, I think, for everyone else preparing for the future. And it is starting to categorize those tasks because not every task needs a fable five. Right? Not every task is gonna need a g b t five six pro, although I'll still use it for every task.

Jordan Wilson [00:37:09]:
Right? Like, that's how you also have to start thinking because eventually, these subsidies are gonna go away. We are gonna have to become a token efficiency mindset. That is the future. And I think is this the Chad GPT moment for, open source? Is GLM five two that thing? I don't know if it is. It's still too early to tell. However, I do think this is, if nothing else, the foundation for open source to have its chat g p t moment. Alright. I hope this one was helpful.

Jordan Wilson [00:37:45]:
If so, please let me know about it. Go to starthereseries.com. There, you can sign up for free access to our start here series, kind of, space inside of our inner circle community where every single episode is there ready for you to gobble up. Listen to it on two x. I'm not gonna be mad at you. Alright. I hope this is helpful. Thanks for tuning in.

Jordan Wilson [00:38:11]:
Hope to see you back tomorrow and every day for more everyday AI. Thanks, y'all.

Midroll [00:38:16]:
And that's a wrap for today's edition of everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit your everydayai.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers, and we'll

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI