Episode Categories:
Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
Maximizing Business Productivity with the Latest AI Advancements: Key Updates from Claude, Copilot Studio, ChatGPT, and More
Recent AI developments have equipped business leaders with new functionality that can be applied for immediate operational gains. The newest iterations of Claude Opus, Microsoft Copilot Studio, ChatGPT workspace agents, and related tools are providing granular controls, improved integrations, and advanced automation. The following analysis organizes these concrete updates by relevance and demonstrates their operational value for organizations seeking measurable improvements from AI adoption.
Claude Opus 4.8 AI Model: Performance, Controls, and Cost Considerations
Anthropic’s launch of Claude Opus 4.8 was benchmarked as a model leader in coding, reasoning, agentic skills, and collaborative tasks. Key details highlight how this release brings both opportunities and caveats for business deployments:
Enhanced Reasoning Controls: The introduction of a graduated “effort control” slider enables precise adjustment of response depth and speed. This allows strategic allocation of compute resources, matching model performance to business task complexity.
Token Efficiency and Pricing: While Opus 4.8 maintains pricing parity with its predecessor, it is notably token-inefficient. Sophisticated outputs require a significantly larger token budget, impacting operational AI costs. On high-volume or resource-intensive tasks, this can rapidly exceed expected usage limits, even for premium-tier subscribers.
Dynamic Workflows: Claude Code’s new capability orchestrates hundreds of parallel sub-agents, supporting tasks such as large-scale codebase migrations and stress-testing plans. However, this parallelization can cause rapid burn of API credits or subscription limits and should be monitored carefully.
Microsoft Copilot Studio: No-Code Agents that Use Your Computer
Microsoft Copilot Studio’s expansion now includes general availability of “computer using agents.” These agents interact with desktop applications and websites through the actual user interface, not just pre-built APIs or connectors. Practical implications:
Process Automation Beyond APIs: Organizations can now automate manual tasks that historically required human intervention due to lack of direct systems integration. Agents can, for example, pull key operational data, update spreadsheets, or interact with web dashboards that previously required repetitive clicks.
Enterprise Security and Flexibility: Newly added controls allow secure management of credentials and dynamic adjustment to interface changes, mitigating the downtime and errors common with brittle automation scripts.
Integration with Existing Workflows: These agents can be embedded in multistep workflows, bringing operational consistency and freeing human capital for higher-value efforts.
ChatGPT Workspace Agents: Customization and Reasoning Power for Teams
Significant under-the-radar updates to ChatGPT workspace agents provide substantive utility for business applications:
Model and Effort Selection: Builders now have fine-grained control to select specific model variants and “thinking effort” levels on agent creation. The ability to match model horsepower to workflow step—choosing lighter models for routine steps and higher-capacity models for complex reasoning—directly translates to optimized team productivity and cost management.
Role-Based Publishing and Guided Setup: Updated permissions allow IT and operational leaders to determine who can publish new agents, while guided setup accelerates onboarding for non-technical staff.
Google NotebookLM and Drive: Real-Time Data Syncing for Knowledge Work
A quietly impactful update now enables NotebookLM (Google’s Gemini-powered AI note-taking tool) to automatically synchronize with Google Drive:
Live Document Sync: Changes in connected Google Sheets, Docs, and Slides are now reflected in corresponding AI-driven notebooks without manual uploads. Deleted or access-revoked files are dynamically removed, maintaining data integrity and compliance.
Agent-Driven Knowledge Updates: Teams already employing automated agents to update documents are now equipped to have these changes immediately reflected in their AI workplace memory. This reduces repetitive administrative work and allows for rapid daily updates, such as dynamic audio summaries or metric tracking.
Microsoft Copilot WorkIQ Redesign: Transparency and Contextual Relevance
A new design for Microsoft Copilot underpinned by WorkIQ introduces:
Progressive Disclosure and Control: Users can control when Copilot uses organizational knowledge, with a visible toggle for including work data such as emails, files, meetings, and chats. The model surfaces thinking traces, relevant data sources, and actionable next steps.
Contextualized Responses: AI support is better grounded in business context, supporting wider organizational processes such as review cycles or cross-department initiatives. This clarity makes Copilot a more effective augmentation tool in daily knowledge work.
Eleven Labs Dubbing v2: High-Fidelity Multilingual Content Localization
Eleven Labs released Dubbing v2, an AI model that dubs spoken content into 90+ languages and accents while preserving original tone, emotion, and delivery:
Audio-to-Audio Translation: Unlike traditional transcript-based dubbing that strips nuance, this model conditions on source audio performance. Organizations can localize video, training content, or internal communications for international audiences without losing the original speaker’s emotional intent.
Automation for Scale: Fully automated pipelines eliminate manual intervention, reducing localization bottlenecks and ensuring rapid distribution of materials across global teams or markets.
Practical Takeaways for Business Value
Drive Digital Transformation with Automation: Copilot Studio’s no-code, computer-interacting agents make automation of legacy manual processes achievable—even in environments with limited API support.
Balance Cost and Capability in AI Model Selection: Fine-tuning “effort” levels and model variants in Claude Opus or ChatGPT workspace agents provides a new lever to align AI costs with expected task value.
Foster Real-Time Knowledge Work: Seamless syncing between NotebookLM and Google Drive enables knowledge teams to maintain always-current project files, dashboards, and meeting artifacts.
Localize at Scale While Preserving Personalization: Advanced dubbing solutions from Eleven Labs allow organizations to expand their reach internationally while maintaining quality and nuance.
Approach Upgrades with Caution: As AI models roll out advanced features, new cost dynamics may accompany visible productivity boosts. Regular monitoring will be necessary to realize ROI.
By translating these targeted updates into actionable steps, organizations can capture immediate operational improvements and position themselves for scalable—and sustainable—AI-powered growth.
Topics Covered in This Episode:
- Anthropic Claude Opus 4.8 Model Performance
- Claude Opus 4.8 Token Efficiency & Pricing
- Claude 4.8 Code Dynamic Workflows Overview
- ChatGPT Workspace Agents Model Controls Update
- Copilot Studio Computer Using Agents Launch
- Microsoft 365 Copilot New WorkIQ UI
- Google NotebookLM Google Drive Auto-Sync
- Eleven Labs Dubbing v2 Emotion-Preserving Translation
Episode Transcript
Jordan Wilson [00:00:18]:
Another day, another world's best AI model. That's right. Anthropic just released their latest version of Opus Opus 4.6. And, well, at least according to most benchmarks, it is the best model out available. And although the reviews are a bit mixed, it looks like at least for now, Anthropic will regain that top spot. So that's probably an AI update that you didn't miss this week, but there's a handful ish that you probably did that flew under your radar. Like, did you know that now you can build an agent in Microsoft Copilot Studio without code that can actually use your PC? Or did you see that Chad GPT's agents got a pretty big under the hood update that I literally didn't see anyone talk about? Like, not exaggerating. I didn't see this anywhere.
Jordan Wilson [00:01:20]:
So, yeah, unless you're glued to a bunch of RSS feeds and Twitter drops or reading company blog posts like I used to read the back of a cereal box on Saturday mornings, you're probably gonna miss some of the most important AI updates that you can start using today. Don't worry. I work for you. I stay up late and wake up early, so you know what AI features are worth your time and which ones you should avoid. So this is our new Friday features. Let's get into it. So on today's show, here's what you're going to learn. You're gonna learn the sneaky way that Anthrapic's new powerful Claude Obis 4.8 might end up costing you more money.
Jordan Wilson [00:02:04]:
You're gonna see why Microsoft's agent updates bring them up to speed with everyone else. And speaking of agents, you're gonna know a small chat gbt agent update that is gonna make a big difference if you start using it today. Alright. Let's get into it. Welcome to Everyday AI. My name is Jordan Wilson, and, yeah, we do this thing every day, Monday through Friday at least. It's an unedited, unscripted daily livestream podcast and free daily newsletter, helping everyday business leaders like you and me make sense of the barrage of AI updates. I tell you what's important, what's not.
Jordan Wilson [00:02:39]:
You use that information and become the AI wizard in your company. Everyone's like, how did this person know so much? Well, it's because you listen to me, and I don't sleep and neither do my agents. Alright. So if that's what you're trying to do, be the smartest person in AI at your company, our website is where you make that happen, youreverydayai.com. Go sign up for the free daily newsletter. We recap the highlights of each day's show as well as all of the other AI news that you need to know to keep up and get ahead. Speaking of keeping up, if you wanna keep up with me, I'll actually be in San Francisco, next week. So I'll be out there, I think, Monday through Wednesday.
Jordan Wilson [00:03:20]:
So if you happen to be in San Francisco, make sure to hit me up in the show notes. I always leave my LinkedIn. Just make sure to leave me a message. Otherwise, I'm not gonna know who you are. Stranger danger. Anyways, I'll be out, checking out the Microsoft Build Conference. So, let's get into it. Our first one, this one sounds very small, but it's actually pretty big.
Jordan Wilson [00:03:43]:
And I'm gonna give you some random ideas, on this one. But notebook l m and Google Drive are now working a little bit better with the latest update. So notebook l m now has Google Drive sync. So here's what it is. That's automatic syncing between Google Drive and NotebookLM. So when a Google Docs, Sheets, slide, file changes, whatever, the notebook updates automatically if you've added, one of those things as a source in any of your notebook without having to manually upload it. So notebook l m also strictly respects file deletions and permissions. So removed or access revoked files also sync and then will drop out of that notebook.
Jordan Wilson [00:04:29]:
Alright. And this is available to all Google Workspace customers and users with a personal Google account that have access to notebook l m. So, essentially, everyone. So it already started to roll out, and Google did say this will take about two weeks. So make sure you go check it out. There's no admin control, no end users for this or anything like that. So if you don't know notebook LM, like, okay. First of all, how? Number one, I use it literally multiple times a day.
Jordan Wilson [00:04:59]:
I used it so much today. I was like, how did I function before this? And one of the things that I, it was actually a little kind of quote, unquote, cheap code because there was a way to manually do this before that most people didn't know about. But now having this automatic, I think, opens up a lot of new, use cases. And think about this. And the quick thirty second primer on notebook LM, if you are super new here, it's Google's amazing, kind of grounded AI app powered by the latest Gemini models, but it is grounded in your data. So you can add different sources, whether those are static files that you upload or dynamic links to your Google Sheets, your Google Docs, your Google Slides, etcetera. So that's the big unlock here because, well, guess what? Those docs are live. They're dynamic.
Jordan Wilson [00:05:49]:
So one thing just to put this little bug in your ear, I constantly have agents updating my Google Docs. Now I don't have to do anything else in notebook l m. So this piece to me is huge. I can have just a daily, and this is what I'm gonna try once it does roll out to my account. I wasn't in the first, kind of tranche of of rollouts here, but I'm just gonna have my, my, agent. So I've been using codecs to go into my notebook LM using the, browser use and just update all of my notebooks. Now I don't have to do this because it's gonna automatically sync. But now all it's gonna have to do as an example is create a new, custom audio overview for me every single day.
Jordan Wilson [00:06:33]:
So as long as I have those connections already set up in a notebook, so maybe it's, you know, a certain Google Sheet I use to track important metrics, a Google Doc that an agent will update every single day. I don't have to configure it each day, add a new notebook, all that. So this right here, y'all, this is that personal assistant. We finally have it. So this is a small feature. It's like, oh, these two things that already work together. Well, now they sync. So just add in one more layer.
Jordan Wilson [00:07:00]:
So I can't wait until Google, and and maybe I'll reach out to, the the team there, and get, you know, give this as a suggestion. If this could just give you automatically a new, you know, audio summary every day, a new audio overview, that would be amazing. There's literally very popular apps that this is all they do. So, this is one. If you couldn't tell, I'm personally geeked about a small one, but if you are a, Google user or just a NotebookLM fan and you love those audio overviews or any of the other, artifacts in the NotebookLM Studio, this is one you should definitely be trying out. Alright. Our next update. And, yeah, this is one I literally forgot where I found this, and it is a random, help article on the OpenAI website.
Jordan Wilson [00:07:54]:
This wasn't on the normal release notes. This wasn't on OpenAI's Twitter accounts, a blog post. It was on the, Chad GPT enterprise and e d u release notes, which is not the one that I normally check. So I almost missed this, but this is actually another one that I'm excited for, a small one that I don't think you should overlook. So here's what's new. New updates to workspace agents. There's new controls and capabilities. So if you don't know, workspace agents are shared and reusable team agents, kind of like an evolution of custom GPTs that run multistep workflows in the cloud across connected business tools, and you can build them, you know, no code, right, conversationally.
Jordan Wilson [00:08:40]:
So right now, workspace agents are only available on team plans. So if you're on a business, e d u, enterprise, etcetera. And I think that these have been slept on because of all the codecs hype. Actually, right now, workspace agents are one of only, like, three things that I use on the chat g b t web app. It's deep, deep research, workspace agents, and Canvas. Although Canvas, just yesterday got kind of removed but replaced with writing blocks. So you're not gonna see the canvas button there anymore, FYI, but the canvas feature and functionality is still there. I was talking to the team and, encouraging them to add a toggle back, and it sounds like they might.
Jordan Wilson [00:09:22]:
Anyways, here's what's new. There's new model and thinking control for the workspace agents. So now there's a model selection that appears in the composer when you are building workspace agents, and there's thinking effort controls that have moved into the model picker. So before, you didn't have this fine con fine tune control when you created a a workspace agent. Right? It just used the latest model, g b d five five, but there wasn't a lot of, information on what version of g b d 5.5 that it was using. Because, technically, I'm doing the math here. There's instant. There's four levels of thinking, and there's two levels of pro.
Jordan Wilson [00:10:03]:
So there's technically seven variants, that you can choose from. So not having that control. Right? I've said this many times, on the, on the show before. Although the, GPD 5.5 instant model has gotten better and it also was updated yesterday, I still don't use it. Right? It's it's like I always use a thinking model, and I would encourage you to do the same. So this is pretty big, just from the model, capability to set the thinking level, is a big upgrade for those workspace agents that you can run around the clock and connected to all of your apps and connectors as well. So, in the agent builder, you can add those tools, apps, custom MCP, skills, files. I mean, the workspace agents are getting absolutely slept on.
Jordan Wilson [00:10:50]:
And I think the reason is is because, codex is just blowing up. So, here's why it's useful. Well, you can pick the model and you can dial in the thinking effort and lets builders, you know, match the horsepower to the task. So you can use lighter models for simple high volume steps, or you can use the thinking and pro at higher effort for more, complex reasoning. And then there's also, updated connector action constraints, that add control beyond right action, approvals by restricting how specific connectors can be used. So here is, the little tidbit here from OpenAI. So they said new controls and capabilities for ChatGPT workspace agents. We're rolling out new model admin app access and responsibility capabilities for ChatGPT workspace agents in enterprise and EDU.
Jordan Wilson [00:11:46]:
It says enterprise and EDU, but I did verify this under, the model capabilities are also available in the normal business plan. So there's normal business plans. There's enterprise and EDU. So although they don't denote it in this help article, the, model, kind of reasoning effort controls are available in the business plan. So aside from that, there is also the, updated role based publishing permissions. So workspace admins can control which roles can publish agents to the shared workspace directory. And then you also have the guided agent setup. Cheggi p t now asks setup questions to help users create useful agents more quickly and a couple other smaller things.
Jordan Wilson [00:12:29]:
Alright. So pretty pretty nice update there as well. Alright. Next, we have a new anthropic update that is not their new Opus 4.7. So here's what's new and why you might wanna keep an eye on your Claude bill. Alright? So, this is a new Claude code dynamic workflows. So this is a new Claude Code capability where Claude can plan the work and then run hundreds of parallel sub agents in a single session, then verify its outputs before reporting back to the user. So for example, you know, Claude Code with Opus 4.8 can carry out code based scale migrations across hundreds of thousands of line of codes, from kickoff to merge with the existing test suite as its bar.
Jordan Wilson [00:13:21]:
So here's who has access. Well, it's available now. If you use Claude code, for enterprise team and max plans. So I have tested this out a little, and my gosh. Goodbye to your rate limits. You thought Anthropic's rate limits were tough before. Try using hundreds of agents in parallel. I have the $200 max plan, and I was able to run one prompt.
Jordan Wilson [00:13:44]:
I saw other people that were on the $100 max plan and tried to test this new feature out. Actually, I don't know anyone that was actually able to, and they hit their limits. So, yeah, like I said, good luck. I think this is only if you're, you know, like, there's that story that I don't know if it's true or not, that some large company accidentally spent $500,000,000, on their, I think it was their Claude bill because they weren't looking at usage. So, yeah, Claude rolled out a lot of things, that they're like, oh, it's you know, look at all these great features, but all it really does is run up your bill and exhaust your limits very quickly. Although in theory, a really cool and useful feature. So if you are Scrooge McDuck or the monopoly man and have play money, this is great. So here's what Anthropic says.
Jordan Wilson [00:14:35]:
They said, today, we're introducing dynamic workflows in Claude code, helping Claude to take on the most challenging tasks end to end. Work you normally plan in quarters now finishes in days. Claw dynamically writes orchestration scripts that run tens to hundreds of parallel sub agents in a single session, checking its work before anything reaches you. Some problems are too big for one pass by a single agent, especially in complex legacy code bases, a bug hunt across an entire service, a migration that touches hundreds of files, a plan you want to stress test from every angle before you commit it. Dynamic workflows can handle all of those end to end. So, yeah, I tested this once, but on a very small code base. Right? Just on, a tool that I've been building off and on for the past, three to four months. Who knows if I'll ever release this? I've thought about it.
Jordan Wilson [00:15:31]:
You know, it's it's actually really cool. I use it every single day. But it's a it's a small code base. So if I had a larger code base, I don't even think I would be able to do this on the $200 max plan. I think this is just for API credits. It's like every time someone hits that feature, you know, anthropic adds another, you know, I don't know, billion dollars to its valuation or something like that. So, why it's useful? I mean, it just collapses those large multi day engineering efforts. So when you're talking about, you know, mass migrations, repo wide, refractors, into, refractors into a single orchestrated session instead of manual subtask management.
Jordan Wilson [00:16:08]:
So, obviously, who's gonna find this useful? Inch engineering teams. Right? And just enterprise development organizations. So, good one if you have the budget and well, if you're, really a dedicated engineering team. Alright. The next one that I think will probably be more widely, used is the new well, new ish, Copilot studio computer using agents. So these were announced actually quite a while ago, but they now are just generally available. Yeah. It's kind of a problem with, you know, Microsoft and Google.
Jordan Wilson [00:16:47]:
Sometimes, you know, their conferences, they release you know, they announce things, and they're released right away. And then other times, they go to these, you know, frontier programs or trusted tester, you know, which is, like, 0.1% of the general public. And it stays there for, like, six to twelve months, and you never hear about it again until yeah. Well, now this is one of those, the computer using agents inside Copilot Studio is something when it first came out. I think I actually went out and bought a laptop, a a PC just to use this, and they just didn't become generally available until now. So I might have to go find that laptop. Alright. So here's what it is.
Jordan Wilson [00:17:25]:
So computer using agents are now generally available in Copilot Studio, letting organizations build agents that interact directly with websites in desktop applications through the user interface to automate processes where the underlying systems lack APIs. So this is definitely a new enterprise capabilities that lets organizations manage credentials more securely, choose models best suited for different automation scenarios, and build automations that adapt to changing interfaces instead of breaking when a screen or web page changes. So computer using agents can now be embedded directly into multistep workflows as well. That feature has moved into preview. So who has access? Well, right now, everyone does. It is generally available. So if you do have a Microsoft Copilot, Studio license, you will have access to it. So why it's useful? Well, I think number one, if your organization only uses Microsoft Copilot and you hear all these cool things happening in cloud code and codex and how I'm talking about my agents are just running the, you know, literally using every single, application on my computer.
Jordan Wilson [00:18:42]:
I'm looking at my current, agent that's in goal mode now. It's at twenty five hours running straight. Right? How do you do that? Well, you give a computer use because, you know, yes, MCPs, these model context protocols are great to bring data from all these different services. You know, Microsoft Copilot, like others, they have apps and, apps, integrators, connectors, whatever you wanna call it that bring your dynamic data sources in. But what about all those other things that don't have MCPs and don't have APIs or don't have, you you know, apps or connectors supported within Microsoft Copilot Studio. Right? The simplest thing for me is, like, as an example, looking at my podcast stats. But sometimes for me, that takes, like, so many clicks. It's almost impossible.
Jordan Wilson [00:19:26]:
Or one thing I like to do is, you know, look at, you know, release threads. Right? So, on Twitter. So when a company you know, that's kind of their preferred channel now. I like to see what people are talking about, what's working, and what isn't. I don't wanna sit there and scroll through all those. Right? Social media is distracting. So that's an example of a computer using agent that can take over your browser and really do anything that there's not a direct connector for already. Here's what Microsoft says.
Jordan Wilson [00:19:54]:
Computer using agents are now generally available, With computer using agents now generally available in Copilot Studio, organizations can, organizations can build agents that interact directly with websites and desktop applications through the user interface. This helps you automate processes that previously relied on brittle scripts or manual workarounds because the underlying systems lack APIs. With the new release also comes new enterprise ready capabilities designed to help you operationalize UI automations more confidently. Organizations can now manage credentials more securely, choose models best suited for different automation scenarios, and build more resilient automations that can adopt two changing interfaces instead of breaking whenever a screen or web page changes. Alright. So, pretty exciting one. And our next AI feature that, hopefully, you didn't miss, this one, there are a couple new features, but it's going to look completely new. That is, and I'm guessing this is in preparation for Microsoft build next week.
Jordan Wilson [00:21:02]:
Microsoft is rolling out. It's live now on some accounts, but it is rolling out, to everyone. A brand new, design for Microsoft Copilot. So for, with WorkIQ. So here's what's new. A new design for Microsoft three sixty five Copilot built on progressive disclosure. So So Microsoft says that Copilot begins with a clear, readable response, then adds structure and next step support as you refine what you need with formatting when it improves clarity, suggested prompts when they deepen work, and follow-up actions when they move it forward. So this progression is powered by WorkIQ, which is now, a little bit easier to see and use.
Jordan Wilson [00:21:48]:
It is a drop down toggle to either include work data work data and memory or to not include it. So there's a little toggle that you can turn work IQ on by default, which is great because then you can see in the responses, the thinking trace, the searches, and what data it is automatically going to pull in, from kind of that WorkIQ graph. So kind of this new progression is powered by WorkIQ, the intelligent layer you can see, when active and directly control. So this draws on your emails, files, chats, meetings, all that. So WorkIQ adopts to the depth your work requires, including the ability to choose between AI models when that can surface more relevant results. So who has access to this right now? Well, it's rolling out now. So this is the consumer slash m three sixty five Copilot app experience, right, which is distinct from Copilot Studio. So here's why it's useful.
Jordan Wilson [00:22:55]:
Well, it aims to make Copilot's controls feel less intrusive while keeping AI assistance close to Word, Excel, and PowerPoint work surfaces. So also by grounding in your broader context and not just individual art, individual artifacts, WorkIQ helps Copilot support significant shifts like performance reviews, cycles that an org can change. So who's gonna find this useful? Well, I mean, anyone that uses, Microsoft three sixty five Copilot daily. Right? And if you didn't like the previous or, I guess, technically, the current because this is a slow rollout, but it is live now. So if you didn't like the other Copilot UI, you know, you might like this one a little bit more. You know, the big thing that I'm seeing again, I talked about this, I talked about this on the show, yesterday. It looks like everything else. Right? Copilot maybe wasn't the best designed previous experience, but it was at least a little bit unique whether you like that or not.
Jordan Wilson [00:24:07]:
So the new design from my, perspective is a little flatter, a little technically cleaner, you know, more monochromatic and less, colorful. So on the surface, it's a design aesthetic, but, also, there are some, you you know, new features and new tweaks, that Microsoft says make it a little bit easier, to use. So, there's that. Let me know if this is something that you guys wanna dig into more once it's released. Alright. Next one. This one, pretty impressive. Dubbing v two from Eleven Labs.
Jordan Wilson [00:24:50]:
So, what is this? Well, as you can guess, it's dubbing the second version. Alright. So, Eleven Labs, obviously, a leader in, text to speech, but this is a new dubbing model that translates spoken content across languages while preserving the original delivery. So dubbing the two conditions on the source performance, not a transcript, that's the key thing here. So your tone, emotion, and delivery carry across every languages. So right now, it supports 90 plus languages and accents, enabling localization for international audiences. So this is already launched, and it's available in Eleven Labs dubbing studio. It's web based, and it works for individual creators through enterprise teams on the Eleven Labs platform.
Jordan Wilson [00:25:40]:
So, here's why it's useful. Well, it's conditioning on the actual source performance rather than a transcripts, meaning the dubbed output keeps the speakers emotion in pacing, which transcript based dubbing tends to flatten. Right. That's also obviously a little bit better for someone that's very emotive like me. Right? Like, if you don't, watch the video version of the podcast, you know, sometimes I'm, you know, flailing my arms in the air and making crazy faces. It's funny. I have people that, you know, sometimes send screenshots when I'm making a very unflattering face or something like that. But, you know, it captures that same emotion.
Jordan Wilson [00:26:23]:
And looking at some of the demos, I haven't had a chance to do this, just yet because this one literally just came out. But from the demos that I've seen from, you know, people that I kind of know or, you know, I I watch or listen to these people, really good. Right? Not just preserving the tone of voice across multiple languages, but just really, capturing the emotion. Right? It actually sounds like this, you know, these people that I've heard their voice before. So, obviously, who's gonna find this valuable? I mean, content creators, podcasters like me, but also just video teams, you know, HR departments, people in learning and development. And, like, I actually think there's a ton of use cases, people doing international business. If you're a global corporation and you want to have a more localized and friendly onboarding for the new, you know, 100 people that you train every single week, whatever it may be, this is pretty big. And results impressive because, normally, I wouldn't put a text to speech update on, you know, this Friday feature show, but this one was that good.
Jordan Wilson [00:27:32]:
So this is how eleven Labs describe it. They say dubbing v two brings high quality dubbing to creators, marketers, and studios, fully automated with no pipeline to build. So they say this supports source audio, source text, and target text. The full pipeline, translation, cloning, dubbing, and syncing runs automatically with no manual intervention. They say it's perfectly synced. It is an audio to audio model, so it doesn't require a transcript or text, and they say close to human quality. So, again, pretty impressive one for me that I'm like, yeah. I think people need to hear about this.
Jordan Wilson [00:28:13]:
Alright. Speaking of hearing about our last big Friday feature update, You probably didn't miss this one. There's a new, not quite undisputed, but probably king of the hill when it comes to AI models, not harnesses, though. So keep that in mind. So this is more of, I think and this goes to my point that I made in yesterday's show, about how, you know, all models are starting to be the same, and it's more about the harnessing, the tool calling, what works under the hood, which I know is technically part of, you know, a model. But as an example, you know, the harnessing of codecs using a g p d 5.5, by all measures, is much better than the harnessing of Claude Code now using Opus 4.8. So just a a quick distinction there, for our audience. But let's talk about Claude Claude Opus four eight.
Jordan Wilson [00:29:12]:
It is good. It is impressive. So Claude four eight is an upgrade over Opus four seven with improvements across benchmarks for coding, agentic skills, reasoning, and practical knowledge work, and is a more effective collaborator. So, there is a new effort control on Claude. Hey. Another sneaky way that you're probably gonna burn through your, usage pretty quick. So there is a new effort control on Claude and Cowork that lets users choose how much effort Claude puts into a response. So a higher effort thinks more frequently and deeply, and a lower effort responds faster and uses rate limits more slowly.
Jordan Wilson [00:29:55]:
So, yeah, this one's I'm actually super glad that Anthropic did this because with 04/07, they took away extended thinking and they introduced adaptive thinking, which adaptive thinking was a disaster. Right? It was a toggle, and it essentially, I would always wanted to think, and I would instruct it to think, and it would never actually use its reasoning and logic abilities, which when it used it, it was great. So the four six toggle, for that thinking was terrible. So, big props to Anthropic for bringing it back, but they brought it back in a way that might have you burn through your token usage. Right? So just keep that in mind. They are not kidding when they say it uses more of your limits. Also, Anthropic says, Opus 4.8 is around four times less likely than its predecessor to allow flaws in code. It has written to pass unremarked reflecting an emphasis on what they call honesty and flagging uncertainty.
Jordan Wilson [00:30:58]:
So, I am not too certain about that in my very little experience, but, I did watch a couple review videos from people that had early access, and they called this out as well. Normally, people that are very, you know, proanthropic called this out. The honesty thing and hallucinations didn't seem to check out from some, you know, initial vibe test. So we will see once it goes through all of the benchmarks where it lands on that. So, who has access? Well, everyone. So, the cool thing is it is marked at the same price as Opus four four point seven. So I think this is, like, the first time in a very long time that Entropic has upgraded the model without an upgrade in price. However, it is token inefficient AF.
Jordan Wilson [00:31:46]:
Right. There's a lot of very helpful charts if you're, you you know, looking at what models to use. I highly recommend looking at artificial analysis, and they have a great intelligence, per, kind of cost. And, yes, Opus does great. Opus 4.8 does great on benchmarks, but it is extremely token inefficient to essentially achieve the level, of intelligence that it does. It uses way more tokens than anyone else, and it is not even close. And then, you know, add in these, you know, these, quote, unquote, new reasoning levels and yeah. To get that, you know, to pass the quality bar for whatever artifact you're working on, whatever output you're trying to create, in Claude, it's gonna cost a ton more.
Jordan Wilson [00:32:36]:
So it's obviously available system wide, on the workbench, in Claude on the web, in Claude code, all that good stuff. So, here is what Anthropic says about Opus 4.8. So they say, we're upgrading Claude Opus to a new version, Claude Opus 4.8. It builds on Opus 4.7 with improvements and cross benchmarks and is a more effective collaborator. Opus four eight launches alongside several new features, which we already talked about. One of them, users on cloud dot a I now have control over the amount of effort cloud puts into a task. Cloud Code has a new dynamic workflows feature that we already covered that allows it to tackle very large scale problems. And there is also, as well, a new fast mode for Opus 4.8 where the model can work at 2.5 the speed.
Jordan Wilson [00:33:31]:
I I like how they phrase this. They said it's now three times cheaper than it was for previous models. So, yeah, it used to be a six x cost. So now it's a two x cost. So, yeah, don't let that, like, oh, it's cheaper. Oh my gosh. Let me do this two x thing, or this 2.5 speed thing. Yeah.
Jordan Wilson [00:33:52]:
It's double the cost. Alright. But the capabilities, yes. They are impressive. So the, benchmarks, for the most part, at least on the ones that Anthropic handpicked, it outperforms. They only included five, no, six benchmarks here, and it is tops except, interestingly enough, on terminal bench, which is the one that they should be winning on, which is agentic terminal coding. Right, g p d 5.5 is still very far ahead on that one. And I do have a more comprehensive benchmark list.
Jordan Wilson [00:34:27]:
But as you'll see from my list, and there's actually two others that I left off this list, where g p d 5.5 was winning. So that's why I'm like, okay. What's the best model in the world? I mean, if you look at artificial analysis, the intelligence index, which I think is the best indicator, not everything is fully out yet, for Opus 4.8, but it does look like it will come in ahead of GBT 5.5, by a slight to moderate margin. But here's the thing. At any point, right, we're we're hearing, that GBT 5.6 could be here any day. The normal codex Thursday release was pushed back, to today. So, maybe by the time you read today's newsletter, we'll see what new, things that we have from OpenAI via codex, or maybe they'll, sneak in 5.6. Not sure.
Jordan Wilson [00:35:21]:
But regardless, we do know that, OpenAI's next model is around the corner. Google already said that their Gemini 3.5 pro is going to be released soon. So interestingly enough here, Infropic strategy, not sure how it's gonna work. Obviously, they have Mythos around the corner, and they did say that that was going to start rolling out in the coming weeks. Right? So whether that's two weeks or ten, we'll see. But the assumption on how this is gonna play out, OpenAI and Google are gonna come out with models that top Opus 4.8 in the next few weeks. And then once they do that, Entropic will then release Mythos. And my thought is, that Mythos will probably have a decent lead at least on most benchmarks.
Jordan Wilson [00:36:06]:
Although it's already even though it's not available, there's already, I think, three benchmarks that Mythos has been passed on even though it's not publicly available. Right? When it came out, you know, everyone's like, oh my gosh. Mythos is is gonna be the benchmark king forever. It's not even released yet, and some of their reported benchmarks have already been surpassed. Anyways, I do expect that to be the case. So at least for king of the hill when it comes to models only, Anthropic probably took the lead back with Opus 4.8. We're gonna see OpenAI take the lead back with five six. I assume Google will recapture the lead, with three five pro in June, and then probably shortly thereafter, we'll have Mythos.
Jordan Wilson [00:36:48]:
And I think Mythos might hold on to it, for, you know, maybe two months. Who knows? Alright. My first impressions with, Opus 4.8, really good in some respects. Absolutely terrible in others. Here's what I mean. And the stuff that you would expect, a a Claude model to be good at, it is really good. Specifically, front end design, amazing. You know, creating, fully functioning games, apps, websites, those things, you know, HTML designs, whatever.
Jordan Wilson [00:37:20]:
So so good. Certain things that I like to use these models for and, you know, I will call out, you know, another benchmark that it is behind on. And I maybe that is one of the reasons on MCP Atlas, which is multi step workflow orchestration tools via the model context protocol, which is interesting considering considering, Anthropic invented this, and they're behind, Gemini 3.5 Flash. Right? So that's one it's behind on, and it's behind on browse comp as well. And it is the worst model actually on browse comp. And that's a model's ability to research questions by searching the live web. And my gosh, did I ever find that out by, you you know, using, Opus, a little bit today. Actually, in prep or, yesterday in preparation for the show.
Jordan Wilson [00:38:11]:
And not just that, but this this new thing where Anthropic says it's, you know, really prioritizing honesty and refusals, which I actually found absolutely infuriating. So I think it was a combination. Right? I told, Opus four eight to do some research on a transcript from one of my shows. I know that I mentioned the different tiers of risk for the EU AI act, and I wanted to see what those were. And it essentially said, nope. Not gonna search the web. Right? Because these tiers don't exist. And I'm like, yeah.
Jordan Wilson [00:38:42]:
It does blah blah blah. Right? So I had to tell it three different times. You you know, and it's and and and it was using this new honesty language and how, you know, it's not gonna tell me something, you know, that's that's not factual. So, you know, it's it's at least for me, in in my very limited use, not loving the new, four eight. I know this is weird. I think I'm still using four six. I really like Ocus four six. I think it's a great model.
Jordan Wilson [00:39:10]:
Just this new honesty thing and, you know, there's some other quirks, but, you know, mainly, the even on a $200 a month plan, you I can't use it how I want to. Right? That's that's the reality. I can't really use any cloud models the way I want to on a $200 a month plan, but maybe that's just the new reality. Right? As we, you know, shift from token maxing to token efficiency, right, maybe even people like myself who have seemingly unlimited, you you know, budgets at least to spend on the subscription side are gonna have to cut back a little bit. Alright. So that's it. Let me know. Should we do an Opus four eight show dedicated, next week? Let me know.
Jordan Wilson [00:39:46]:
If so, you know, what if you're still listening at this point, that's how I know. Right? Sometimes I leave little things at the end. But just put Opus four eight in the comments on Spotify, in the comments on, LinkedIn. I got my arbitrary number. You know, if if we get that numb that many, I'll I'll do a dedicated show, and I'll take my time and put it through the, put it through the ringer for you guys. So that's it. That's a wrap. A lot of new things, you know, from chat g p t agents to Copilot Studio, you know, agents that can use your computer to small little things like Google Drive and notebook l m syncing.
Jordan Wilson [00:40:21]:
I think that there's a lot of new capabilities that were unlocked, and you probably had no clue. I hope this was helpful. If so, make sure you go to youreverydayai.com. Sign up for the free daily newsletter. And remember, if you are in San Francisco, check out the show notes. I always leave my link to my LinkedIn. Hit me up if you wanna chat AI, whatever it is, or if you're gonna be at the build conference, let me know. Like I said, I always put my LinkedIn link in the show notes.
Jordan Wilson [00:40:48]:
Make sure to tell me you're from the podcast, though. Otherwise, I'm gonna assume stranger danger. Thank you for tuning in. Hope to see you back next week and every day for more everyday AI. Thanks, y'all.
