Resources:
Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Start Here Series in our Inner Circle Community: Join for free access
AI Productivity Advances: Seven Features Business Leaders Can Implement Now
The latest episode of Everyday AI provided actionable intelligence on a slate of AI advancements released within a compressed timeframe—each available for immediate deployment by organizations. This article dissects each feature’s business utility, clarifying where unique value emerges and how operational workflows stand to benefit.
ChatGPT Health Integration: Unified Medical and Wellness Data Analysis
OpenAI has rolled out a dedicated health integration directly into ChatGPT, available to all U.S. users across free and paid tiers and supporting iOS. This enables individuals to connect Apple Health and medical records, centralizing insights—such as tracking lab result trends, monitoring medication compliance, and summarizing longitudinal health data—within a siloed, privacy-oriented space.
From an operational perspective, this eliminates data silos and the need for multi-portal logins, allowing managers and HR professionals personalized dashboards for wellness initiatives and more effective care coordination for chronic health matters, while explicitly stating that diagnostic and treatment functionalities are not supported. The utility expands to pre-appointment preparation, dietary recommendations for restaurants, and summarization of bloodwork for actionable meetings 04:35.
Claude Voice Mode Expansion: Enhanced Productivity for Power Users and Executives
Anthropic’s Claude has extended voice mode functionality to its efficient Sonnet and advanced Opus models. Now, those using paid accounts can operate hands-free with connectors enabled, integrating live data streams into AI-guided workflows.
Productivity gains materialize for executives multitasking during travel or meetings, and for operators seeking streamlined hands-free ideation. Unlike basic walkie-talkie interactions, this upgrade allows switching models mid-conversation and automated connector-driven insights, though the interface remains turn-based rather than real-time bi-directional like some competing solutions 09:15.
Microsoft MAI Image 2.5 Pro: High-Fidelity Visual Asset Generation
Visual content creators within organizations now have access to Microsoft’s MAI Image 2.5 Pro, a high-precision text-to-image generator embedded in Microsoft 365 Copilot, PowerPoint, Foundry, and APIs. The model delivers controllable edits, object removal, and in-image text rendering, with pricing at $106 per 1,000,000 image output tokens and lower costs for inputs.
This tool is especially relevant for organizations operating solely within the Microsoft ecosystem, offering a step-change over earlier image models for branded presentations, product shots, and marketing collateral—without the compliance concerns tied to non-Microsoft tools 13:12. While quality may not rival leading alternatives, the improved integration directly within PowerPoint eliminates unnecessary software switching and preserves workflow continuity.
Gemini 3.6 Flash and Flashlight: AI Cost Efficiency for Development-Heavy Teams
Google’s release of Gemini 3.6 Flash and the even faster 3.5 Flashlight model focuses on token efficiency, delivering up to 17% lower token consumption compared to predecessor models. The key appeal for organizations is lowered operating cost for agentic workflows such as bulk document processing and complex search, especially where high throughput and low latency are essential 16:17.
The Gemini 3.5 Cyber variant caters to enterprise cybersecurity contexts, though with more restricted access. As high-volume automation becomes synonymous with cost savings, these developments furnish IT teams and business process leaders with tangible economic returns 18:20.
Gemini Spark Agent: Multi-Step Task Automation and Workflow Synchronization
Gemini Spark, Google’s AI-driven personal agent, is now broadly accessible to professional subscribers, automating multi-step business tasks across the Google Workspace suite. It performs research, composite workflow coordination, and direct edits to Docs, with extensions planned for Sheets and Slides. Automation spans tasks such as travel booking management, real-time monitoring for scheduling conflicts, and auto-generation of actionable updates across integrated apps 22:23.
The absence of immediate workspace integration limits broad business utility, yet teams with personal Pro accounts access unparalleled automation within the Google tech stack, minimizing manual transfers between tools and accelerating project management cycles 24:10.
Claude Cowork “Record a Skill”: Transforming Process Knowledge into AI-Driven Automation
Anthropic’s Claude now enables users to screen record workflow processes and converts these recordings—complete with keystrokes, mouse movements, and narrations—into shareable, repeatable skills. Unlike comparable products, voice narration is captured natively, lowering the barrier to translating complex tacit knowledge into structured agentic automation 26:12.
For operational leaders, this directly supports SOP documentation, onboarding, and repetitive process automation. Recorded processes can be shared or ported between Claude, Google, OpenAI, or Microsoft ecosystems, unlocking interoperability for standardized tasks and knowledge capture 30:18.
ChatGPT Voice on Desktop: Natural Language-Driven Computer Control
OpenAI’s desktop deployment of ChatGPT Voice, powered by GPT Live, provides full computer control via natural language—enabling voice commands to direct agents, files, programs, and cross-tool coordination. The system works with “app shots,” a functionality that translates screenshots and app context into actionable instructions, allowing the AI to reference and act on information across active windows and documents that extend beyond visible screen real estate 33:38.
This unlocks a new interface paradigm for knowledge workers—enabling hands-free direction of workflows, error identification, report generation, and data extraction. Integration with remote control means workers can direct desktops from anywhere, enhancing flexibility for hybrid and travel-intensive teams 32:07.
Each of these features, all released and available for immediate use, provides quantifiable workflow optimization targeting discrete business functions: wellness program oversight, executive productivity, visual content production, DevOps cost control, process automation, and natural language desktop orchestration. The convergence of these tools signals a new competitive battleground where speed, interoperability, and seamless human-machine collaboration offer concrete operational and strategic value.
Topics Covered in This Episode:
- ChatGPT Health Syncs Apple and Medical Data
- Claude Voice Mode Adds Opus and Sonnet
- Claude Voice Mode Supports Connectors
- Microsoft MAI Image 2.5 Pro Launch Details
- Microsoft MAI Image Model Benchmark Preview
- Google Gemini 3.6 Flash and Flashlight Release
- Gemini 3.6 Flash: Token Efficiency Upgrades
- Google Gemini Spark Agent for Task Automation
- Claude Cowork "Record a Skill" With Voice Narration
- ChatGPT Voice on Desktop: Full Jarvis Mode
- ChatGPT Voice Controls Apps via App Shots
- Cross-Platform AI Skills Sharing (Claude, Codex, GPT)
Episode Transcript
Jordan Wilson [00:00:16]:
If you listen to our Monday AI news that matters segment, I told you we were kind of in for a wild week. That's because after a somewhat slowish week last week, I said on Monday that this week was going to be a banger in terms of new AI features you can actually use. Sure. I knew ahead of time a few things that were gonna come out, but even I was kind of blown away just yesterday at the sheer amount of new updates that all dropped within hours of each other. I mean, no joke. On Thursday alone, we got huge new AI features from Google, Microsoft, Anthropic, and OpenAI in a matter of hours. That's why on Fridays, we break down what's new in AI that you can actually use, and we do it in a very specific way. We give you the AI features that are available for you today, not rumors, leaks, or wait list.
Jordan Wilson [00:01:14]:
These are instant AI upgrades that you should be putting into play today, and we've got a lot to cover on today's show. So stick with me for the next twenty five ish minutes, and I'm gonna tell you how chat g b t went full Jarvis mode and the new app shot secret that unlocks the future of work. I'm gonna let you know the new Claude feature that took a page out of the codex playbook, but actually made it a little bit better. And I'm gonna fill you in on the kind of new Google agent that you can start using today, but with one caveat. Alright. Let's get into it. Welcome to Everyday AI. My name is Jordan Wilson, and we do this every day.
Jordan Wilson [00:01:57]:
It's your unedited, unscripted daily livestream podcast and free daily newsletter helping business leaders like you and me keep up with the avalanche of AI news and updates. I tell you what matters, how to use it. You take that information. You're the smartest person in AI, and you can grow your company and your career. So it starts here, but make sure you go to the website at youreverydayai.com. We're gonna be recapping the highlights from today's show. So if you miss any of the details, maybe you're out on a walk and you wanna go back and say, wait. How did this work? It's all gonna be in the newsletter as well as all of the other AI updates you need to know.
Jordan Wilson [00:02:34]:
Alright. Let's get straight into it. I'm a big fan of the Friday features show. I wish I would have started it, like, three and a half years ago. We've been doing it now for about six months. But just for our podcast audience, I'm always showing, on my screen some basic, you know, information from the company. You're not missing anything. So, just FYI.
Jordan Wilson [00:02:54]:
But let's dive in live. I'm gonna be sharing my screen, just kind of showing the, windows for the different announcements that we're gonna be going over. Alright. So let's start first with your health. Yeah. This is actually, one I'm personally excited about. I know some people aren't gonna be, you know, diving in headfirst necessarily to give their health data to anyone, let alone an AI company. For me, I'm absolutely wildly gonna be using this.
Jordan Wilson [00:03:28]:
So here's what's new. ChattGPT Health is finally rolled out to everyone. It is launched for all US users today, and it's expanding the feature that was first introduced in January, but it was introduced as a wait list, feature, and it connects your Apple health and supported medical records directly into chat GPT. So what the heck can it do? Why is it useful? Well, it can compare your lab results, summarize changes since a prior appointment, track your medications, factoring your sleep activity and workouts, essentially, anything that, you know, your Apple health might track and all the different applications that connect in there. Well, it can bring in that data and anything else that you can update or upload into chat should be help. So OpenAI says that health conversations obviously live in a dedicated space siloed from all your other chats. Right? So, yeah, you're not gonna have all of that information popping up via, like, a memory, into your other chats. So like I said, this is already rolling out to US users.
Jordan Wilson [00:04:35]:
You do have to be 18 or older, but the cool thing is this is available on free plans, on paid plans, and also on iOS. So one small caveat though, free users, you're just gonna be using kind of the model that you have the most access to, which is GPD 5.5 instant. And then for other paid users, you'll be able to use the more powerful GPD five six sold. Alright. So why is this useful? Right? Because right now, obviously, there's no one place for all of your health information to live. Like, I can't tell you the amount of conversations I've had with people about this over the last few years. And, you know, there was a company out there that was kind of, you know, AI native. They were called Forward Health.
Jordan Wilson [00:05:22]:
They are no more. And I was kind of excited about that concept because I'm like, there needs to be, you know, something that's not tied to a single provider that allows you to just essentially bring in all your health records and to have a smart system that just knows all that. So, obviously, there is nothing previously that would have stopped you from doing this inside of a large language model like Claude or Chad GPT, Gemini, Copilot, etcetera, but it wasn't really set up for this. Right? So think of this as a specialized version of ChattGPT that's siloed from everything else. So it keeps all of your health data, separate, but it is literally built for this. So, you know, right now, your health info is scattered across different patient portals, EMRs, apps, and it just pulls it all into the conversational layer. And there's also some kind of practical prompts baked up for you. Right? So how your cholesterol might be trending or how you can summarize your blood work before an appointment, what you should ask your doctor tomorrow, etcetera.
Jordan Wilson [00:06:22]:
And that context kind of bleeds, usefully into everyday tasks like factoring a dietary restriction into your restaurant picks. So who's gonna find this valuable? Well, this is obviously a consumer play, but I think this just goes to show where OpenAI has started to focus more recently. Right? Not necessarily just on the models, but how to make their models more useful for more people in more ways. So, you know, whether you are just someone that wants to get a little bit more serious about your health, you just have questions, or if you're managing chronic conditions, if you're a caregiver trying to track something in your family health or just preparing for your upcoming appointment. So, obviously, OpenAI says this is not intended for diagnosis or treatments, and health conversations aren't used to train foundation models. So pretty big one. I'm excited to get this. I signed up for the, the wait list.
Jordan Wilson [00:07:23]:
Previously wasn't, wasn't kind of, selected to be part of that first group. So this is one I'm gonna be jumping into head first. Alright. Our next AI update, a new update well, new ish feature from Anthropic, just kind of borrowing from the codex playbook, but I think they made it a little bit better. So Claude has some new voice mode upgrades. Oh, no. This is not the one. Sorry.
Jordan Wilson [00:07:48]:
That's a little bit, better than, chat g b t's. So this one is voice mode. Not quite as good, but also this one is a little different. Yeah. Just so many new updates. Even I'm getting confused, and this is all I do every single day. Alright. So this new update from InfraBiq is essentially expanding how and where their voice mode can be used.
Jordan Wilson [00:08:17]:
So Infropic has upgraded the Claude voice mode to run on Opus and Sonnet for the first time. That's because previously, it ran only on their least powerful model, which is Haiku, which is why for me, I absolutely never used voice mode, inside Claude on my phone for that very reason. But now I probably will. The cool thing as well is you can switch models mid conversation in voice mode now works with connectors. So that is the big difference right now with the, the one thing that's maybe a little bit better is that their upgraded voice mode works with connectors, but it's not bidirectional. So it's still kind of the, quote, unquote, older dumber, version of a voice mode. So you you've heard me kind of rave about the new GPT live. That's because GPT live can listen to you, and speak at the same time.
Jordan Wilson [00:09:15]:
Right? You can interrupt it. It might interrupt you, and it's kind of listening, thinking, and doing all these things at the same time. So this is still the old school walkie talkie mode where you talk, you wait, Claude responds. But the big benefit here is it can use connectors. So pulling in your live data. So that's not something that GPT live can do, although you can individually upload files into GPT live, and OpenAI did say that they are working on bringing connectors. So here's who has access. You have to have a paid Claude account to get that upgraded Sonnet in Opus voice mode.
Jordan Wilson [00:09:52]:
Otherwise, you will be defaulted to Haiku. So free accounts do have a very limited usage and only a single connection, and all prompts then obviously go through Haiku. And right now, voice mode does support 11 languages. So why is this useful? Well, if you want to be able to chat with your connectors and you have a paid quad plan, this is really good, especially if you don't mind kind of working in walkie talkie mode. For me, I've really enjoyed kind of the, the bidirectional or duplex, new way of talking. It seems like that's the future of the voice interface. So So as long as you don't mind kind of this more kind of waiting and holding and, you know, only being able to talk or listen at once, not that bad. So who's gonna find this useful? I mean, anyone that's a power, user of Claude, if you're doing any hands free work, commuters, you know, executives between meetings, or just operators who want to think out loud, pretty big update and release here from Anthropic.
Jordan Wilson [00:10:58]:
Alright. Next, we have a yes. Another new image mode from Microsoft. Alright. So here's what's new in MAI images or sorry, MAI image 2.5 pro. All these image models are always a mouthful. Right? We should just I don't know. Microsoft get a fun, like, nano banana, name, and then we can call it you know, why not, like, Clippy, Clippy Pro, you know, 2.5? Alright.
Jordan Wilson [00:11:28]:
So this is launch, and it's launched in a lot of different places. So it is in boundary, preview. It's in co it's in Microsoft Copilot PowerPoints, and it's in also a variety of other places. But the new image mode, it's pretty good. I don't think we have a lot of benchmarks on it just yet, but I would assume that we get those probably, within a couple of days. So this is Microsoft's newest highest fidelity professional grade image model built, they say, for superior high quality imagery, detailed editing, and precise in image text rendering. So this does, Microsoft says, handles text to image generation plus controllable edits, object removal, replacement in painting, text updates, all while preserving composition. So who has access? So, yeah, it is now available in Microsoft Foundry.
Jordan Wilson [00:12:28]:
If you do have a Microsoft, three sixty five copilot plan, you can use this right now inside of PowerPoints, or you can use it via the API. The pricing is $5, per 1,000,000 text input tokens or $8 per 1,000,000 image, input tokens. Alright. And then, $106 per 1,000,000 image output tokens. So why is this useful? Well, I'll say this. Right now, if you're one of those organizations that can only use Microsoft Copilot and you can't use anything else, then this is great. This is a nice upgrade over the previous MAI image two point o, but it's still presumably right. We'll see what the benchmarks say.
Jordan Wilson [00:13:12]:
At least from my trained eye, this is still fairly far behind, GPT images too, which is the best AI image, you know, model in the world and also the Nano Banana two. So those are kind of the top two, and then you have some, some more, you know, Chinese open, versions of these. But the best two still are GPT images and Google's Nano Banana. So not really in the same tier. We'll see what the benchmarks say, at least by my kind of quick trained eye. I still think that there's probably a quality drop off. But regardless, if you are someone, especially, I think, in if you are building decks in PowerPoint via Copilot, being able to use this image generator in there is gonna be big. Right? So yeah.
Jordan Wilson [00:14:00]:
I know that so many people out there are kind of living inside of PowerPoints and maybe your visuals and the old clip art just won't do. I think this is gonna be pretty big. So, you know, if you're trying to create product imagery, marketing visuals, brand assets, big. It is big for those purposes when you can't use anything else. Alright. Our next AI update you can use today. Yeah. It was this big of a week that I didn't even really honestly get a chance to play with a brand new model from Google Gemini.
Jordan Wilson [00:14:35]:
Aside from just some daily driving tasks, I haven't really given the new Gemini 3.6 flash a run. So here is what's new with Gemini 3.6 flash. There's actually three different models, but I say most people are focused on Gemini 3.6 flash, but they also unleashed, Gemini 3.5 flashlight. I know. Confusing. So they released both, the 3.5 upgrades and updates and new variations in the 3.6 flash, which is the first of the 3.6 flash series. Even more confusing, the pro series is still stuck on 3.1. So, yeah, now your latest and greatest Google Gemini models technically have three different step changes.
Jordan Wilson [00:15:21]:
You have 3.1 pro because we are still waiting for 3.5 or maybe they'll skip straight to 3.6 pro. So you have 3.1 pro, then you have the 3.5, flashlight in 3.5 flash cyber, and then the only one on the 3.6 tier is flash. So extremely confusing. They're they're all for different purposes, but let's break down at least, what's new in these models. So, Gemini 3.6 Flash. Essentially, not that much smarter per se, but it is cheaper and more token efficient. So Google says that Gemini 3.6 flash consumes 17% fewer tokens than 3.5 flash while taking fewer reasoning steps in tool calls in multistep workflows. Then you have 3.5 flash light, which is the cheaper and faster version of flash.
Jordan Wilson [00:16:17]:
Alright? So flash is like your, you know, cheaper version, and then flashlight is the cheaper and faster and more lightweight version. So the 3.5 flashlight is the fastest model in the 3.5 series at 350 output tokens per second built for low latency, high throughput work like agentic search, and document processing. And then you also have that third model, Gemini 3.5 cyber, which is tuned specifically for cybersecurity tasks. So who has access? Well, everyone. If you're a paid user, you can go into Gemini. You know, your gemini.google.com, the Gemini API, inside AI Studio, Android Studio, anti gravity, Gemini enterprise. Right? Literally, wherever you're using Google Gemini for the most part, you can find the new 3.6 Flash there. Flash Cyber though is limited access only.
Jordan Wilson [00:17:11]:
That's the, you know, government and trusted partners similar to, you know, a a Mythos or OpenAI has their, Daybreak, I believe, is that what their version is called. So why is this useful and who should be using it? Well, if you are a developer, especially if you live inside of the Google ecosystem, if you're working on, you know, just document processing in bulk, if you're working on, any agentic tasks, that's where it's gonna pay off. Right? It's cheaper and more token efficient. You pay less per token and burn fewer tokens in the process, and that's kind of the, agentic economics one zero one story there. So 3.6 does also 3.6 flash does deliver higher precision coding with fewer unwanted edits in reduction, reduction in loops, essentially. Right? So if you're a dev running production agents inside Gemini or if you found the previous version of 3.5 flash to work in terms of price per performance, I think that 3.6 flash, although it's not a big bump in terms of capabilities, it is just a little faster and cheaper and more efficient. So, any developers that have already been using the flash series is gonna be good, for you. And then also Flashlight, important here, will power the future Google search AI overviews where speed really matters the most.
Jordan Wilson [00:18:34]:
So the big news here is, well, we still don't have a new pro version. Yeah. We were, kind of reports this week. If you follow our newsletter, we covered this. Said that the pro version may get delayed a few more months, and then we similarly, got news from, frequent, open a, frequent, everyday AI guest, Logan Kilpatrick, found the show a handful of times, did say that Google is pretraining Gemini four. So who knows? Maybe we technically won't get a Gemini 3.5 pro or a 3.6 pro. Maybe we'll just be in for a longer wait, and we'll get a Gemini four pro. Not sure, but Google did say that they're pretraining Gemini four, and we saw reports that the pro series is delayed.
Jordan Wilson [00:19:24]:
So if you put two and two together, maybe that just means there will be no 3.5 or 3.6 pro, and we'll just be going straight to four pro. We'll see. Alright. Let's go to our next one. And, yeah, this one, I'm both excited, but a little bummed. But this is kind of how Google rolls out some of their products. Maybe I'm just being, a little snobbish in, you know, wanting the best tools for where I actually need them. But, yes, we finally have a new twenty four seven personal agent from Google that just works.
Jordan Wilson [00:20:02]:
I'm talking about Gemini Spark. So not technically new, but new to most everyone, because this did launch in May, but only those two, on the Ultra plan, which is that 200 a month plan, from Google Gemini. So we didn't cover it on the show in May. Right? For the most part, I only cover, on the Friday features, those that come out to the base paid plans. Right? Your $20 a month plan, not your plans that are $102,100 dollars a month. So that was launched for Ultra subscribers in May. But now Google has released, Gemini Spark to all pro subscribers as well. So that's Google's personal AI agent for automating complex tasks.
Jordan Wilson [00:20:50]:
It's rolling out. I have access to it right now. So it performs multi step tasks on your behalf, research, planning, complex workflows while you focus on, well, whatever else. So it's kind of built on these three different concepts. So you have your tasks, which is your high level goal. You have your skills, right, that open skills standards. Those are your, reusable instructions with context. And then schedules.
Jordan Wilson [00:21:15]:
So those are automated triggers like times or conditions. Alright. So like I said, if you are on that base, paid plan, you should get access to it now, if especially if you're in The US. Google did say that other countries are going to be getting access to it soon. So here's the downside or the caveat. I have no clue when it's gonna be rolling out to workspace plans. This is usually one of those things that they say soon. So if your business runs on Google, on a paid Google Gemini account, maybe you have access to it, maybe you don't.
Jordan Wilson [00:21:50]:
But if you are on a standard paid, you you know, Google Gemini plan like me. Right? So my personal Gmail, I have a, paid Gemini account for my personal Gmail. I have a paid Gemini account for my work emails. So on the work side, we don't have this. So for me, it's like, well, I don't know how much I'm gonna be using Gemini Spark in my personal Gmail. Right? I just use it more for testing and for certain purposes in Gemini that I don't have access to in my workspace account. So why is it useful? Well, it works inside of the Google apps. Right? So, any axe, any app that you give Spark access to inside of the Gemini workspace, it can essentially go and, you know, monitor for updates in those apps, you know, push, updates from, you know, docs to slides or from your Gmail to your calendar.
Jordan Wilson [00:22:45]:
Right? So this also works, with your personal intelligence, that new feature from Google. There's a remote browser, and computer with code execution, chats, Canvas, all those different things. So an example that Google gave is, you know, saying, Google Doc, The task is to plan and manage your business trip coming up. It'll look through your, you know, your Gmail, your calendar, and your schedule, and it will tell you, oh, your flight is delayed or, you know, a travel booking skill, you know, plus a Gmail writing skill, you can rebook the room and send a confirmation or something like that. So Spark can now open, read, and edit Google Docs directly with that expanding to spreadsheets and presentations, including reading comments and editing shared team docs. So I think Gemini Spark was previewed a super long time ago. And at the time, I was personally very excited, but it took, obviously, many months for it to actually roll out, and it's still, at least according to my research, not available for every single paid workspace plan. So, obviously, this is great if you run, Gemini in your business, but not all business paid accounts have access to this.
Jordan Wilson [00:24:07]:
All pro accounts on the personal Gmail side do. So if I'm being honest, six months ago, if this rolled out to my workspace account, I would be talking about GeminiSpark every single hour because it's great. But right now, you can do most of these things anyways, with codex or quad code, quad desktop. So that's why I'm not, like, gonna be, you know, raving about Gemini Spark. You know, it's kind of interesting that in a lot of cases, OpenAI, Anthropic, and the crazy thing is even sometimes Microsoft, offers some of these features and functionalities before Google does with their own products. Right? So even Microsoft rolled out a, essentially, an OpenClaw ask integration, and you can connect with Google. So, although it's exciting, technically, Google a little bit behind even in their own backyard. Although, I do know that millions of people will like this update, and that's why I decided to cover it on the show.
Jordan Wilson [00:25:10]:
Alright. Two more, and we are getting to our big one. But before we talk about how, Chad GbT can now be your Jarvis, we have to talk about the this is the one where, Anthropic took a, page out of the codex playbook and made it a little bit better. That is the new record a skill feature, that just shipped this week in Claude Cowork. So this is essentially, you can screen record yourself doing a task once, and then Claude analyzes the recording and turns it into a reusable, re, runnable skill. So Claude captures screen activity, cursor movements, keystrokes, and the new addition that we don't have just yet in codex, at least not by default, is your voice narration. So then, it processes all of that together, in the recording, and it turns it into a structured skill that saves it to your library. So who has access? It is rolled out to all paid plans now.
Jordan Wilson [00:26:12]:
So, free users can still use existing skills. Right? But you can't obviously record new ones. So it was a little confusing to find this. So it's only available on co work. So you have to go into the chat interface. So now, obviously, Claude desktop changed their interface a little bit. It used to be the three different tabs. It was chat, cowork, and code.
Jordan Wilson [00:26:38]:
Now it says chat and code, but in chat, there's a cowork tab. Right? So you have to be in chat. You have to go find your cowork tab. You hit the plus button, and then you can find, the new record a skill. So it actually took me a little bit of clicks to find it. I'm like, wait. Where is this? So you do need to update this. This is desktop only.
Jordan Wilson [00:26:56]:
So update your Claude desktop app. Go into chat. Click the cowork tab. Click the plus button, and then and actually, super useful. So I did a little bit of testing. It's really good. So why is it useful? Well, until now creating a skill meant either chatting with Claude and having it create one for you. You know, you can obviously create those skill MD files, by hand.
Jordan Wilson [00:27:21]:
But so many times, it's just like at the end, and and and you should be doing this, FYI, if you're not already doing this. The way I look at new chats or tasks in any, you you know, large language model, especially any agentic one. When I'm done and I get an output where I'm like, this is good, you should be turning that into a skill. Right? So this is instead of having to work through it or to hand hold, right, Claude or Kodak's or anti gravity or anything, right, at the end, you would normally say turn this into a skill. So the difference here with the teach Claude a skill is, well, you don't have to work through it. You just click start recording. You do your work. You dictate through it.
Jordan Wilson [00:28:04]:
You say, here's what I'm doing, a, b, and c. I'm opening this site. Here's what I'm looking for on this page. I download it. I open it in this program. Right? And it'll just record all of that and turn it into a skill. So like I said, technically, taking a, page out of the codex, replay, or no, record and replay skill, that did this exact same thing. But the downside with codexes is by default, there is no voice narration.
Jordan Wilson [00:28:32]:
Right? There is kind of some workarounds, around that, but it is nice that in this new one for Claude Cowork, it is just enabled by default. You don't have to have any workarounds to actually narrate your way through that skill, which is very helpful, obviously. So who's gonna find this valuable? Well, literally anyone. So if you're a cloud desktop user and you use cloud co work, it's gonna be great. I would love to see, Anthropic roll this out for Cloud Code as well. I think that would be really helpful, just because seemingly, I'm guessing Cloud Cowork, is not kind of riding that same wave of of popularity that it was in earlier, 2026 just because the the the interface kind of got changed over a little bit. And I think more and more people are starting to use Claude Code and just more of the agentic features and the normal chat section. So you can only set this up in the co work, which is a downside.
Jordan Wilson [00:29:30]:
But literally, anyone that uses Claude and you're doing the same processes over and over, I mean, this is big. Right? I love I absolutely loved and use the this codec skill, all the time for my workflows. So I'm personally happy, to see this in, quad code. The other thing, here's the here's the big unlock y'all. You can do this in I talked about this in codex as well. You can do this in Claude Code or codex, go through, record the skill, and skills are shareable, and they're fairly, openly supported across all different platforms. So you can just create a skill, very detailed skill as an example using this new feature from Anthropic, and then use that skill inside of Google or use that skill inside of codecs or chat g p t. Alright.
Jordan Wilson [00:30:19]:
And our last big feature in y'all, this this one, this is one of those instances where I'm like, okay. This is crazy after I tried, this. So, yeah, hasn't even been out a full twenty four hours, but OpenAI went to full Jarvis mode, and they launched chat GPT Voice on desktop. And it's a lot more than it sounds like, but let me go through the bullet point details first. So OpenAI launched chat GPT Voice in the desktop app, powered by GPT Live, which I already referenced, that lets you talk through work and coordinate tasks across chat, work, and codex. So per OpenAI's announcement, you can control your computer and direct multiple agents running in chat g b t work or codex using just your voice with g b t live speaking, listening, and coordinating work simultaneously. So here is how open AI kind of describes the main benefits. So it says you can work across projects.
Jordan Wilson [00:31:21]:
You can start a new task, coordinate work in progress, and pick up where you left off without micromanaging. You can work across your desktop. You can use your files, apps, and connected tools from Slack and GitHub to Notion and beyond, and you can explore, plan, and learn. You can talk through a poll request, ask questions about a code base, or learn a new topic through natural back and forth. So who has access? Well, if you are on any paid plan and you're using the desktop version of chat g p t work slash codex, it's the same thing. It is out now. And it works also with, chat g p t's remote feature, which is pretty cool because you don't even need to be in front of your computer. You can just have your phone.
Jordan Wilson [00:32:07]:
Right? I have one, right. I'm traveling right now. I'm not in Chicago, but I have essentially a home Mac studio that never gets shut off, and I can be controlling it right now in just with my voice, and it can be doing literally anything, and I can be watching and seeing it go in real time. So this just kind of brings that, GBT live mode that was announced about two weeks ago, to the desktop. So before, that was only available via mobile. So it doesn't just bring it to the desktop. But this is the first time where I'm like, wait. You technically don't even need to type or use your mouse anymore.
Jordan Wilson [00:32:49]:
Then, yes, I haven't been able to run it through many hours of demos, but, the possibilities for this one are absolutely bonkers y'all, and it is so good. So let me talk about a a couple things. Number one, it works with Chat GPT's remote feature, which is great. That is on your mobile app. You can click remote, and you can technically, control any computer that you have connected to that chat g b t account. So that's cool. But here's the other thing, app shots. Alright? So what app shots are, if you're, you know, a power, you know, codex or chat g b t work user, The default, you know, is the two command keys, and it takes essentially a screenshot, but not just a screenshot, but every single thing it brings into context that you have, in that program.
Jordan Wilson [00:33:38]:
So as an example, right now on my screen, I have up some notes, some bullet points that I wanted to go over. Right? But it's a very long list. So, obviously, what's on my screen is probably only, like, 5% of what's in that document. So if I use the app shots, it's gonna obviously take a screenshot of what is actually on my screen, but it brings in the context of everything else in, that program that I have up that is not even on the screen. So the cool thing is with this new, you know, I kinda wish we coulda got a fun name for it. I know they can't use Jarvis. I'm just using Jarvis. Right? Like, you can have this app shot feature and functionality while you're using chat g p t, live voice on desktop.
Jordan Wilson [00:34:22]:
So in my kind of playing around, I had, you know, six or seven different programs open on my computer, hands free, not typing, and just telling chat g b t work or codex what to do. Open this file. Okay. Great. Can you find this change that I made? It's probably somewhere in the middle, and it can go it can find that using the app shots, right, bringing in all that other context without having me to always tell it, oh, somewhere in this document, you know, I I I think that I outlined, something about that that certain KPI. Right? It's all there, and it's absolutely bonkers. So, you know, why we keep saying Jarvis is this is literally the scene out of Iron Man. Right? When Tony, Tony Stark is at his computer just saying, hey, Jarvis, you know, pull up this.
Jordan Wilson [00:35:09]:
Go do this. Right? Now we actually have this for the first time, and it works. So not just having the remote is cool. The app shots is kind of that, that hack, but what I'm really excited about is for the future of this. So, obviously, OpenAI's, computer use with the new g p d 5.6 model took a huge leap forward. But the thing I'm actually excited about is when we get that super blazing fast, cerebras, integration. So, you know, OpenAI did say that it was coming in July. So when we can bring this with, this ridiculously fast, version of g b d 5.6, What you are able to accomplish in front of your computer without even sitting down is amazing.
Jordan Wilson [00:35:58]:
So I don't know. Maybe maybe this is gonna bring back the popularity of having, like, a treadmill desk or something like that or just being able to go touch grass. I'm like, okay. Is this gonna also, you know, the XR glasses and all these other things? I'm actually geeking out and excited about this one because I don't know. Like, I I I don't like sitting in front of a screen all day. I like getting up and walking around a lot. Right? And, obviously, you can have voice dictation apps, but that's different. So this is just a completely different way to work, a different way to interface with smart AI.
Jordan Wilson [00:36:30]:
It's one I'm super excited about. So this is useful to anyone. So if you're using, the new ChatGPT work desktop app, which is essentially the, developer or the non, technical version of codex, this is in theory a way to completely change how you interface with the computer. Not just how you use chat g b t and how you work, but literally how you interface with the device. So, I'm geeked out about that one. I am too, or I hope you are too. Maybe I'll I'll do a a Wednesday demo of that one. So that's a wrap on all the things that are new this week.
Jordan Wilson [00:37:05]:
Let me give you a very quick update of our seven AI features you can use today, if you have paid plans. So ChatGPT Health is out and launched. Claude has some nice voice mode upgrades. You no longer have to talk to the dumber Haiku model. You can talk to the smart, Opus and Sonnet. We have the new Microsoft MAI image 2.5 pro. Google released Gemini 3.6 Flash and Gemini 3.5 Flashlight. We got the new Gemini Spark.
Jordan Wilson [00:37:38]:
Google's always on twenty four seven AI agent that is available now to pro users. Anthropics, very useful. Claude, record a skill, including that voice, being able to narrate your skill, which is great. And then last but not least, OpenAI went full Jarvis mode with chat GPT voice on desktop. I hope this was helpful. If so, do me a favor. If you're not already, please subscribe to the podcast on Spotify or Apple, and then go to your everydayai.com. We're gonna be recapping the highlights from today's show, anything you missed and everything else in today's newsletter.
Jordan Wilson [00:38:13]:
So thank you for tuning in. We'll see you back Monday and every day for more everyday AI. Thanks, y'all.
