Ep 743: The future of AI? 7 New AI Features that Bring us Closer to On-Demand AI Assistants

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


7 New AI Features Poised to Reshape Business Productivity in 2024

Staying ahead in a rapidly advancing AI landscape is not just about tracking new models—it's about leveraging specific features that make workflows faster, smarter, and more efficient. This week, several under-the-radar AI updates emerged that offer actionable value for organizations and teams looking to refine their operations. Here’s a focused breakdown of the most consequential AI releases and how they translate into business impact.

AI Music Generation: Google Lyria 3 Pro Brings Usable Custom Soundtracks

Historically, AI-generated music has been limited to novelty use cases, offering brief 30-second clips barely suitable for business applications. Google’s release of Lyria 3 Pro changes that, allowing up to three minutes of music creation, accessible to paid Gemini account holders. The system provides granular control over song structure—intro, verse, chorus, and bridge—directly via prompts, making it practical for content creators in search of royalty-free, structured, and tailored soundtracks for presentations, ads, or video assets.

Critically, Lyria 3 Pro excels when paired with Gemini’s ability to extend and iterate on music files, such as generating longer tracks for continuous usage (e.g., in-lobby music or focus playlists). Enterprise video teams and developers building creative tools through the Gemini API or Google Vids ecosystem stand to benefit the most, streamlining content pipelines and reducing dependence on licensed music libraries.

Mobile AI Task Automation: Copilot Task Comes to iOS for On-the-Go Scheduling

Microsoft’s Copilot Task, previously a web-only research preview, is now available on iOS. The tool enables multi-step, agent-driven workflows from any location—users can trigger tasks like sorting critical emails, executing competitive research, or building slide decks, all from their phone. Importantly, these actions are completed via Microsoft’s cloud, not the user’s device, and the app integrates with Google Drive, Gmail, or Outlook, connecting AI workflow orchestration to the business’s core document stack.

This update reduces friction for mobile business owners who frequently handle workflows in transit, allowing for real-time scheduling and routine task management even outside the office. For organizations looking to increase productivity from distributed or hybrid teams, this is a scalable, no-cost feature (for now).

Real-Time Multilingual Communication: Google Translate Live in Any Headphones

Communication barriers in international business meetings, travel, or remote collaboration are being lowered by Google’s Live Translate feature, now accessible on iOS for any headphone type—not just high-end models. Powered by Gemini’s speech-to-speech translation, this feature offers real-time in-ear interpretation in over 70 languages, maintaining tone, cadence, and speaker emphasis. Modes cater to both private (listening) and group (conversational) scenarios, instantly turning standard headphones into highly effective interpretation devices.

This free update represents a practical gain for companies with global suppliers, partners, or employees, enabling seamless multilingual interaction without expensive hardware or external translation services.

Conversational Search and Voice Agents: Gemini 3.1 Flash Live Accelerates Web and Voice Workflows

The introduction of Gemini 3.1 Flash Live marks a new phase in voice-driven web search and audio agent development. Supporting over 90 languages and boasting twice the contextual awareness compared to prior models, this update now includes global availability through the Gemini Live API and Google AI Studio.

For developers and enterprises building customer-facing voice agents, Gemini 3.1 Flash Live provides near-instant audio-to-audio capabilities, better filtering of background noise, and nuanced response handling. Applications include customer service bots, hands-free research, and live video search—such as using a mobile camera to query environmental objects or foreign signage in real time.

Persistent File Storage in ChatGPT: The Library Feature for Organized Insights

OpenAI has introduced the ChatGPT Library for paid subscribers, consolidating all uploaded and AI-generated files persistently in the sidebar, accessible across separate conversations and projects. For business owners and teams, this simplifies document management—no more duplicate uploads or lost reference material. The library enables one-click access to past spreadsheets, presentations, or AI-assembled artifacts, immediately answering project questions or providing historical analytics without context loss.

While currently limited to paid tiers, expect broader adoption and advertising integration in OpenAI’s anticipated "super app," making this feature central to AI-augmented knowledge work.

Genspark Real-Time Voice: Multi-Platform Agentic Control for Business Ops

Genspark’s new real-time voice mode empowers users to manage emails, calendars, and other connected apps by simply speaking complex tasks—even while driving or multitasking. Unlike earlier voice interfaces that struggled with file access or context, Genspark connects voice commands directly to an authenticated ecosystem, enabling actionable agent workflows (not just transcription or note-taking).

Organizations investing in productivity tooling for mobile workforces, or those unsatisfied with legacy voice AI, will find this especially valuable for managing distributed operations where hands-free, context-aware task execution accelerates decision-making and follow-up.

Anthropic’s Computer Use: Direct Agent Control of the Desktop—A New Agentic Layer

Anthropic’s much-hyped "computer use" feature (currently in research preview for Mac) represents a bold step: AI agents can now directly control desktops, mimicking human actions like mouse clicks, keyboard input, and application launching. Unlike virtual browsers, this feature operates at the OS level, interacting with legacy or non-integrated software directly.

The immediate implication for enterprise users is enormous. Tasks requiring complex desktop workflows—especially on apps lacking API connectors—can now be controlled by Claude, with fallback to screen control when necessary. This has the potential to breathe new life into legacy tools and unlock AI assistance for highly specialized, non-cloud software. Security remains a consideration, but for organizations needing automated solutions for old systems or complex cross-app operations, this is an inflection point.

Takeaways for Business Decision Makers

The most significant value of these specific AI features is not in their novelty, but in their tangible impact on day-to-day business operations:

  • Persistent file access in ChatGPT eliminates organizational chaos and accelerates knowledge work.

  • Task and workflow orchestration via Copilot and Genspark makes mobile productivity practical.

  • Lyria 3 Pro and Gemini’s media tools streamline content production, reducing the need for external assets.

  • Real-time translation and voice agents open global communication and AI-augmented customer service.

  • Anthropic’s desktop control signals the opening stage for direct AI agent deployment beyond the browser.

Vigilance and early adoption of these targeted capabilities can unlock measurable productivity gains, competitive differentiation, and flexible employee experiences—especially relevant as the complexity and scale of business digital operations continue to grow.


Topics Covered in This Episode:

  1. Google Lyria 3 Pro AI Music Generator
  2. Suno 5.5 vs. Google AI Music Models
  3. Lyria 3 Pro Extended Song Duration Features
  4. Google AI Music Prompting & Structural Control
  5. Microsoft Copilot Tasks AI Agent Launch (Mobile)
  6. Copilot Tasks Multi-Step Agentic Workflow Features
  7. Google Translate Live in Headphones
  8. Gemini 3.1 Flash Live Audio Model Update
  9. Gemini Live Voice AI & Search Expansion
  10. ChatGPT Persistent File Storage Library Update
  11. Genspark Real-Time Voice AI Agent Capabilities
  12. Anthropic Claude Computer Use Agent Feature Release



Episode Transcript 



Jordan Wilson [00:00:16]:
Each week, there's usually a dozen or so smaller AI updates that slip under the radar yet can completely change how we can interface with AI. This week was no different. We recently started this new Friday features series here on everyday AI to better showcase smaller AI updates in fresh features from all the biggest AI companies. But to be honest, this week, we could have literally just focused on anthropics updates. They had nearly 10 noteworthy new AI features, though their biggest one, their new computer use tool went absolutely viral, and I think it's the first mainstream look at the future agentic layer. And Anthropic wasn't the only company going completely crazy this week with useful AI updates as Google dropped new features across music, real time voice AI, and even translation that will go over all of those today. Oh, and there's new AI powered ways to search the web, better ways to chat with your files on chat g b t, and a new fairly capable live voice agent. All of these new AI features waiting to potentially disrupt your old manual workflow.

Jordan Wilson [00:01:34]:
Did you miss any of these? Yeah. Chances are you did. Don't worry. Sit with me for the next twenty ish minutes, and you're going to be up to speed and the smartest person in AI in your department or company. So here's what we're gonna be going over on today's show and what you'll learn. Well, first, you'll learn why the new Genspark super agent might be worth paying closer attention to with some of its new real time capabilities. You're gonna learn a fast and free way to schedule work from your phone with one of the biggest names in the tech, and you'll learn why Anthropic just showed us the future of AI, even though right now it's a little bit buggy. Alright.

Jordan Wilson [00:02:17]:
You ready to get into it? Let's get featuring on this Friday. Welcome to Everyday AI. My name is Jordan Wilson. If you're new here, this thing's for you. It's your daily livestream podcast, free daily newsletter, helping everyday business leaders like you and me keep up with these bevy of AI updates. I tell you what matters, what doesn't. You take that information and grow your company and career. So if that sounds like what you're trying to do, sweet.

Jordan Wilson [00:02:42]:
Starts here. But for the real cheat code, that's our website, youreverydayai.com. So there, you can not only sign up for our free daily newsletter where we recap each day's podcast and everything else going on in the AI world. But also you can go watch the video version of every single episode as well. So sometimes these, you know, Friday shows a little bit more visual. It should be fine for our podcast audience. You should be fine even if you're just out walking your dog or on the treadmill or whatever. I'll keep you going and make sure to go check out today's newsletter as well.

Jordan Wilson [00:03:14]:
Alright. Let's get into it. Live stream audience. Let me know. Can you see my screen? Hopefully, you can. Alright. First, Lyria three Pro from Google. Google got super musical, this week.

Jordan Wilson [00:03:29]:
And you know what? It might be one of those instances where it's like, okay. Are they competing with Suno here? Does Google even need to compete with Suno? Probably not. But does the new Lyria three pro get there? Not quite. You know? Because also this week, Suno released their newest version, Suno 5.5, which is extremely impressive. But I think for the most part, now even if you were a a casual AI, music user, well, now you have access to the new Lyria three Pro if you have a paid Google account. So even if you were like, you know, this Suno looks great or, you you know, UDO looks great, but I don't wanna pay for it. Well, now you just have access to it with Google's new model. So let's go over what's new.

Jordan Wilson [00:04:16]:
The biggest thing with Lyria three Pro because Lyria three has been out for a couple of weeks, but Lyria three Pro takes the quality and the duration up. The biggest thing is previously, Lyria three only gave you thirty seconds, which is like, what's the use? Right? Yes. It just came out, like, five weeks ago, but, you know, I tested it around. I'm like, what can you do with thirty seconds? Right? There's there's nothing, you know, unless you need an intro for a podcast or something. Right? Thirty seconds isn't doing too much. At least if you want to actually, you know, use the music in a meaningful way. But now you can go up to three minutes. Some other big updates kind of under the hood.

Jordan Wilson [00:04:57]:
Google says that Lyria three pro better understands song structure, and you can prompt for intros, verses, choruses, bridges, etcetera. So, you have a little bit more granular control at least via your prompting, over how the, kind of, music turns out. You know, one thing I didn't see that I wanted to do some testing on, that I did, which it actually worked pretty well. You know, I said, okay. Three minutes is great, but what if you want something longer? Right. One thing I love, I love lo fi. Right. I always listen to, lo fi when I do work.

Jordan Wilson [00:05:31]:
It just helps me kind of space out. And, you you know, all these AI music generators, they make probably better music at their core than these lo fi stations. But, you know, I need, like, an hour of lo fi to really lock in. And I'm like, okay. I wonder if I can upload a Lyria three, music piece of, you know, an m p three into Google Gemini and have it pick up from where it left off. So Google didn't say this, but it's actually really good at that. So if you have, like, a very random use case, it actually did a good job. Right? I I generated something with Lyria three pro.

Jordan Wilson [00:06:04]:
I uploaded it and said, hey. Can you essentially pick off where the end of this song left off? Right? So a lot of these you know, I I know that Google is great with, anything multi like, multi modality, and they have kind of first frame, last frame for video. So I'm like, let me do the equivalent of this for music, and it actually worked pretty well. So here's how you can access Lyria three. So if you do have a paid account on Google Gemini, you already have access, and you can access it within the Gemini app. So if you're on the pro, the $20 a month, well, this makes it easy. It's 20 songs a day. If you're on the Ultra, you get 50 songs a day.

Jordan Wilson [00:06:36]:
You can also access it via Vertex AI. So if you are wanting to use this in production, you have it that, route inside Google's AI Google AI Studio via the Gemini API, Google Vids, and Producer AI. If you are a free Gemini user, yeah, you're not gonna get, access to it right now. So, here's what I think it's useful for and maybe who can find it valuable. I think it's useful. It just moves AI music from kind of these novelty clips to actually usable content. Right? You know, something I used to do a lot way back in the day, like, fifteen to twenty years ago. Right? I've made a lot of videos.

Jordan Wilson [00:07:15]:
You know, and sometimes you would just spend a lot of time looking for the right, you know, music to transition from one clip to another. So if you're a content creator and needing royalty free custom music, this is great. Right? You can, kind of dictate it with the lyrics that you want, all of those things. You know? And just the the new structural awareness means that outputs just sound like real songs and not just random loops. Right? The thirty second clips weren't too helpful or useful in my opinion, but now with three minutes, they are. And like I said, use my little hack that I tried. Let me know if it works well for you, and then you can actually maybe even string something a little bit more together. Also, I think this is good for enterprise video teams that are using Google Vids.

Jordan Wilson [00:07:56]:
Google Vids is actually pretty slept on. I think it's really good. And also just for developers, you know, building creative tools via the Gemini API. Also, it is important to know that all outputs are watermarked with Google's Synth ID. So you can tell that it is AI generated, and Google says it is trained on licensed permissible, permissible data from YouTube and Google partners. Alright. So let's get going to our next one, and this is from Microsoft Copilot. Alright.

Jordan Wilson [00:08:27]:
So, on Wednesdays, FYI, and maybe I'll quickly explain kind of our weekly, rhythm here, and hopefully, you guys are are liking it. I'm kind of enjoying the flow myself. But on Mondays, we bring you the AI news that matters. But generally, with the way that AI has just kind of come to take over the enterprise and the business world, a lot of times, we are not talking a ton about new features unless they're big AI model updates. Right? So on Mondays, we go over the AI news that matters. On Wednesdays, we go hands on and in-depth with one thing. We do live demos, you know, really getting under the hood. And then Fridays here, it's kind of the in between.

Jordan Wilson [00:09:02]:
So on Wednesday, we did go over the new Copilot tasks, which I was very impressed with. But the new feature here for feature Friday is Copilot task rolling out to mobile. So you will have to update, you know, on iOS your, Copilot app, and then, you should be good to go. The good thing here as well is it's free. One thing that's confusing, and I went over this, a little bit more in-depth. So let me see, what what episode, that was our Copilot, tasks episode. If you wanna go listen to that, it's 07:41. Just keep in mind, this is for Copilot on the web, in the Copilot app.

Jordan Wilson [00:09:43]:
This isn't for Microsoft three sixty five Copilot, so not the enterprise version. Although I do think and hope that they'll be rolling this out. And I did, I have been chatting, with the head of this project over at Microsoft. They've already, shipped out some of the features or some of the bugs that I found. So, you know, if there is something in Copilot task that you would want, you know, let me know. I'll see if I can, get the Microsoft team to to cook it up. But a little bit more about Copilot task, you know, and this will explain maybe why it's super helpful on the web. And I think we've seen a lot of this over the past couple of weeks, specifically with everything that Anthropic has been shipping, right, which is just bringing kind of that full, you know, agentic capability, but in a remote way.

Jordan Wilson [00:10:28]:
Right? So, now you can take advantage of the, you know, everything that Copilot has to offer, but via mobile. But specifically, when it comes to, agentic, orchestration. Right? Because that's really what Copilot, tasks is. So this is Microsoft's new agent feature that executes multi step tasks in the background using its own cloud computer and browser. So it is a research preview, and like I said, it is free to use now. And the coolest thing that I think aside from, yes, now Copilot tasks are in a mobile app, which is super helpful. I like being able to text Copilot tasks. Right? Simple things, you you know, hey.

Jordan Wilson [00:11:09]:
What are the most important emails I have right now? You know, it you know, being able to go do some competitive research and create a deck for me. Right? Being able to text something like that for me is really cool, and then being able to schedule these things as well. So right now, this is, available in a research preview. Alright? But it is free. And now with this rolling out to the, the mobile app, the Copilot, app on iOS, I think it's really, really helpful. So like I said, I did go into this a little bit more in the dedicated Microsoft tasks show, but, you know, why it's useful? I mean, come on. This is like a true agentic powerhouse. I was actually kinda shocked at how good this is just because for me, you know, Copilot has never been something that I'm, like, the first to rush to, right, unless I really need to do something inside Excel, inside PowerPoint for a certain reason.

Jordan Wilson [00:12:04]:
Right? Because then they do have, some of the new integrations with both Claude and with Chad GPT, that I think are really, really good. But for the most part, I'm not, you know, rushing to use Copilot as, number one or number two in my stack, if I'm being honest. But with this one, I am, just because it did a really good job, and it was fast. Right? And go check out that, entire episode. Right? Seven forty one. So one thing I really liked about Copilot tasks is you can even edit the slides that it creates. Right? In the same way that, you know, Google has kind of an annotate feature, in some of their different products, you can literally, it can go and do a bunch of research for you on its own. Agentically, you can schedule something every day or, you know, once a week, right, and connect it to your data as well.

Jordan Wilson [00:12:49]:
That's the other big thing. Right? So connecting it to, you know, your Google Drive or your Gmail or your, you know, your Outlook email. Right? If there's something that you do routinely every week that lives inside of those apps. Right? You can schedule it. It can grab that. It can create you new documents. Right? You can browse the web. It has access to its own browser, and now it's also, accessible via the app.

Jordan Wilson [00:13:10]:
So, pretty pretty big one there that I think a lot of people are gonna find useful. Alright. Next. This is a smaller one, but I think this is one where technically, like, billions of people, are gonna find, I think, just immense value out of this. This is because Google Translate live is finally available in headphones. Alright. Stick with me. I know this sounds small.

Jordan Wilson [00:13:38]:
So essentially, this has been a feature that had previously been rolled out to Android, but is now available, on the Google or or sorry. Inside the Google Translate app on iOS. And here's why I think this is really important. Well, first, just kinda let me tell you what it does. So the new live translate with headphones feature, is available on iOS. And, essentially, any, headphones that you have, you know, whether they're wired or Bluetooth, right, now they are can kinda be a real time personal interpreter. Right? So some of the newer, you know, AirPods, as an example, have this feature kind of built in, but they're pretty expensive. Right? So now it's literally any pair of headphones.

Jordan Wilson [00:14:19]:
If you're using Google Translate can translate live. Right? So this is, I think, a, you know, one of those things. Yes. There's plenty of business use cases, obviously. Right? If you do any international business, if you're traveling internationally, I think this is really gonna help, especially if you don't speak the language. But, you know, obviously, on the personal side, this is huge. Right? And so this is powered, by Google Gemini's speech to speech translation, and it also preserves tone, cadence, and emphasis of the speaker. And the best thing is, well, it's free.

Jordan Wilson [00:14:53]:
Right? And it does work with any headphones. You don't have to have, you know, any AI powered headphones, any of the new, Google or Apple headphones, literally anything as long as you, have access to data on your phone, and you have the, the latest version of Google Translate, that's all that is needed. So, there's three different modes. So there's a listening mode that's real time translation in your ears. There's conversation, so you can hear the translation in the headphones and others hear it out loud. And then there's text only that is just the on screen text translation. So right now, it supports more than 70 languages, and while there's no more you know, I've I've had to do this before in certain instances. There's no more kind of passing a phone, you know, back and forth.

Jordan Wilson [00:15:39]:
You can just wear the headphones now and listen. So, obviously, if you're visiting different countries, this is super helpful. You know, families with members that speak different languages or just business professionals in multilingual meetings. This one, I think, is going to be very helpful. Alright. Our next new feature, this one also from Google. You might be saying, wait, Gemini 3.1 Flash, isn't this an old model? Well, yes, but no. Because it's brand new.

Jordan Wilson [00:16:07]:
Because this is now Gemini 3.1 Flash Live. So here is what's new, and this does change the way that you search. So just stick around for that for a second here. So this was just released yesterday. This is Google's new audio to audio model, and it powers the new Gemini live experience and search live. Alright. And search live, so this is just a way to essentially well, it's the way you search the web with live video. Right.

Jordan Wilson [00:16:38]:
So this is also expanding globally to 200 plus countries, and this was previously in US only. So, here's who has access. Well, literally, anyone. Anyone can if if you're using AI mode, in Google search, you will be able to benefit from this. Obviously, developers, you know, you can use this on the Gemini live API or in Google AI studio or enterprises, you know, in the Gemini enterprise for customer experience. So here's why it's useful and I think who will find it valuable. Well, I think it just follows conversation way better according to Google. Right? And it also holds context for two times as long as the previous model.

Jordan Wilson [00:17:24]:
So it's better at filtering filtering out background noise, right, traffic, TV, also being able to, recognize kind of tonal nuances so we can adjust responses when users sound frustrated or confused. Right now, 90 plus languages. So this is obviously developers if you're building voice agents, customer service agents, this is huge. Right? So the flash version, obviously, a little, cheaper than the normal Gemini three one pro, you know, which is multimodal by default. So this is specifically made, for audio to audio. And I think that where, this is really gonna get popular aside from being able to use it on the web is just with developers being able to, create voice agents. I think this is one area where, Google's, Google Gemini models are really good. Also, I think enterprise contact centers are gonna find a lot of utility out of this and just international users, who couldn't use search live.

Jordan Wilson [00:18:20]:
Right? This is only available in The US, but now anyone can use it. I think there's a little, kind of video demo. It might have been on a different page. Right? But if you haven't used, search live before, it is literally right. If you see something out in the world and you're like, what the heck is this? Well, it's a live version of search. So you can click the, the live button, and it will, see your camera, and you can talk to it as well. So if you see something strange or if there's a sign, again, kind of, you know, going back to the multilingual, aspect. Right? If you're traveling and you're like, okay.

Jordan Wilson [00:18:57]:
Like, am I in the right spot? You know, it will be able to see your camera. You can talk to it. Right? But this is where it's really helpful. It's just this faster, inference, the faster, model, and, you know, being able to, you know, work across different languages is obviously helpful as well. So right now, also, anything that's generated does have the synth ID, just FYI. And developers, if you are migrating from 2.5 Flash, this does use the thinking or make sure to use the thinking level instead of thinking budget. Some things for developers, it did change a little bit under the hood. So pretty impressive model.

Jordan Wilson [00:19:37]:
So here's kinda what Google said about it on their announcement post. They said today, we're advancing Gemini's real time dialogue capabilities with Gemini 3.1 Flash live, our highest quality audio and voice model yet. It delivers the speed and natural rhythm needed for the next generation of voice first AI, offering a more intuitive experience for developers, enterprises, and everyday users. And you can access it in the Gemini Live API, Google AI Studio, or the, like I said, the Gemini enterprise for customer experience, and it's available in search live in Gemini live. So, another thing. Right? The Gemini live, aspect is helpful as well. Right? So if you are someone that inside, you know, Gemini as an example or you're someone that loves to use, you know, chat g b t voice mode, you know, make sure to check out this new version inside Gemini with Gemini live. Alright.

Jordan Wilson [00:20:29]:
Next, this one's gonna seem small, but chat g p t has 900,000,000 weekly active users. This is a small thing that a lot of people are going to find helpful, but you probably didn't see anything about it because it's not even really technically a feature. It's just a new way that OpenAI is essentially storing files inside chat g p t, but it's important. So stick around. So the chat g p t library has officially rolled out inside of chat g p t. This is a new persistent file storage hub in chat g p t's sidebar. So it automatically saves every file you upload or that chat g p t creates. So that's documents, spreadsheets, presentations, etcetera.

Jordan Wilson [00:21:17]:
And here's the cool thing. Files persist across conversations until you manually delete them. So who has access to this right now? Well, sorry. It's not actually 900,000,000, but, 900,000,000 people have the ability to go use this if you want to because it is only available right now to paid subscribers. So you do have to be on the ChatGPT plus, pro, or business plan. It's not available yet on the Chat GPT free or go plans. Although, I do feel that eventually, OpenAI will roll this out because on the free plan, right, I think that's where they're really gonna benefit, from being able to serve better ads. And if you can, well, access people's uploaded files, you can serve them way better ads.

Jordan Wilson [00:22:02]:
So my guess is eventually, this will roll out to free, consumers when we get the, the new super app, from chat g p t. Right? If you listen to, the show yesterday, episode seven forty two going over, OpenAI's most chaotic week ever, We did talk about that OpenAI is eventually gonna be rolling out a kind of single app to rule them all. Right? They got rid of Sora, but they're essentially gonna be rolling their chat GPT, codex, and Atlas all in one app. So I do see this new library feature, being very useful in the future super app. But, I mean, when it comes to just utility on how you can use it today, I mean, here's here's the biggest thing. You can ask chat g p t about any files that you've saved. Right? And that's the biggest thing, and that's something I'm gonna immediately find helpful. There's a lot of different ways that you can organize your files.

Jordan Wilson [00:22:59]:
Right? You can use projects. You can use GPTs. You can upload a file to an individual chat, to an individual conversation. Right? But sometimes, I'm, like, even thinking in my head, especially when I'm on the go. I'm like, oh, you know, I I know I had JetGPT go. You know, I'm usually using, GPT five four pro for this. I know I had to do a deep dive on, you know, this big spreadsheet, but where did I save that? Was that in a project? Was it in a GPT? Was it in an individual chat? Now you just if you've uploaded the file, if you know what it's called, or if you have an idea, it doesn't matter. You can just ask about it.

Jordan Wilson [00:23:34]:
So this is great, especially for people like me who try their hardest to stay very organized in chat GPT, but sometimes you don't for whatever reason. This is this is pretty big, I think. And I think this eliminates one of the biggest friction points of Chat GPT, which is uploading the same files across different conversations or just losing place. Right? So for these artifacts that ChatGPT does, produce, that's big as well. So, yes, this is helpful for the files that you upload, but also what if right, people? Right? You need to be using ChatGPT for this, y'all. You need to be using Chat GPT to make you spreadsheets, to make you PowerPoint presentations. Right? So now when Chat GPT makes you documents, it all goes in the library as well, and you can reference those from anywhere. So again, this might be one of the smallest, updates on paper, but if you are a power chat g p t user, this is actually a pretty big update.

Jordan Wilson [00:24:31]:
Alright. Moving on next. So Genspark, one of the super agents. Right? Obviously, a growing category, you know, since Open Claw has really popularized the, you know, super agent or the, you know, all in one agent. But Genspark and Manus as an example had been around for a while. And I think that GenSpark just got a new update that a lot of people are gonna like. Right? So when we talk about real time voice and access to your data, that's huge. Right? That's one thing for me personally.

Jordan Wilson [00:25:06]:
I don't know. Chat GPT's voice mode hasn't really done it. I do have to give the new version of the Gemini three one, a little bit more testing since it literally just came out. But I might actually I don't even know if I have a paid GenSpark account. It might be one of the only, platforms I don't have a a paid account for. I pay for literally all of them. But this one, you you know, they just released a demo video, this week. Essentially, you know, this woman who, you know, was using GenSpark real time voice on her drive.

Jordan Wilson [00:25:36]:
You know, she's having it check her calendar, move things around, send emails, you know, check her, you know, her teams. Right? So anything that GenSpark connects to, at least when this works and if it does work correctly. Again, I think I've tried all these other features, it, on today's show except this one, FYI. So I don't have any personal experience with this one, but in the demos that I've looked at, it looks pretty good and pretty helpful. But, you know, the true, I guess, feature here, I think this could shine in areas. Right? I think chat g p t voice mode is probably the ones that, you know, most of us might use or maybe, you know, many of us are just so kind of disappointed with the state of voice AI, although I do think it's getting better and better. But one thing that I've always realized is sometimes at least ChatGPT's voice mode, you know, used to be called advanced voice mode. It's still running a super old model.

Jordan Wilson [00:26:30]:
It's running a version, I believe, still of the GPT four o series, so not the smartest model. In in my testing, I've done a lot of testing in the past. It usually had problem accessing your files. So that's with Genspark real time voice. I do think that's helpful in being able to access the different connectors if you did give it, access to. So this isn't just transcription. Right? This is voice commands that you can trigger, you know, multi step agent workflows across your different apps. You know? So you can just speak tasks and get them executed, manage your calendar, check your email.

Jordan Wilson [00:27:07]:
Right? This is one thing. When I saw this demo, I'm like, oh, wait. I might actually really benefit from this. So it is also kind of available via their Speakly app. So Genspark has kind of a dictation or, you know, real time agent voice app called Speakly, but it does connect via Genspark as well. So you do get a some use case on the freemium model, but not a whole lot. But if you are on the, the plus, account or the pro account, I think you're gonna get a lot more access out of this. So here's why it's useful.

Jordan Wilson [00:27:40]:
Well, you can just speak a complex task, and Genspark's multimodal agent can break it down and execute it. Right? And it claims four times more efficiency over keyboard inputs for complex tasks. And like I said, it does connect to the same agent ecosystem that handles, you know, your research, email, contents, and even phone calls. So I do think who's gonna find this useful. If you're someone that's on the go, right, give it out. Give it a try on the on the, on the freemium plan. But if you're someone that's on the go, it's gonna be helpful. If you've tried, you know, other, kind of AI apps or AI agents that connects to your, you know, all these other tools that you use, your Slack, your calendar, your email, and it struggles for whatever reason, I think this is at least I'm not saying this is gonna, you know, be the tool that solves your problem, but it's at least, a contender that could potentially help.

Jordan Wilson [00:28:31]:
Alright. And last but definitely not least, this is the one that went mega viral. Yes. Anthropic released so many new features this week. Alright. Some that didn't even make our list. Right? They have the new work tools that are available on mobile. They have the new Claude Code channels.

Jordan Wilson [00:28:50]:
Right? A very open claw, feature being able to talk to Claude Code via Telegram, Discord, iMessage. They have their new co work projects. Like I said, Anthropic literally released, like, 10 pretty big updates this week. But the biggest one by far and I'll say this, this might be the most, like, viral AI update ever that wasn't a new big model. Right? That's because at least by Twitter vanity metrics, this thing got like over 75,000,000 views. Everyone had eyeballs on this. So this isn't one you miss. Alright.

Jordan Wilson [00:29:24]:
But this is their new computer use. Right? Unfortunately, they didn't really give this a good name because at first, they were kind of tying it to dispatch, which is the new way to kind of, control the desktop version of Claude CoWork with your phone. Right? Dispatch, we covered this on the show last week. Really cool. Right? It's a way that you can now use your phone, right, the iOS app, inside Claude, and you can control your desktop computer, via dispatch. Right? So before the computer use, you couldn't actually control the computer. But with computer use, you can. I'm talking about mouse clicks, you know, keyboard inputs, not just, you know, a a a web browsing agent.

Jordan Wilson [00:30:14]:
So this is completely different. Right? What we've seen, a lot with the the the the current kind of, agentic, tool use stack. It's virtual browsers. Right? Like, a lot of these, you know, ChatGPT agent mode as an example. It has a virtual machine, and it works in your browser. Computer use literally is computer use. Claude can see and use your computer. Right? I did a lot of testing on this and, you know, let me know.

Jordan Wilson [00:30:45]:
Let me know if we should do this for Wednesday's show. I'm thinking maybe. Alright? But if you're listening on the podcast or on the live stream, just say computer use. Alright? I always have an arbitrary random number, in my head. And if we hit that, I'm like, yeah. We'll do it. So I did do some testing on it. It's super helpful.

Jordan Wilson [00:31:02]:
I mean, here's one of the reasons. So we did technically already have a lot of this functionality, whether it was via co work, you know, being able to run certain things in the terminal. But now, like, some of the testing I did, well, I was just having it launch and use other apps. Right? I was kinda having fun with that. So with the new computer use, I had, you know, Claude launch Atlas and, you know, that's, OpenAI's, agentic browser. And I was having it do some, I think I was having it use perplexity inside of Atlas. There's stuff like that. Right? So I think there's some huge benefits.

Jordan Wilson [00:31:41]:
But first, let me just tell you a little bit more, kind of about how it works. So it can now control your entire Mac. So that means open apps, navigate browsers, click buttons, fill spreadsheets, anything that requires a mouse click, right, which is something that, even their, kind of Chrome extension would sometimes struggle with, right, because it was using screenshots and computer vision. So now this is a completely different ballgame. This is a new kind of layer. So it does use connectors first, FYI. So if you are asking something about Slack as an example, it's first gonna check a Slack connector if you have it. But then it'll fall back to screen control if a connector doesn't exist or if it can't find, something in that connector.

Jordan Wilson [00:32:27]:
So right now, it works in both Claude CoWork, but also Claude CoD as well. So this is confusing because it says, on their kind of blog post here, it says, let Claude use your computer in CoWork. And then originally, they were just tying it to dispatch, but you don't need to use dispatch to use it. So, I'm actually kind of quite shocked. Maybe Anthropic didn't understand how quite viral this would go to give it a proper name. So I think at first, people were calling it, computer, computer control via dispatch, and then they were saying it's co work, computer control. And I think now most people are just calling it computer use. So, yeah, this is really, really helpful.

Jordan Wilson [00:33:11]:
So you do have to be on a paid plan right now, and it is only available for Mac, and it is a research preview. It is buggy in my limited experience so far. I've been using it for a couple of days since it was announced. But when it works, right, when it works going back to how I started the show, I think this is the, not the end interface of AI agents, but I think this is the next big interface of AI agents, because I think eventually, you know, especially like, voice commanding, agents, they're not gonna need our computer per se. But I think we may, you know, need a couple of years until agents just have a kind of stronger and more reliable protocols, you know, versus just a two a, MCPs. Right? Eventually, I think agents will just be able to, be able to nicely translate everything, and they won't need to necessarily use our desktop as an operating layer. But for now, I think a 100%. And the other thing, if this is used in enterprise, and that's a big if.

Jordan Wilson [00:34:17]:
Right? Because obviously, there's a lot of new dangers with this on the prompt injection side. Right? But this could be something very useful, right, for, companies that are using old legacy software, right, that just doesn't work well with any kind of AI. Well, now you can just keep this open, give it a pretty difficult task, and it can technically operate that legacy, that legacy software. So, who else is gonna find value? Why else is it useful? Right. So like I said, this can still pair with dispatch. So that's the cool thing. You can literally control your entire desktop computer from the Claude app on iOS, which is really cool. And I think this kinda fills the long tail.

Jordan Wilson [00:35:02]:
Right? It covers the apps that will never have a dedicated connector, and you can you know, if you're running late to something, right, I was doing this a lot. I was just out on I don't know walk. Right? And I'm like, oh, you you know, for me, it's not like I needed something, but I was just kind of testing it in that scenario. Oh, like, I know I download this file on my computer. I'm not sure where it is. Right? Like I said, a lot of these things were, available before. If you had things set up correctly via co work, and it could access certain elements of your computer, but now it can literally control anything on your computer as well. So, I mean, just who's gonna find this valuable if all of us.

Jordan Wilson [00:35:38]:
Right? Literally anyone. If if you're on Mac right now, if you have a iPhone, this is huge. So knowledge workers who are just juggling a lot of apps, especially those apps that don't have dedicated integrations, that's the big thing. Yeah. You can use this to showcase certain things, and it's fun and it's cool. But if there's a dedicated connector integration, you're just wasting your time by doing it this way because it is still just using computer vision and clicking around. But for all of those things that don't have dedicated apps, within Claude, this is huge. Right? So anyone that just wants to delegate desk work from their phone while they're on the go or just power users who are already using coworker dispatch, like I said, this, I think, signifies the next step in agentic work.

Jordan Wilson [00:36:22]:
So that's a wrap for today's show. I hope this was helpful. Are you liking the new Friday features? It's a well, I say it's a quicker show, but I geeked out and accidentally talked for thirty six minutes. But, hopefully, this is a good way. You know, it's so hard to keep up with everything that's going on. Yeah. We do the AI news show on Monday, but like I said, a lot of that starts to seep over into just big enterprise, politics, society, culture, all these other things. So I think Friday is our shows.

Jordan Wilson [00:36:51]:
And let me know if you find it helpful that, you know, are like, hey. You probably missed this. Here's why it's helpful. Here's how to go find it useful and to get you going on your way. So if this was helpful, if you are on the podcast, please do me a favor. Leave us a rating if you could. Right? If this is helpful, go tell other people about it. Right? Subscribe to this podcast if you're listening on Apple Podcasts, Spotify.

Jordan Wilson [00:37:13]:
We'd really appreciate that then. Go to youreverydayai.com. Sign up for the free daily newsletter. Thanks for tuning in. We hope to see you back next week and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI