Ep 815: New ChatGPT Voice model, Grok 4.5 drops, Meta’s ai comeback and 7 more New AI features to use Today

Resources:

Join the discussion on LinkedIn: Got something to say? Let us know on LinkedIn and network with other AI leaders


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Start Here Series in our Inner Circle Community: Join for free access


AI Voice Models and Next-Gen Tools: Business Value from the Latest AI Updates

Recent advancements in artificial intelligence were scrutinized in a detailed breakdown focused on tools that are immediately shaping how companies operate. Instead of abstract hype, the analysis outlined seven specific, newly released features now accessible to teams of any size. Critical attention was given to why and how these models differ from earlier versions, and what business functions they specifically enhance.

Business Impact of Real-Time AI Voice Assistants

The conversation focused on the release of OpenAI’s GPT Live voice assistant model, contrasting its capabilities against prior iterations from major AI vendors. Previous voice modes were limited by significant latency, robotic interruption patterns, and an inability to reference real-time web data during dialogue. The latest duplex architecture solves these shortcomings by enabling the model to listen and speak simultaneously, actively using phrases like “mhmm” or “got it” for a more lifelike interaction 06:03.

Business value emerges from the model's capacity for real-time search and reasoning using GPT-5.5, which supports genuine hands-free workflows for activities like brainstorming sessions, impromptu tutoring, and live Q&A. Broad access—extending to both free and paid users and enabling a billion users to activate these tools—democratizes this capability. Practical scenarios include voice-based business queries while commuting, managing logistics during travel, or fielding complex stakeholder questions without a screen 09:35.

Developer Tools: Enhanced Voice API Flexibility

One concept discussed was the simultaneous release of developer-targeted voice APIs (GPT Real Time 2.1 and 2.1 Mini) 10:48. These models offer adjustable reasoning depth, native tool calling, and improved support for tasks involving alphanumeric data, silences, or interruptions. The business value lies in the lower latency—reduced by 25%—and real-time response flow, making the technology suitable for cost-sensitive custom voice agents used in automated call centers or live business support systems.

Product teams can now build persistent, always-on agents that execute diverse commands with real-time feedback, improving operational responsiveness and enabling the automation of routine employee or customer interactions 12:37.

AI Model Releases Optimized for Software Engineering

A key theme that emerged was the introduction of GROC 4.5 by XAI, specifically tuned for software engineering and complex agentic workflows. Unlike all-purpose models, GROC 4.5 is engineered to use fewer tokens for coding tasks, claiming up to fourfold efficiency gains compared to peer solutions 13:33. Benchmark rankings indicate strong positioning in coding, mathematical, and scientific knowledge tasks.

For organizations with costly software development pipelines or with a mandate to maximize compute efficiency, this model presents a targeted option. Integration with Cursor—an acquired tool specializing in code workflows—further positions the stack for enterprise-scale, compute-intensive engineering projects, leveraging deep reasoning at a fraction of the historical cloud cost 16:08.

Dedicated Mobile AI Agents for Document Management

The discussion explored Notion’s launch of a standalone iPhone app dedicated to running its AI agents 18:23. Unlike generic mobile offerings, this tool allows teams to capture tasks, ideas, and updates via text, voice, or photos and trigger AI agents to generate pages, draft communications, and process workflows directly from mobile devices.

This approach eliminates friction for mobile knowledge workers and aligns especially well with distributed teams managing high-velocity documentation and project coordination. Importantly, the model-agnostic design supports integration with GPT, Claude, Gemini, and other enterprise models to fit existing corporate AI strategies 18:51.

Communication and Call Management with AI-Driven Notes

Several points were raised, including Google Voice’s AI note-taking capabilities, enabled by Gemini models on paid business plans 21:35. This feature supports real-time call transcription, summarization, and action item extraction, distributing post-call notes automatically via email.

For small businesses and field teams who spend significant time on the phone, this allows instant conversion of voice conversations into searchable, actionable digital records. It directly reduces manual documentation costs and ensures operational continuity by embedding call intelligence into everyday tools 22:13.

Multilingual and Data-Driven Visual Content Creation

Attention was given to ByteDance’s release of cDream 5.0 Pro, an image model positioned as a contender for data visualization and infographics. The discussion explored advanced capabilities such as dense text-to-visual reasoning, interactive layer editing, and native support for 10 languages 25:24.

The key value for content teams working in design, marketing, or multilingual environments lies in the ability to convert structured data or complex ideas into branded, layout-ready visuals. Precise output for international campaigns and automated infographic production can translate into measurable gains in campaign velocity and asset consistency.

Meta’s AI Image and Video Models: Social-Centric Automation

A notable business trend identified was Meta's renewed push with Muse Image and Muse Video models 29:26. Distinct features include agentic, self-refining workflows and the ability to generate or manipulate images of public Instagram accounts via simple tagging.

For entities focused on social media marketing or content management, the advantage is the seamless ability to create contextually relevant visual assets inside familiar platforms such as Instagram, WhatsApp, or Messenger—minimizing the need for external tools. Free daily usage within these apps opens advanced generative features to teams prioritizing rapid content iteration 32:15.

Conclusion: Immediate Enterprise Applications

The discussion systematically broke down seven actionable updates spanning voice, developer APIs, software engineering models, mobile agent tools, call intelligence, visual content tools, and social media automation. The throughline was practical deployment—each new feature offers a specific, measurable improvement for modern organizations adapting to the accelerating adoption of AI in real business contexts.


Topics Covered in This Episode:

  1. OpenAI GPT Live Voice Model Launch
  2. OpenAI GPT Live Duplex Conversational Features
  3. GPT Live Access Free and Paid Tiers
  4. OpenAI Developer API: GPT Real Time 2.1 Release
  5. XAI Grok 4.5 Model for Software Engineering
  6. Grok 4.5 Token Efficiency and Benchmarks
  7. Notion Agents Standalone iPhone App Release
  8. Google Voice AI Gemini-Powered Call Summaries
  9. ByteDance SeeDream 5.0 Pro Image Model Launch
  10. ByteDance Image Model Infographic and Layer Features
  11. Meta Muse Image and Video Model Rollout
  12. Meta AI Agent Image Generation Through Instagram




Episode Transcript 




Jordan Wilson [00:00:18]:
Most AI tools at work don't know anything about you. Every time, you're stuck reexplaining your projects, your team, what actually matters. Slack just fixed that. The all new Slack bot is a personal AI agent built into Slack, and it starts with your context, your messages, files, and channels, translation, and AI that already knows how you work. Check it out at slack.com. If you're a heavy AI user like me, you always have your eye on the next big model release from one of the big players, And we've actually seen plenty of action on that front the last few days since the big model makers have started to work with the federal government and get the models actually released. As an example, we have Fable five now for a few more days, and we're going to have GPT five six SOL, in a couple of hours. And maybe we'll have Gemini 3.5 pro any day now.

Jordan Wilson [00:01:20]:
But when I think about it, in those instances, we're actually just going from having a really great model to a model that's just a little bit better than the great model we had before, which when it comes to tracking AI releases, sure, that's helpful. But are any of these a step change or, you know, really unlocking something that we didn't have before? And even with these great models, the honest answer is probably not. But this week, we actually got one of those AI features that I am beyond excited for because we're going from having honestly not a good product at all when it came to real time, voice AI to now something that is finally good. So we're gonna be going over that, the new OpenAI chat g p t real time voice model that's available now as well as a bunch of other AI updates that you might have missed this week. So on today's show, in our weekly Friday features on Thursday roundup, you're gonna learn why a new voice model may be a bigger deal than Fable five or GPT five six Soul. You're gonna learn how Meta may be quietly making its AI come back. A lot of big releases from Meta this week, and you're going to see the new image model that's pushing the once untouchable chat g b t images too. Alright.

Jordan Wilson [00:02:44]:
Let's get into it. Welcome to Everyday AI. My name is Jordan Wilson, and we do this thing, well, every day. This is your daily, unedited, unscripted, live stream podcast, and free daily newsletter helping business leaders like you and me keep up with the nonstop avalanche of AI updates. I tell you what matters, what doesn't. You take that information. All of a sudden, you're the smart AI person in your company. You're welcome.

Jordan Wilson [00:03:06]:
Hey. To show appreciation, do me a favor. Go ahead. Subscribe to the show here if you're listening on Spotify or Apple Podcasts, and make sure you go to our website at youreverydayai.com. We're gonna be recapping the highlights from today's show. So if you hear one of these seven AI features and you're like, wait. What was that? Yeah. It's gonna be in the newsletter, so don't worry.

Jordan Wilson [00:03:26]:
Alright. And, hey, FYI, we're doing this a little different this week. Normally, we do the Friday AI feature show on Friday. But in our newsletter, I put out a poll, yesterday, and, well, we know that chat g b t's new model, g b d five six soul, is going to be released, well, today. So on our Friday show, we're gonna be going over that new model. So today, we're doing our Thursday feature. So you guys voted for it overwhelmingly. For the most part, I just work for you guys.

Jordan Wilson [00:03:56]:
So, yeah, if you are a podcast listener and not on the newsletter, you gotta get on the newsletter because you're just directing me. I'm like your human agent that just, you know, you know, goes to work for you. Alright. So let's get into it and go over the new features. Alright. So, the big one. And I think I mean it. This one, technically, in terms of unlocking something, some new AI capability for our work that we didn't have before, yes.

Jordan Wilson [00:04:26]:
I mean, you can make the argument Fable was such a big jump or, you know, the original, like, o one and o three models from OpenAI were a big jump, and, yes, they were. And, yes, people are excited to have Fable five back for a couple more days under subscription plans, and, you know, we're gonna be getting GBT five six soul any minute here. But I think the new GBT live voice model may actually be one of the biggest step forward, in day to day AI use that we've had in a very long time. That's, I think, for a couple of reasons. One, the models before were just really bad. Right? Chad GPT's original voice mode, when he came out, it was like, oh, kind of cool. And then you had your advanced voice mode for the last, like, two years, and this is really not good. Right? Gemini live, not really good.

Jordan Wilson [00:05:17]:
You you know, Claude has a voice mode, not really good. At least, especially from the big players, there just hasn't been a great, voice, you know, AI voice assistant that you can talk to that can pull real time information from the web that can work with the conversation and the context that you give it. We just haven't had that, which is crazy. Right? Because people are getting more and more comfortable dictating versus typing. So now this one exciting GPT live. It is out now. It's new. Let's talk.

Jordan Wilson [00:05:45]:
So here's what's new. OpenAI killed the the kind of turn based, you know, kind of walkie talkie voice mode and replaced it with the new GPT live, which is a duplex model, what they're calling it. It listens and it talks at the same time. So it also does this, like, back channeling like a human. So as you're talking, you know, it'll say things like, mhmm. Got it. Right? And then it redecides several times per second whether to speak, keep listening, pause, interrupt, or call a tool, like to, you know, look in the context of your, conversation, to look on the web, etcetera. Also, it pulls in g p t 5.5 at the same time for web search and heavier reasoning.

Jordan Wilson [00:06:26]:
So before, obviously, the voice models, you know, what they worked on, I think, over the last, like, you know, 2023 to 2025 was just making the voices sound more realistic and a little less latency, which is, like, that's cool and all. But if they don't know left from right, if they don't know what day it is, if they can't look up real time information, what good is it? So that's what's different with this new model from OpenAI. So here is from OpenAI's blog post what they said. They said we're launching GPT Live, a new generation of voice models that make talking with AI feel much more like having a real conversation. GPT Live is built on a full duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT can show it's paying attention with phrases like, yeah, engage in a quick back and forth, or just stay quiet when you need a moment to think. The result is a voice experience that is refreshingly easy to talk to. GPT live is also our smartest voice model yet.

Jordan Wilson [00:07:26]:
For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the results back into the conversation when it's ready. While it works, GPT Live can keep talking with you and maintain the flow of the conversation. At launch, GPT Live will, use GPT five five in the background as we release new frontier models. We'll continuously update the models used by GPT Live. So like I said, huge news. My gosh. I've been waiting for this because I've tried. I've tried to find actual business.

Jordan Wilson [00:08:02]:
Right? Whether I'm taking a walk or, you know, out in the car or if I just, like, wanna take a break from looking at a screen. I I've tried every single, you know, voice model from all the big players. I think when Google's Gemini live came out, it was okay. But Gemini live actually does a bad job, strangely enough, at querying information from Google. Right. So I am excited. I've gotten to play with this one a little bit, but it is brand spanking new. So who has access? Well, everyone.

Jordan Wilson [00:08:31]:
So if you are on the free plan, even on chat g p t, you do have access to this. It does use a, technically, the mini version of GPT Live One, but even free users. So a billion people yeah. A billion. That is how many weekly active users OpenAI has inside of chat g p t. So a billion people just got a brand new huge upgrade in what you can do with AI. Right? And then on the paid tiers, you do get the full GPT Live One model. So it is live now.

Jordan Wilson [00:09:02]:
Make sure your app is updated, and then you just go and tap tap the voice button, and you're off to the races. So why is this useful? Well, it removes the dead air that made the old voice mode feel robotic. Right? If you would always kind of, you know, interrupt each other, you're waiting. It it cuts in when you don't want it to. Right? It was just all voice models were really, really bad. Now not the case anymore, and this is genuinely usable hands free for brainstorming, tutoring, and live q and a where interruption actually matters. So who's gonna find it valuable? Well, anyone using chat g p t while you're driving, walking, cooking, you know, packing for a trip. I think there's a lot of, you know, cool use cases, that with it being able to reason and look things up in real time that we just didn't have before.

Jordan Wilson [00:09:51]:
Right? And we've seen rumors over the last year or so that, you know, OpenAI was working on some hardware devices. Those things, you know, allegedly got tabled. The Johnny Ives, you know, IO acquisition. So maybe this is the model that will power future, OpenAI hardware when and if that comes. If so, that makes a lot of sense, because I think right now, Google Home, with Gemini is the only one that has a true smart AI home speaker. And even the difference now, between GPT Live and Gemini Home I mean, GPT Live is in another another field, at least for now. We'll see what the rest of the labs do. So, pretty big news there.

Jordan Wilson [00:10:30]:
Let's go on to our second update, which is technically related. So this is kind of the developer version of this model. So from OpenAI, they also released, the day before. So, they released this on Tuesday, two new API voice models. These are called GPT real time 2.1, and then the budget friendly GPT real time 2.1 mini. So this allows you to add adjustable reasoning, depth, native tool calling, and better handling of alphanumerics. Right? Things like, you know, ordering numbers, codes, etcetera, silence, noise, and interruptions. So OpenAI did say they also cut latency in this voice model by 25%.

Jordan Wilson [00:11:15]:
And this is for developers. Right? So if you, are using the back end of the OpenAI platform, this was launched, like I said, this week, and it is generally available. You can use it in the, OpenAI Playground, but this isn't like a consumer feature. So this isn't what you're gonna, you know, log in to chat gbt. In that aspect, you're using the new g p t live model. For, builders, you are using the g p t real time now 2.1. So, who's gonna you you know, why is this useful? Who's gonna find it valuable? I think developers, product teams who are building voice agents, right, call center automation, anyone working on a cost sensitive voice start up. So even if you're not a true developer, right, this is something I'm gonna be using, because I've wanted to build something.

Jordan Wilson [00:12:07]:
It probably doesn't take long. Maybe I'll do it as a live show once. It's just hard to sometimes use my phone and audio when I'm trying to do things that require audio. Anyways, you know, I've wanted to build a simple, computer control agent that I can talk to and say, hey. Open this. Open that. Right? So kinda like what I do in codex, I dictate to codex, but I thought, hey. It'd be cool if I just had an assistant that lives and it's always on and always listening, and I can just talk to it and it can do anything that, you know, the, OpenAI models can do.

Jordan Wilson [00:12:37]:
So that's an example of something that you could build, and you can do it fairly easily just, you know, using codex. So developers, product teams are gonna find it useful, but, I mean, why it's actually useful is, well, they just the lower latency and then the built in tools mean that any voice agents that you or your team build will actually feel real time instead of laggy. Alright. Let's go on our third one. We have a new model. Is it in the frontier? Maybe. We'll see. We don't have a lot of benchmarks yet, but let's talk about what we do know about GROC 4.5, the new model hot off the presses.

Jordan Wilson [00:13:15]:
So this is XAI, or are they called SpaceXAI now? I I don't even know. I think they're just XAI. Right? So this is their newest flagship model, tuned hard for software engineering and multi step agentic work. So the headline claim is this. XAI says that it uses four times fewer tokens than comparable models on average. So although this isn't something that's got that's benchmarking off the charts, it is somewhat comparable. You know, as an example, I'm looking at their announcement post here, on, you know, DeepSuite one one, which is, I think, one of the best, software engineering benchmarks. You you know, it's not great, but it's, you know, at least in the tier with, you know, Fable, g b d five five, and Opus 4.8.

Jordan Wilson [00:14:08]:
Same thing with terminal bench. It's much more competitive. It's actually ahead of Opus 4.8, and pretty close there to g b d five five and Fable as well. And then Sweebranch Pro, same thing, a little ahead of the, Opus models in, GPD five five there. So great, model specifically, I think, for software engineering. So here is what Grok says. They said, today, we're launching Grok 4.5. Yeah.

Jordan Wilson [00:14:37]:
There we go. Space x AI's see, is this x AI or Space x AI? They say it both ways. Anyways, this is, Space x AI's smartest model built to excel at coding, agentic task in knowledge work. It is our strongest model ever and was trained alongside CURSUR. More on that here in a second. So they say real world engineering excellence, RAC 4.5, was trained on datasets spanning knowledge in coding, science, engineering, and math with both intelligent and efficient reasoning. ROC four five excels at real engineering tasks and exceeds comparable leading models at these tasks. So a little bit more to know.

Jordan Wilson [00:15:19]:
Well, right now, where can you use this? So, well, you can use it in the Grok app, but the access depends on if what plan you're on. So free users get, like, hardly nothing. So, if you're on a paid plan, you will have some limited use of GROC 4.5 or if you're using GROC build. So why is it useful? Well, hey. We'll see how the third party benchmarks, look. Right? I'll be looking at artificial analysis. Cost to run the test is always one of the most telling benchmarks. But if this is true, if you are able, at least on the software engineering side, able to get, you know, Opus, you know, five five level, output level intelligence at, you know, about a fourth of the cost.

Jordan Wilson [00:16:08]:
That's pretty big. Right? So this is probably not gonna be a model that, you know, your whole team jumps in on and you're using it for all kind of things. But this could, in the future, aim to be a specialist dev software engineering model, especially with Cursor. Right? So SpaceX, acquired, Cursor as part of, you know, x AI, SpaceX going public. And the thing that they have that's a huge advantage is they have compute. Right? Them and Meta have tons of compute. They're one of the probably the only, you you know, quote, unquote, AI companies that just have compute to spare. That's why they're running it out.

Jordan Wilson [00:16:43]:
They've become kind of a NeoCloud. But this and that combined with Cursor is going to be a big deal. So Cursor working with, the xAI or SpacexAI team. I'm sure that Grok and, Cursor's model composer, 2.5 are gonna start to intertwine and, you know, learn from each other. But I I would probably be more interested in composer three, than Grok 4.5. But regardless, for a lot of people, if you're out there building, you do at least have a new model to take a look at. Alright. We have a lot more to get to, but first quick break for a word from our partners.

Jordan Wilson [00:17:24]:
Here's what most AI tools still can't do, work outside their own little box. The all new Slack bot just changed that. It's your AI teammate inside Slack, and now it can read, write, and act across the other apps your team already uses. No more tool switching. No more reexplaining yourself every time you open a new tab. One ops team at Engine says the summary feature alone saves them fifteen to twenty minutes of use. See what Slackbot can do at slack.com. Alright.

Jordan Wilson [00:17:58]:
Let's keep it going for our Friday features on Thursday. Yeah. During during that ad read I don't know if you've ever done this where you have an old coffee, but you didn't know it was old, and then I almost spit it out. Yeah. That that'll wake you up. Alright. Our next AI feature, Notion agents, a dedicated iPhone app. So here's what's new, why it matters, and who's gonna find it valuable.

Jordan Wilson [00:18:23]:
So this is a brand new standalone iPhone app from Notion, which is separate from the main Notion app. And this is out now. It is released and is built purely for running Notion's AI agents. So you can capture by text, voice note, or photo, and then answers, agents answer from your docs and connected tools and can take actions like creating pages, drafting updates, and running team workflows. And this is obviously model flexible, so it can work with Claude, Gemini, ChatGPT, whatever your team uses. So who has access? Well, it is live on the App Store. So it is iPhone first right now, and it does run on iPad and Mac, but isn't really designed for them and no Android announcement yet. So no pricing dedicated pricing, but if you do have a paid, you know, your company has a paid, Notion plan, you're you're in.

Jordan Wilson [00:19:20]:
Right? So, it does Notion's agent in AI features, for the most part, do require a paid plan, and those go anywhere from, you know, about 10 to $20 per user per month. So it does tie into your existing paid plan. So why is it useful? Well, it lets you fire off workspace tasks and capture ideas from your phone instead of waiting to get to a desk. And then the voice and the photo capture kind of kill the friction of logging things on mobile. So who's gonna find this valuable? Well, if you are a team that uses Notion, and Notion has been, I think, crushing it on the AI side. Right? I think Notion and Slack are two companies that, you know, a year and a half ago, if you would have said, hey. Would they be in this kind of, like, second tier? I would have said probably not. But sure enough, you know, I think Notion, Slack, and some others are, you know, really creeping up in terms of capabilities and what they can do.

Jordan Wilson [00:20:12]:
So, you know, if your team is on Notion and you haven't started yet taking advantage, of their agentic features, they're actually really, really good. Right? And, met with couple weeks ago when I was out in San Francisco, actually, met with the Notion team out there. And, you you know, they're they're working on some, pretty capable, offerings, so it is something to take, to keep an eye on. So, although I I I'm not super sure. Right? I'm not a a power Notion user myself. I've used it, you know, off and on throughout the years. I'm not certain about the separate app. So I'm not sure why it's not just built into their normal app and then having a, an agent section.

Jordan Wilson [00:20:55]:
I'm sure there's some, you know, technical reason why. But we'll see, if that sticks, if it gets popular, if it doesn't, etcetera. Alright. Next. This one, again, this might be for a smaller audience, but it's something that I think could make a big deal. So Google Voice AI. Yeah. So I think for a lot of small business, small business owners, if you don't know Google Voice, it is a way to just get a, phone number, a phone line for your company in, like, two minutes and four.

Jordan Wilson [00:21:29]:
Super cheap. And now they're rolling out a ton of new very helpful AI features. So Google Voice has added Gemini powered take notes for me. Right? So you can just tap notes mid call and it records and transcribes everything automatically. And then after the call, you get a summary with key points and action items plus the full transcript and audio, and then the notes arrive as an email link. So here is what Google, said about this new, feature. So, they say Google Voice launched over fifteen years ago as a call forwarding and texting service within The US. Since then, it has evolved into a trusted global communications platform for users and businesses alike.

Jordan Wilson [00:22:13]:
We're taking the next step in that evolution. We're bringing the power of AI note taking directly to your phone calls and introducing new plans designed to help users scale their businesses with a professional grade phone system at an industry leading price. So this does require a paid Google Voice standard or premier plan. So the free starter tier does not include it. Right now, it's English only for users in The US. I'm sure that will change here soon. Alright. So, well, why is it useful? So if you which I do know a lot of small businesses that use, Google Voice.

Jordan Wilson [00:22:50]:
Right? So I guess the downside is, well, you they it's been behind some of the other, you know, VoIP, providers, out there that have had some AI features, before them. So not saying this is going to be a differentiator for businesses, but especially for small businesses that are out in the field, they have physical locations. This one's a a big one. So, you know, it is able to, you know, turn all those phone calls into searchable records and, you know, ready made follow-up task with zero manual note taking. So the standalone pricing also, you know, if you're a solo operator, you have a side hustle, small business, I think it's a way to cheaply, bring some AI, especially if you're, a type of business that's on the phone all the time. Alright. Our next AI update. Yeah.

Jordan Wilson [00:23:41]:
Chad GBT images, their new images two point o model, that's been absolutely dominating since it came out. No one has been able to touch the, OpenAI GBT images two model when it comes to all benchmarks, but they may have company near the top, not passing them yet, but at least now a more serious, contender. So this is ByteDance, their new model, c dream five point o pro. So this is ByteDance's new image model built to understand design, not just generate images. So that at least from some of the samples that I've seen, this is seems to be kind of where they're pushing heavy in. So adding that understanding or reasoning layer to image generation and not just having it be, you know, this diffusion technology that just decides what pixels go where. Some of the more impressive examples that I'll kinda scroll to for our our livestream audience here. You know? So as an example, showing, like, the six major types of tea.

Jordan Wilson [00:24:42]:
Right? But it has to be able to reason and think, and break that down. There's a brewing temperature breakdown, you you know, kind of a, oxidation level. So there's a lot of reasoning and logic and thinking and planning that goes into some of these images. Right? So if you think about things that might be used as an infographic for your company, you know, just strong imagery that can really support a more visual story, this, you know, is probably now in that top three, I think, you know, along with Nano Banana, Pro and GBT images too. You now have seed dance, or sorry, seed dream, not seed dance. Seed dance is the video. Seed dream is the photo version. Alright.

Jordan Wilson [00:25:24]:
So it also adds that data in infographic visualization from dense text, interactive precision editing, and stronger photorealism, and in, native generation in 10 languages. But Here's what most AI tools still can't do, work outside their own little box. The all new Slack bot just changed that. It's your AI teammate inside Slack, and now it can read, write, and act across the other apps your team already uses. No more tool switching. No more reexplaining yourself every time you open a new tab. One ops team at Engine says the summary feature alone saves them fifteen to twenty minutes of use. See what Slackbot can do at slack.com.

Jordan Wilson [00:26:17]:
One of the standout features, I think, is a layer separation piece. So it edges toward it being kind of like a design tool, not just like a prompt box. So if you've ever used, you know, Canva or Photoshop and having all these different layers and being able to adjust it via layers versus as a flat image is gonna save a lot time and give you more granular control. So this is announced, and it is available inside ByteDance's own apps right now. So these are probably not apps that you use, but it's in DreamMania. And then for, Chinese users, it's in Jimmeng. So for international users, you can download the, Dream Mania app from ByteDance. Right now, there is no public API available yet, but there are some third party platforms, you know, that, you know, are conglomerates where you can use a bunch of the different I think fall is one of them.

Jordan Wilson [00:27:14]:
Is it fall or fall, f a l, or is it f a I? Right. So there's a bunch of these third party ones that have it. So if you really do want to, get your hands on it, if you use one of those third party, you know, AI image generators that have model selection, you're probably gonna see it in there. So, why is it useful? Well, like I said, it it is instantly a top three, maybe top two image model. So it turns data and dense text into clean on brand layouts. So if for whatever reason, if you can't use chat g b t or if you are a heavy into marketing, design, advertising, and you're not getting the results that you like out of nano banana or GBT images too, you now have another player to at least try. So also, multilingual teams, if you need text accurate output across different languages, designers, marketers, content teams who are producing, you know, infographics and branded visuals, this is definitely a new AI feature you can use now that you're gonna wanna check out. Alright.

Jordan Wilson [00:28:19]:
And then last but not least, Meta is back. Back again. Meta's back. Tell some friends. Alright. So let me just get this out there. Meta, when they came out with their, you know, they have their MSL, their Meta Superintelligence Labs. They spent, you know, billions of dollars hiring all this talent, billions of dollars on compute, and then, you know, everyone thought, oh, we're gonna get a new llama.

Jordan Wilson [00:28:46]:
Instead, they went a different direction with the muse, that is their large language model. And there are some rumors that they're coming out with a bigger version of it. So we have MuseSpark, which is presumably, that came out about two months ago, and that's their smaller model. And, actually, at the time it came out through a lot of benchmarks, it was maybe like a top four, five model. Not anymore. Right now, it's probably in the, you know, eight, nine, 10 range after a lot of new updates from all the other big players. But you can knock Meta for going from, like, zero to top five with their first large language model in Muse. So they've actually done the same thing with Muse Image and Muse Video.

Jordan Wilson [00:29:26]:
Because if you're looking at Arena, right, which is kind of the blind taste test for image and video models, these are top three models. Right? So going from zero, never having at least under the Muse family, I know they did, have, you know, their image models previously, but, you know, they scrapped everything, started, you you know, gotten rid of the the llama baseline. So pretty pretty interesting results for our, you know, livestream audience. You can see some of these examples. We shared them in our newsletter. Pretty good both on the photo and the video side. So here's a little bit what it is, what's new. So Muse Image is Meta Superintelligence's Labs first or MSL.

Jordan Wilson [00:30:07]:
Their first in house image model built as an agent that can use search and code execution to self refine its output. So kind of the, maybe controversial but signature feature is you can generate or manipulate images of a public Instagram account just by tagging it. Right? So you kind of had this originally in Sora from OpenAI, which is no longer r a p Sora. They had this thing called Cameos. Right? But people had to open up. Right? Like, yes. You can use my Cameo, but it looks like if you're a certain level of public Instagram account, people can do it just by tagging you. Alright.

Jordan Wilson [00:30:51]:
So Muse Video is built on the same foundation with native audio, but is still in development. So no release on that. So this is, on the images side, it is free for everyday creation if you are using the Meta AI app, or Meta dot AI inside Instagram, WhatsApp, Facebook Messenger, all those things, you should see, the new, Meta Muse image model available. But heavy users are going to pay. So there is a kind of a free daily usage threshold, and then it rolls into the new Meta page subscription plans. They didn't, you know, disclose the exact limit or price. Again, this is pretty new, but there's no public API right now. So, if if you're not on a paid Meta plan, which I don't know anyone that's on a paid Meta subscription.

Jordan Wilson [00:31:45]:
I know they just rolled that out a couple of months ago. You know, unless your business, like, runs on Meta in Instagram, I don't know anyone that's actually subscribing to Meta for AI features. Right? So it will be, worth tracking when and if they come out on the API side because at least the samples of the video model and the photo model look really good. Right? But if it's fairly limited to what you can use for free, I don't know how much traction this is gonna get. Right? Especially when, presumably, we may be getting within, you know, three to six months, We may be getting brand new updates to GBT images and to Nano Banana. So at that point and, you know, we just talked about, the new, you know, image model, seed dream from ByteDance. So I don't know. If I was meta, I would have came out a little more hard, like, hey.

Jordan Wilson [00:32:38]:
Use this thing. Use it a ton because the results are pretty good. And like I said, arena, I believe once they were released, they were both top three, top four, you know, image and video models. So fairly impressive. So why is it useful? Well, number one, it's free right now, limited use, and it lives inside apps that billions of people already use. So this is no new tool to to learn. It's just in the apps that people use. Also, interesting to see how this plays out, but the agentic self refinement aims for just kind of more accurate, less hallucinated images.

Jordan Wilson [00:33:13]:
So I think, you know, the people or the groups that are gonna find this most valuable, especially early on, is if your job right now requires, right, if you're a social media manager, social media marketer, then you obviously are living inside and, you know, living and breathing on your, you know, meta metrics. That's this is gonna be for you. Right? So versus having to work with an external tool, Meta's obviously made this, you know, this earlier technology available in their ads. I would assume that we'll see that here soon as well. So if you are a content creator, social media marketing, social media advertising, and you lay heavily, on any of the meta, you know, Instagram, Facebook, WhatsApp platforms, this is gonna be one worth checking out. Alright. So that is a wrap y'all. Our seven big features.

Jordan Wilson [00:34:05]:
Let's do them in reverse order. So the new MetaMuse image and video models, ByteDance's new c dream five o pro on the image side, Google Voice AI, great for small businesses, Notion agents dedicated iPhone app for, teams using, Notion. Grok 4.5, new model could be good for software engineering. OpenAI real time 2.1, the update to GBT real time two point o can do a lot of things agentically. And then last but not least, which I think may ultimately end up making the biggest difference for millions of business users is the new GPT live model. So I'll probably end up doing, a review on that. But thank you for tuning in. Thanks for the curve ball again.

Jordan Wilson [00:34:54]:
Right? Last week, we had the, July 4 observed holiday. This week, we're switching it up. So tomorrow, we can do the show on Soul, g p t five six Soul. So make sure to join us for that. Thank you for tuning in. Hope to see you back tomorrow and every day for more everyday AI. Thanks y'all. Here's the thing about AI at work.

Jordan Wilson [00:35:19]:
You can't use what you don't trust. The all new Slack bot runs inside Slack security boundary, only sees what you've already allowed it to see and never trains on your data. Personal AI agent with the trust to actually put it to work, Slackbot from Slack. Learn more at slack.com.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI