Ep 435: How 50X cheaper & faster AI transcription is changing enterprise work

Revolutionary Advancements in AI Transcription

Years ago, human transcription was preferred for its high accuracy, especially when dealing with names, places, and unclear speech. However, the field of artificial intelligence (AI) has seen rapid development recently, particularly in the area of Automatic Speech Recognition (ASR). One such groundbreaking invention in this domain is Whisper by OpenAI.

Whisper is an open-source ASR model that offers high accuracy in transcription, thereby greatly reducing the number of errors made compared to earlier models. Furthermore, it supports multiple languages and is MIT licensed, meaning it can run on a wide array of platforms, unlike some proprietary models.

Technology Advancements: From V1 to V3 Turbo

Since its initial release, Whisper has undergone several changes and enhancements. With its improved versions, V3 and V3 Turbo, the transcription model has dramatically increased in both speed and accuracy. Whisper V3 Turbo, in particular, has aided in near-instantaneous transcription processing, a feat previously unheard of.

The advancements in real-time factor mean that audio can be transcribed significantly faster than real-time, making it advantageous for various applications. Such innovations highlight how technology is continually streamlining and optimizing processes.

Lived Experience: A Deeper Appreciation

Many businesses today recognize the challenges and time consumption associated with manual transcription. The advent of advancements like Whisper not only decreases the time taken for transcription but also drastically improves its accuracy levels.

However, the industry is still exploring AI’s ability to capture tone and inflection in speech, hinting at potential areas for future growth and innovation.

Evaluating and Preparing for Future Improvements

In rapidly evolving fields like AI, it's crucial for businesses to stay ahead of the curve by continually assessing their current systems and practices. If current AI models don't serve your needs, keep a close eye on upcoming developments. Don't shy away from building prototypes or evolving current models. This practice will allow your business to be prepared when game-changers arrive in the market.

The Bottom Line: Unlocking New Business Potentials

Transcriptions convert spoken words into text, a format that is far easier to process by both humans and AI. This transition opens up a wealth of information previously trapped in audio format which can be used for search functions and data analysis, among other applications.

Advancements in AI models have also resulted in dramatic cost reductions, making transcription services more accessible than ever before. As a result, a significant business potential lies in unlocking, analyzing, and integrating large audio data into existing systems.

The future of work is pointing towards live advancements in AI, such as real-time transcription and voice modes. New technologies are enabling more dynamic and interactive business communications, and it's only a matter of time before these tools become internalized within industry practices.

The industry's disruption through AI transcription could potentially redefine fields such as court reporting, reducing reliance on human transcribers. Nevertheless, human oversight will always be a crucial facet in ensuring transcription accuracy, particularly in sensitive areas like legal proceedings.

Fasten your seat belts as we embark on this exciting AI journey, reshaping the landscape of modern enterprise work.

Topics Covered in This Episode

1. AI Transcription Benefits
2. Whisper Model by OpenAI
3. Cost of Transcription
4. Business Applications for AI Transcription


Podcast Transcript


Jordan Wilson [00:00:17]:
Every word that you say, every meeting, every speech, that's gold. I think, so oftentimes when we get caught up in implementing generative AI in our business, we think about other large language models that exist. Right? And we think about, like, oh, we're we're limited by their training data. You you know, hey, hopefully these models get better, but what about your words? What about all of those meetings? What about that big seminar that you're speaking at? That is unstructured gold. I think, something that we don't talk about enough on this show or in general is how the words that we speak, the conversations that we have, how valuable those are, and how, the AI surrounding that is getting cheaper, faster, more accurate, and what that really unlocks for businesses of all size. Alright. I'm excited to talk about that and a lot more today on everyday AI. What's going on y'all? My name is Jordan Wilson.

Jordan Wilson [00:01:26]:
I'm the host of everyday AI. This thing's for you. It is your daily livestream podcast and free daily newsletter, helping people like you and me, us everyday people, catch up with everything that's happening in the world of AI and how we can use all this information to grow our companies and our careers. Is that you? If so, welcome home. And your other home is our website, your everydayai.com. So, if you find value in today's conversation with our guests, we're gonna be recapping and sharing a lot more insights in our daily newsletter, as well as keeping you up with everything else that's happening in the world of AI. Also, there's like, I don't know, a 1000 hours of audio content text on there, exclusive interviews from the smartest people in AI in the world, all for free on our website. Alright.

Jordan Wilson [00:02:12]:
Before we get started, let's first go over the AI news. So Anthropic is set to close a $2,000,000,000 funding round as its valuation source to 60,000,000,000. So Anthropic, one of the biggest startups in the generative AI space, is reportedly nearing the completion of a $2,000,000,000 funding round led by Lightspeed Venture Partners. So, this investment will significantly increase its valuation from 18,000,000,000 last year to an impressive 60,000,000,000. So the latest funding round is part of a broader $6,000,000,000 initiative for Anthropic followed by an earlier $4,000,000,000 investment from Amazon. So, yeah, they've, I think raked in, like, $8,000,000,000 in commitment so far in this round. So in philanthropic's annualized revenue has reached an approximate $875,000,000 driven by its model of selling access to its advanced, API or sorry. Its advanced AI systems to enterprises and through platforms like Amazon Web Services.

Jordan Wilson [00:03:16]:
So, who knows? Maybe with this extra, cash that Anthropic just, you know, put in its pocket, maybe their rate limits will go from unusable to kind of usable. We'll see. Alright. Next, NVIDIA CEO Jensen Huang is claiming that his new AI chip performance on his GPUs surpasses Moore's law. Yeah. Yeah. We're breaking science breaking science in the face. So, the NVIDIA CEO stated in an interview that their latest data center superchip is over 30 times faster for your AI inference workloads compared to its predecessor, which could significantly lower the cost of running AI models.

Jordan Wilson [00:03:56]:
He emphasized that by innovating across the entire stack, architecture, chip design, design, systems, libraries, algorithms, etcetera, NVIDIA can achieve advancements at a pace that exceeds Moore's law. So Wong introduced the concept of hyper Moore's law. Yeah. Now we gotta learn new scaling laws, suggesting that AI development is not slowing down, but is instead governed by 3 active scaling laws, pre training, post training, and test time compute. So Wong also claimed that NVIDIA's AI chips today are 1,000 times better than those produced a decade ago, indicating a rapid evolution in technology that could benefit various industries. Alright. Last but not least, Apple is facing a ton of backlash over its inaccurate AI news alerts and has promised an update. So Apple is under scrutiny after its AI feature that essentially summarizes news alerts is generating some false and misleading news headlines raising concerns about the accuracy of information in its new Apple Intelligence.

Jordan Wilson [00:04:59]:
So Apple announced that it will release a software update in the coming weeks to clarify when news notifications are generated by its AI system known as Apple Intelligence. So the misleading alerts have sparked criticism from various media organizations, including the BBC and ProPublica, which reported similar inaccuracies in AI generated summaries of their content. Alright. A lot more on those stories and everything else you need to say ahead, not just keep up. Stay ahead on our website. So make sure you go check check that out and sign up at your everydayai.com. Alright. Enough chitchat.

Jordan Wilson [00:05:37]:
Let's get to the bulk of today's conversation. AI transcription. You probably don't think about it, but it is a boon for business. So I'm excited to have this conversation. Hey. Live stream audience, help me welcoming to our show. We have Philip Keeley, the head of developer relations at Base 10. Philip, thank you so much for joining the Everyday AI Show.

Philip Kiely [00:05:58]:
Hey, Jordan. Thanks for having me. Super excited to be here.

Jordan Wilson [00:06:01]:
Let's let's chat about transcriptions. Before we do, can you tell everyone just a little bit, about Base 10, what it is you all do?

Philip Kiely [00:06:09]:
Absolutely. So Base 10 is an AI infrastructure platform. We take open source, fine tuned, and completely custom models for our customers, and we help them deploy those models on worldwide auto scaling GPU infrastructure. We also assist with the model performance efforts so that we can get them, you know, lower speeds, higher throughput, lower cost, better quality. Our customers are AI native startups and enterprise like Rydo, Bland, Patreon. And one thing that we've been working a lot with recently is the Whisper model. We recently released the world's fastest, most accurate, and cheapest Whisper inputs.

Jordan Wilson [00:06:46]:
So let's I mean, I want I do wanna dive into Whisper, and I'm sure it's something that a lot of our audience is familiar with. But before I even go there, what's the main benefits. Right? Like, you know, when people talk about transcription and I kind of started the show out on it, I'm a firm believer. Right? Every word I speak on this podcast, it's instantly transcribed and fed into a large language model. But what's the benefit of of capturing your company's words, and and using those? I think sometimes people just overlook it.

Philip Kiely [00:07:19]:
Yeah. Well, it's just another stream of data. So if you think about all of the YouTube videos in existence, all of the podcasts in existence, all of the phone calls, that maybe have been made into your company's call center, there's just tons of data floating around out there that takes a long time to process. You know? Maybe if you're some sort of super speed listener, you can listen to a podcast on 1 and a half or 2 times speed. But when you think about how fast a human talks, we only speak at, you know, maybe up to a 150 words per minute. I know I'm not supposed to actually speak that quickly when I'm doing a podcast, so I'm always trying to slow it down a little bit. Maybe you listen at 2 x speed, you're getting, what, 300 words a minute. But if you think about how fast someone can read, you know, the fastest speed readers can read at 500 or even a 1000 words per minute.

Philip Kiely [00:08:10]:
So audio is actually a fairly low signal channel. There's there's not a ton of bandwidth, in talking. But if we can transcribe that audio and then we can get it in text, not only is it much easier for us to process as people, we can read a lot faster, but it's also easier for machines to process. You know, we can feed it into large language models like you said, or we can do simple find and replace. We can do simple search. There's a ton of things you can do on text that is really hard to do on audio.

Jordan Wilson [00:08:39]:
Mhmm. It's it it is, you know, I I hate floating around the term like game changer. Right? But, it is. Right? Being able to capture everything that's said, you know, I like to say that is your first party or first company gold, all the words that you talk about. Live stream audience, thank you for joining us. You know, if you do have any questions on AI transcription on on what that means for your business, get them in for Philip now. But maybe let's not, whisper, but let's talk about whisper. Philip, what what the heck is Whisper?

Philip Kiely [00:09:13]:
Yeah. So Whisper is an open source model that was created by OpenAI a couple years ago. And I'll actually give, like, a kind of little history lesson here. So in 2019, I was working on a blog post about speech to text, which can also be called transcription. It can be called ASR, which is automatic speech recognition. And I was kind of doing a survey of the state of the art. And one of the best things I found back in 2019, was something called Amazon Transcribe. It's like an AWS thing, and it was pretty impressive back then.

Philip Kiely [00:09:48]:
You know, it was able to take some, some segments of text, and it was able to create, you know, a reasonably interesting transcript out of them. But there was definitely a ton of errors, especially around things like names, places, popular nouns, as well as just if I kind of mumbled a little bit, then it really didn't know what was going on. And so, actually, a, a year later, I was working on a book. And when I wrote that book, I did a ton of different interviews with experts in the field. These were audio interviews that I needed to transcribe, and I ended up having to transcribe them by hand because I did all of these, you know, I did the survey of all this technology. It, it wasn't really good enough for for, you know, publication. And so I just spent, like, a month at the keyboard typing out these 50,000 words, from these expert interviews. So, you know, I I've always kept my eye on the space since then.

Philip Kiely [00:10:44]:
You know, when open source models like wave 2vec came out, I was really excited. I wanted to try it, but nothing really approached the quality of my, you know, amateur, but still human transcription. So September 21, 2022, OpenAI released a model called Whisplote. And what's really exciting about this model is it's actually MIT licensed, which means you don't have to go through the OpenAI platform to get it. You can run it on your computer. You can run it on a cloud service. You can run it wherever you want. And the first Whisper model was really exciting because it offered much higher accuracy.

Philip Kiely [00:11:19]:
Also, it offered that accuracy across a bunch of languages. So when we talk about ace ASR model and accuracy, we wanna think about WER, which is word word error rate. So for how many you know, for a 1,000 words, how many of those words are gonna be wrong? You want that word error rate to be as low as possible. And so this model came out, it's got word error rates of, like, 10. You know, maybe maybe 1% of the words are gonna be wrong versus, you know, much higher for for other, for other models. And since then, these models have gotten better. You know, now we're on Whisper V3 here in 2025. We also have Whisper V3 Turbo, which is a little less accurate than V3, but much faster.

Philip Kiely [00:12:04]:
So we're we're able to get, you know, faster and more accurate transcription from these open source models, in a lot of different languages.

Jordan Wilson [00:12:12]:
Yeah. And and what you said there, I don't know if anyone else in in our audience, that that hits them, but that hit me because I remember, right, I was a journalist back then. So I literally had taped interviews on a little tape recorder. Right? I had one that was digital, but I think early on, it was an actual tape not to date myself. And I remember hitting play, stop, rewind so many times because especially when you're quoting people for big news publications, you had to get every single word right. You know, I'm I'm even curious as someone that did this as well. What was your first, reaction to you seeing something like Whisper back in 2022? What what what was your reaction when using it at first?

Philip Kiely [00:12:54]:
My I mean, my first reaction was, man, I wish I had this a couple years ago because, you know, I I, my fingers were hurting. I had my mouse on the floor so I could kick it with my toe to start and stop the audio recording. I was thinking, wow, my life could have been so much easier if this had been released a couple years ago.

Jordan Wilson [00:13:17]:
Hey. This is Jordan, the host of Everyday AI. I've spent more than a 1000 hours inside ChatGPT, and I'm sharing all of my secrets in our free prime prompt polish ChatGPT course that's only available to loyal listeners like you. Listen to what Lewis, a business owner, said about the PPP course.

AI [00:13:36]:
I can tell you that when I went in, I I understood a little bit about ChatGPT. I understood some of the stuff. I was able to use some of the prompts, but what I discovered going through, Jordan's webinar was that there is so much more I don't understand that ChatGPT can do, and I really should be using it. And, if anything, I got that from the webinar. I would highly recommend this to anybody from beginner to advanced. You will absolutely learn something from this from this experience.

Jordan Wilson [00:14:02]:
Everyone's prompting wrong, and the PPP course fixes that. If you want access, go to podpp.com. Again, that's podpp.com. Sign up for the free course and start putting ChatGPT to work for you. So, you know, when when we talk about some recent advancements. Right? Because, yeah, I even remember I I used Whisper when it first came out in in 2022, and I I I didn't think it was slow. Right? But now when I'm using it because, yeah, I I run it locally. I have, you know, plenty of of programs that that run-in on the back end as well.

Jordan Wilson [00:14:40]:
Now I'm like, oh, wow. It was slow. What does the the recent speed and the cost, right, when we look at Whisper, V3 Turbo, you know, maybe whenever we see a Whisper V4, what do these advancements actually mean when it's faster and cheaper? Yeah.

Philip Kiely [00:14:58]:
So when we think about speed and cost with Whisper, we talk about real time factor. So if you have, say, an hour of audio, how many times faster than real time can you transcribe that? And my real time factor as a person is, like, 0.3 or something, 0.2. It takes me 4 or 5 hours to type out an hour of audio because I'm constantly starting and stopping it and going back. Maybe if I was a faster typer, maybe if I was a professional, I could go a lot faster. Out of the box, you know, Whisper might get you to, depending on the hardware you're using, I don't know, 50 times, a 100 times real time factor. So maybe that hour of audio, you're able to transcribe it in a minute, and that unlocks a ton. But you're actually able to take it way further through various optimization techniques that we can get into, and you can get that real time factor all the way up to, say, like, a 1000 times where instead of that hour of audio taking a minute to transcribe, it might only take, you know, 5 or 6 seconds. And the other factor in, performance optimization is, you know, if you're trying to do some kind of streaming use case where you're transcribing the audio not as a file after the fact, but live during the conversation.

Philip Kiely [00:16:16]:
And so for that, you care about the round trip latency for a single 30 second chunk of audio. And for that, you can get down to about 200 milliseconds. So I'm a martial artist. For me, reaction time is super important. I don't have the best reflexes in the world, but, you know, the average reaction time for a human is is about 200 milliseconds. And so if you're able to process that audio round trip in the time that it takes someone to sort of, like, react to something happening, then to your end user, that's going to feel like it's basically instant.

Jordan Wilson [00:16:53]:
A a lot of good, comments here from our from our, livestream audience and a couple of questions too. So, you you know, Samuel's asking, is there any effort to capture tone and inflection during transcription? Spoken language has a lot of context components beyond grammar and vocabulary. That's something I was thinking myself, Sam. So thanks for that question, Philip. Are we gonna see that in in future AI transcription? Right? Like, I I sometimes talk very quickly. Sometimes I talk with emotion. Right? Like, is that something that future AI transcription will be able to tackle?

Philip Kiely [00:17:26]:
That's a really good question. Emotion inflection, that kind of stuff is more of a factor right now when we're going in the other direction, When we're going from text to speech and we want an AI model to be able to do speech synthesis, there's a lot of work that's been put into, you know, making that sound much more natural, and that's where, you know, those context components in spoken language are super important. Generally, right now, when we're going the ASR route, when we're going from speech to text, that is going to be just the sort of raw contents of the of the file, or the raw contents of the conversation, but that would definitely be super interesting to to look at. Like I said, it's it's a big area of research going in the other direction, but it's not such a big factor right now in transcription.

Jordan Wilson [00:18:20]:
What what has, you know, what have all of these updates, done to cost? Right? Because I, yeah, I remember even originally, I was happy to pay, you know, you know, a dollar an hour or whatever it was, you know, in the earlier days of of, you know, kind of AI transcription. What is the cost now? And, you know, what does that mean in the grand scheme of things as businesses are trying to leverage all of this data? Right? They're recording Zoom meetings very common now. Right? I I I think people have this, you know, gold mine of data that they're may be sitting on. So can you walk us through the cost changes and then what that actually means?

Philip Kiely [00:18:58]:
Absolutely. So, you know, a couple a cup a few years ago, you're looking at a dollar or 2 per hour of audio, and that's generally how it's measured is. How how much input time are you putting in, that that's how much you're paying. So if you're putting in a 1 hour audio, say, like a podcast and you wanna get back a transcript, it's gonna cost 1 or $2. But today, it's gotten a lot faster. And when AI models get faster, they also get cheaper. The sort of thing that makes an AI model expensive to run is that you have to run it on a GPU. GPUs are very expensive.

Philip Kiely [00:19:32]:
So if you use less time on that GPU to accomplish the same task, then that price goes down. Today, you're able to do these transcription jobs for it it, you know, it depends it depends on exactly how fast you want it to run. It depends on the exact type of transcript you're trying to generate. But if you're doing the simplest, most basic transcription and you're okay with, you know, waiting a couple extra seconds for it to generate, you can get down to just a couple cents per hour. So we're looking at, you know, a 50 to a 100 x reduction in the cost of doing this transcription, and that's massive. You know? Now for the same price that you were transcribing 1 hour of audio before, you could transcribe 50 or a 100 hours, and that just unlocks so much for business.

Jordan Wilson [00:20:19]:
Yeah. And and speaking of that, let's dive into it because, I I I still think this is one of those areas just like I started the show off. I think, you know, so many when we talk about business use cases and advancements in generative AI and large language models, Right? I think everyone looks at using a ChatGPT, a Gemini, a Meta Lama. Right? Like, people look at using these models, but they don't necessarily look from within. Rather creating, which a lot of times is meetings. It's conversations like this. Right? Can you talk a little bit about, maybe some new and exciting business use cases that have maybe just begun, to be become a little bit more unlocked because of that cost and that speed?

Philip Kiely [00:21:03]:
Absolutely. So a big business is going to generate just so much audio. Lot of that's going to be internal. You know, sometimes you might not wanna transcribe literally every single thing that happens, but there are a bunch of places where it is really valuable. So one of those is, you know, any kind of customer facing situation. You know, if you're doing call center, if you're doing a, you know, teller service, anything where you are interacting with a customer and and from the customer perspective, you know, you get on the line there, and you hear, oh, this call may be, you know, monitored for quality assurance. So that quality assurance monitoring, historically, is like a a manual process. You have some supervisors who are maybe listening to a few calls and making sure that everything's going well.

Philip Kiely [00:21:50]:
Now you could just transcribe every single call that's coming into your business, and then you have a fully searchable database. You can do quality assurance. You can also maybe analyze those transcripts to figure out patterns and what your customers are asking for. You can do content moderation at scale. You know, if I post a, you know, some something with text on on a platform and it has, you know, some some stuff the platform doesn't want on there, that's super easy to to identify and and flag, that that I'm using it in words. If I'm posting, say, a podcast on Spotify or something, then that's a lot more difficult. Or if I'm posting a YouTube video because, you know, you you can't really just listen to all of the podcasts and all of the YouTube videos. But if you can get that from audio to text, then you can run it through those same, moderation algorithms.

Philip Kiely [00:22:42]:
You can also do stuff like media subtitling, closed caption generation. You can do that in real time. I know sometimes if I'm watching, like, a sports game on silent, I see the announcer's words, but it's always, like, 5 or 6 seconds after it

Jordan Wilson [00:22:56]:
plays after it. Far behind.

Philip Kiely [00:22:58]:
About it. It's so far behind. Right? And so if we can get that, you know, down to something that's more real time, that's super awesome. And you can also do real time translation with that as well. Yeah. So, yeah, there's just so many different, use cases where you have these massive volumes of audio being generated that before it just wasn't cost efficient to process these or just took too long. And now with this cheaper, faster AI transcription that's more accurate, you can get a lot more value out of these big audio, corpuses.

Jordan Wilson [00:23:33]:
So, you you know, Cecilia brings up a good point because there's entire industries, right, that for many decades have thrived around just typing words, what people are saying. Right? Like, she's asking about how is AI transcription disrupting industries like court reporting. Right? Like, are are are we gonna see some of these, you know, traditional roles where people were just transcribers? Are they just gonna go away?

Philip Kiely [00:23:58]:
Well, you know, you do still have to verify these transcripts. When I talk about accuracy in an AI transcription and that word error rate, you know, that word error rate is not 0. There was a lot that you can do to make your transcripts more accurate. You know, for example, you can look at you can you can have a model analyze them. You can look at, say, chunks that are silent and, you know, replace them or rerun them. But, you know, at the end of the day, if you're doing something like quote reporting where you need 100% perfect accuracy, it's important to have systems beyond just a single transcription model that are going to guarantee that accuracy. And, you know, I think that there's still a major law for human in the loop in these kind of systems where you're able to, you know, go in and and verify these transcripts and make sure that they're completely accurate.

Jordan Wilson [00:24:54]:
Yeah. So you you you talked a little bit, about how, this, you know, advancing technology whisper models, you know, in general, are helping change, how we've done business in the past. But as we look to how these advancements might change how we work in the future, what might we see change? Because everything's going live. Right? You you know, you have your live advanced voice mode from ChatGPT. You have Gemini live. You you know, you can talk to Copilot. Right? Like, how will more accurate, faster, cheaper, transcription change how we

Philip Kiely [00:25:35]:
work? So one thing with that, like, live, voice mode from ChatGPT is it's really cool, but it's also really expensive. Right? That, you know, that sort of capability costs, what, like, you know, 10:10 plus dollars an hour. And this transcription is is only a few cents an hour. So if you're a clever developer, you're able to kind of put this model in front of some other models and build these sort of chains of models, for these compound AI use cases where instead of having one gigantic model that cost a ton to run, and is able to do it end to end, you chain together a few small cheap models and run the same pipeline much faster and much cheaper. One place where that's really important right now is AI phone calling. So if you want to say, like, have a automatic pizza order take, that you're gonna build where a customer can call it up and just say what they want on their pizza, and it's gonna say, alright. I've got this pizza for you, that kind of thing. And you can build, you know, that AI phone calling with, these faster, cheaper transcription models.

Philip Kiely [00:26:41]:
Another big aspect is wearables. So a big trend right now is, you know, having a pin or some some speaker microphone combination on on your body that's able to sort of record your daily context, so that you have, you know, better information for your decision making, for that kind of stuff. And so if you're, you know, wanting to record your life 12 or 16 hours a day, again, if that's, you know, going through that historic transcription algorithm where it's costing a dollar an hour, well, that's, like, $16 a day. That's just not a sustainable business. But if you're able to do it for, you know, at while you're sleeping at night for for a couple of pennies and it's costing a few cents a day, then, you know, now we're in the realm where this can make sense as a consumer product.

Jordan Wilson [00:27:25]:
Mhmm.

Philip Kiely [00:27:25]:
So wearables, you know, local influence, phone calling, all these sort of things are these sort of real time multimodal user experiences. They'll get that are getting unlocked by these transcription models.

Jordan Wilson [00:27:38]:
Yeah. And and I do think that we are going to see those in the wild that actually makes sense. Right? If you've listened to this show, I'm I'm never one to just, you know, things like the humane pin and and the Apple Vision program. Like, no. Not really. But I think some recent advancements, right, the meta's, meta's Ray Bans, you know, some of Google's new products. You know, I think wearables are going to be a thing whether you think they're going to or not. I do think that is kind of the next iteration.

Jordan Wilson [00:28:07]:
But, you know, one thing I'm I'm curious about, and it's something I've always thought about this concept of typing versus talking. Right? Like, I can talk really quickly, but also I don't blame y'all if you listen to this podcast on 2 x. I would too. But, you you know, might we see something in the future where it it becomes less and less common to type and we're just interfacing with, you know, I don't know, autonomous AI agents and multi agent environments and all we're really using is our voice. And if so, you know, what part of this technology has to improve or what advancements are we kind of waiting on until that future is finally here where we're just sitting back, kicking our feet up and just talking, you know, to our AI agents?

Philip Kiely [00:28:56]:
Yeah. So the, the future is now actually for that. You have all those agent use cases and stuff, that are still coming. But if you just want to control a you know, control your computer, if you want to type an article without without using your fingers, that's actually possible. I have a, a colleague who actually had to have surgery, recently, on their hands. And so they went and, used a voice transcription app for a few days to do lighting while they couldn't type as much. They use something called Whisperflow, which which is a an application out there for that. But, yeah, it's, you know, the future is now, in terms of controlling your computer with voice.

Philip Kiely [00:29:43]:
It's not something that's gonna be practical in every situation. Like, if I'm on the train, I don't wanna be talking to my computer, and everyone else is talking to their computer. That doesn't sound so good. But it can definitely be helpful, if you have, you know, limited typing ability. I I don't type particularly quickly. I can definitely talk much faster than I can type. So it's something that I'm super excited about.

Jordan Wilson [00:30:07]:
Yeah. It's it's it's a good point. And I think, you know, having conversations about these type of things is important because I, yeah, I do think, yeah, whether we're talking wearables, whether we're talking, you know, you know, talking to your computer, it it is becoming more and more common more, I I I think part of how we work in the future. One other thing, you know, what what what part of this, Philip, like, why are, you know, if I'm talking to Siri, if I'm talking to Alexa, right, I see a big difference than when I'm talking to as an example, a Gemini live or a, you you know, ChatGPT advanced voice mode. Why is there still this kind of divide, even between the the big tech conglomerates on on which ones can accurately understand our words and sometimes they just can't?

Philip Kiely [00:30:53]:
So what you're observing there is the difference between on device influence and cloud influence. So if you're taking an AI model and running it on the user's device, that's on device of edge influence, and, you know, your your user device is not gonna be as powerful as, like, an NVIDIA H100 GPU sitting in a data center somewhere. It's not going to be able to run as big of a model or run the same model at as high of a quality. And so because of that, for, you know, these voice transcription things, you're probably seeing, you know, a little bit worse results when you're using it on a local device versus when you're using it on the cloud. However, that's changing really quickly. You know, these models are pretty small. They can be just a couple 1,000,000,000 parameters, and so those are actually a really good candidate for local influence even on stuff like smart speakers, or, you know, maybe that next generation of smart speakers that has those upgraded GPUs, upgraded VLAN capabilities so that they can run these small models. And so I I definitely think you'll see that gap close in the transcription space, pretty quickly.

Jordan Wilson [00:32:00]:
Alright. So, Philip, we've covered a lot in today's conversation. I mean, we talked a little bit about Whisper, what this technology is, the cost savings, how, you know, faster and more accurate, you you know, voice transcription AI has led to many new use cases. But, you know, as we wrap up today's show, what is the one most important thing, that you want our audience to know when it comes to how cheaper and faster AI transcription is changing enterprise work?

Philip Kiely [00:32:33]:
I think the most important thing to understand is the trend. You know, in the last couple years, these models have gotten much more accurate, much cheaper, much faster. And there was, of course, the massive leap from 2022 to maybe, like, a couple years before that. I think this is gonna keep happening. So even if you see a use case today where it's like, you know, Philip, actually, like, 5¢ per hour, that's a little too expensive for what I'm trying to do. Or, oh, you can only do 200 millisecond round trip time. Like, yeah, that that that doesn't cut it. We're not done optimizing these models.

Philip Kiely [00:33:07]:
And even in the, you know, last couple quarters of of work on these models, we've gotten much better at running them, being been able to run them much faster and much cheaper, and that's that's a trend that's continuing. So I would definitely, you know, look at these use cases that you're that you're considering today and say, okay. Does this make sense today? If yes, go for it. If no, still maybe go for it because it could make sense in 3 months, 6 months, 9 months, once the once the technology gets even better, and you're gonna be pretty far ahead. You know? You said, for example, Jordan, that you don't always love some of these wearables. You know, that's that's a case where having the prototype today is what's going to set you up to be able to use the, you know, polished version next year, for those companies. And so I'd say for in the same vein, if you're building some kind of speech use case, if you're building some kind of transcription use case, And if it doesn't work today, still build that prototype, put it in your back pocket, and keep an eye on the technology as it advances because it's getting better fast.

Jordan Wilson [00:34:11]:
Mhmm. That's great advice, and I think words that we should all listen to. Alright. So, Philip, thank you so much for taking time out of your day to join the Everyday AI Show. We appreciate your insights.

Philip Kiely [00:34:24]:
Hey. Thank you so much for having me. I had a great time.

Jordan Wilson [00:34:26]:
Alright, y'all. Quick reminder. We covered a lot, and there's a lot more. So if you found something valuable today, please, if you're listening on the podcast, make sure to subscribe and rate the platform. Go back and listen to our library of of episodes. Like, we literally have thousands of hours of content on our website, 100 of episodes. Also, go to your everydayai.com. We're gonna be recapping today's conversation.

Jordan Wilson [00:34:52]:
Yeah. I'm gonna upload it in in 10 seconds. I'm gonna have it all transcribed, but I'm gonna be writing about it, A real human telling you more info and insights to take away. So thanks for joining us. Hope to see you back tomorrow in everyday for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI