Ep 436: AI You Can Trust – How reliable data makes it happen

Resources:

Join the discussion: Ask Jordan and Barrr questions on AI and data


Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup

Connect with Jordan Wilson: LinkedIn Profile

Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course! 


Harnessing the Power of Reliable Data in Generative AI

In a world that has seen a historic shift in data utilization and an essential need for accurate and timely data, day-to-day individual services like ride-hailing and food delivery have taken the lead in data integration. The current wave in data usage trends involves generative AI, with an impressive 100% of data leaders engaged in some form of this emerging technology. Confidence in the data used for these AI models, however, is disappointingly low, with only a third of leaders trusting their data.

Reliable Data: Key to Competitive Advantage

High-quality data gives a competitive edge, significantly influencing the personalization of products in generative AI. A key concern, though, remains the slow implementation rate of AI solutions in enterprises, stressing the need for dependable data as a foundational element. The reliability of data holds immense importance; inaccuracies can lead to brand reputation damage and revenue loss.

Challenges and Opportunities for Small and Medium-Sized Enterprises

Small-to-medium size companies face hurdles when organising and leveraging their data for AI. The solution? Prioritizing the use of high-quality data over unreliable data sets. Interestingly, these smaller organizations could potentially have an edge over larger enterprises due to less complex data systems, allowing for swifter decision-making and movement.

Data Solutions and Adoption Forecast

Solutions like Monte Carlo are stepping in, enabling data teams to be ahead of data issues and thereby preventing unwelcome surprises. This company alone is associated with esteemed businesses like Fox, Roche, and Credit Karma. The field of data observability is growing rapidly, with predictions from Gartner stating that over 60% of organizations will adopt it in the next five years.

Benefits of Agile Data-Centric Teams

Small teams can efficiently utilize data to innovate more swiftly than larger organizations, experimenting and pivoting based on successful trials. Companies of all sizes often mandate trial runs prior to forming centralized strategies, providing a unique advantage to smaller teams.

Emerging Applications of Data and AI Integration

Good data has become a fundamental pillar for AI integration, influencing use cases across sectors. AI improves engineering productivity with coding assistants while enhancing the code review process. It's also implemented in compliance reporting in industries like pharmaceuticals, where it speeds up processes that usually take months.

Exploring Generative AI and Large Language Models (LLMs)

Generative AI significantly reduces time spent on manual report writing by training on past reports, thus generating drafts quickly. LLMs, for instance, can be used to analyze customer support chats and convert unstructured data into structured data with reliability scores. There is potential for synthetic data to supplement real-world data, as current sources are being maximized.

The Continued Importance of Quality Data

The success of generative AI depends heavily on the quality of data used; poor data results in poorly informed AI outcomes. This emphasizes the critical need for businesses to invest in data quality management to fully realize the returns on their AI investments.


The world of data and AI is rapidly evolving, and it's more crucial than ever for businesses to recognize the value of investing in high-quality data. It is the backbone for AI technologies like generative AI and LLMs. Stay updated with the latest trends in data technology and AI by signing up for a free AI-focused newsletter.

Data isn't just numbers; it's knowledge, power, momentum. Recognize its potential, invest in its quality, and drive your business forward with well-informed, data-driven decisions.

Topics Covered in This Episode

1. the Importance of Data
2. Challenges and Opportunities in Leveraging Data
3. Adoption of Data Practices
4. Data Use Case Examples
5.Generative AI, LLMs, and Data Integration


Podcast Transcript


Jordan Wilson [00:00:16]:
When it comes to generative AI, I think sometimes we just overlook some of the most important things. Right? And so we just wanna hit that big red easy button and have it spit out hours of work, and we're just like, yay. Good. We're done. But there's that one most important part. It's the data. Do you trust where your data's coming from? Is it reliable? What happens if it's wrong, and, well, why should you even care? We're gonna be talking about that and hopefully answering a lot of those questions today on everyday AI. What's going on y'all? My name is Jordan Wilson.

Jordan Wilson [00:00:56]:
I'm the host of everyday AI. This thing is for you. This is your daily livestream podcast and free daily newsletter, helping everyday people not just learn AI and keep up, but how you can leverage it and become the smartest person in AI at your company. So if that sounds like you and what you're trying to do, this is your new home. If it's the first time, please make sure if you listen on the podcast, check out your show notes. In there, you will see a website, your everyday AI.com. But before we get started, have to first give a quick shout out to our partners at Microsoft. So why should you listen to the WorkLab podcast from Microsoft? Because it's the place to find research backed insights to guide your org's AI transformation.

Jordan Wilson [00:01:40]:
Tune in now to learn how shifting your mindset can help you grasp the full potential of AI. That's worklab, no spaces, available wherever you get your podcast. Alright. So thanks to our partners at Microsoft. And as a reminder, if you haven't already, make sure you go sign up for our free daily newsletter on our website. We're gonna be recapping today's conversation as well as going over the AI news. Yeah. Technically, prerecorded one here that we're debuting live, but, a lot's happening in the world, with AI, everything happening at CES.

Jordan Wilson [00:02:14]:
We got some open AI rumors swirling, so we'll have that all in today's newsletter. Alright. But enough chit chat. Let's build some more trustworthy AI. I you don't have to hear me ramble on any longer. Have a great guest, today lined up for you all. So please help me welcome to the show. There we go.

Jordan Wilson [00:02:34]:
Barr Moses, the cofounder and CEO of Monte Carlo. Bar, thank you so much for joining the Everyday AI Show.

Barr Moses [00:02:40]:
Thanks for having me, Jordan. It's a pleasure.

Jordan Wilson [00:02:42]:
Alright. Let's let's do this. Well, first, people don't know, what is Monte Carlo?

Barr Moses [00:02:47]:
What is Monte Carlo? Great question. So, Monte Carlo's mission is to help accelerate the adoption of data and AI by reducing what we call data downtime, which is basically periods of time when data is wrong or inaccurate. You can't trust it. I don't know if this has ever happened to you, but you wake up on a Monday morning and you see that one of your data products is wrong. Like you were staring at a report and the number's off, something's wrong. You're like, what the fuck? Like, why is it off? And, oftentimes it's not only really difficult to catch the issue. It's actually also really hard to understand what's the root cause in terms of it. So MoneyCall helps solve all of that.

Barr Moses [00:03:22]:
We're fortunate to work with some of the world's best data teams ranging from companies like Fox, Roche, Credit Karma, and men, many others. It's probably the part of my job that I love the most in getting to work with amazing customers, on some of their hardest problems.

Jordan Wilson [00:03:36]:
Mhmm. So in a nutshell, right, a company comes to you before, what happens after they come to you? Right? Like, if everything goes right, they just better understand their data and and how it works with AI. Like, what's the end result?

Barr Moses [00:03:50]:
Great question. So I would say, you know, there's lots of people today, like data analysts, data scientists, data engineers, machine learning engineers building what we call data products. A data product can be a generative AI application, or it could be a report that your CMO is looking at every day, or it could be a pricing recommendation, algorithm. It it can really be a, you know, sort of variety of of data products, and those data products are often wrong. The biggest issue are sort of based on wrong data, and the biggest issue is that oftentimes data teams are the last to know about that. And so, you know, the very sort of the very first kind of table stakes fundamental thing that we do or that we help organizations is be the 1st to know about data issues. So no longer the days where data teams are sort of caught by surprise by by data issues and sort of hearing about it from someone else. Like, that's the worst thing that can happen that you didn't sort of catch that.

Barr Moses [00:04:45]:
Never happened to me. You know? I'm sort of asking for a friend, if you will. I'm kidding. But, you know, that's sort of, like, fundamentally the very first thing, and that's really the thing that sort of that Monte Carlo kinda set out to really solve when we founded the company, 5 years ago. And that's been sort of the you know, I wanna say, like, the first frontier. I think since then, what's become even more apparent is that that's only, like, first half and, in a sense, maybe even the easier half of the of the work. The actually, the really big challenge, sort of where I think the the sort of AI reliability and industry is heading is not only knowing that the issue that the issue existed, but also answering why and should I care? What should I do with this information? Because oftentimes, data team are just inundated with alerts. Like, this is broken.

Barr Moses [00:05:31]:
This is off. You know, this is this data is late. This data has never arrived. This field is, you know, looks a little bit off. This number is missing. And but in those instances, like, the hard thing is acts actually the answer. I have all these systems working together, but what is actually the root cause? Is it something that went wrong in the data? Is it that the job wasn't completed? Is is that there was a change in the code? Those answers are really, really hard to answer. And so a lot of the things that Monte Carlo does, not only Monte Carlo, just observability more broadly.

Barr Moses [00:06:02]:
So the the field of data observability, if you will, is about answering or helping data teams answer the question of something went wrong, should I care? And if so, why? And how do I resolve that? And that, honestly, is a lot of what observability actually, sort of stored out with. So observability, we didn't make it up in data. We actually sort of, you know, borrow the concept from, borrow the concept from engineering team from our software. So observability in software engineering is very well understood with organizations like Datadog, obviously. You know, who doesn't have Datadog today or something like Datadog? Every single engineering team has something like Datadog and relies on a solution like Datadog to make sure that the software that they're building is, reliable and can be trusted and sort of up and running, if you will. And so data teams, in my opinion, should be doing the same. It's a little bit of a new area observability, if you will. You know, I think Gartner sort of forecast that's over 60% of, of organizations will will have data observability in some sort of fashion form the next 5 years, but but it's a new area.

Barr Moses [00:07:10]:
So, you know, it's been it's been, only recently been defined.

Jordan Wilson [00:07:15]:
I mean, just in the first, like, 3 minutes there, I think you answered, like, my first five questions. I like I I wanna hit rewind just a little bit on the, like, why should we care. Right? And and I think at least in my viewpoint, and maybe this is, for for smaller and and medium sized businesses, but they often don't even really take the time to fully understand, you you know, how their data even works. So they're like, okay. Well, we know we need RAG. Right? We know we need to bring in our own data, you know, to to to work alongside, you know, a back end API. Right? But why does it ultimately matter whether they get their data right, if they are using, you you know, a different, you know, Claude, anthropic, Gemini, etcetera?

Barr Moses [00:08:03]:
It's a good question. And let me just take us back, like, 10, 15 years ago when, honestly, maybe it didn't really matter. Like, it just didn't. You know? We weren't really using data so much. Definitely didn't have any generative AI. And so we could kinda get away with, like, data being wrong most of the time. Worst case, someone just told you you had to go ahead and fix it. Like, no big deal.

Barr Moses [00:08:22]:
You moved on with your life. Right? But I think a lot of things that have changed since there's these been these sort of various, sort of, eras, if you will. And I think the first era was where a lot more people started using data. And so you can no longer get away with, like, looking at the data only once a quarter. Now you have, like, millions of users, you know, pressing, ordering an Uber. And so you can't get the the the time of when your car is coming to be wrong, or you can't get the price to be wrong. Like, for example, if I see that the Uber car is gonna be arriving in 30 minutes, I'm not gonna be waiting for 30 minutes. I'm gonna be signing off and going to a different platform.

Barr Moses [00:08:59]:
Right? And so, yeah, it matters that the time for which for when Uber is gonna arrive, that data needs to be accurate because, otherwise, you're gonna lose your users. So that was sort of the first wave where people just a lot more people were using data, and data products became a lot more important. Then there was a second wave of generative AI, which is now sort of happening. And, actually, interestingly, you know, we recently did, a survey among about 200 or so data leaders. And we basically asked, you know, how many of you are sort of deploying generative AI or building generative AI in production? Can you guess the answer? How many how many are doing that today?

Jordan Wilson [00:09:40]:
I'm guessing a small amount.

Barr Moses [00:09:43]:
Actually interesting. A 100% of them said that.

Jordan Wilson [00:09:45]:
Oh, okay.

Barr Moses [00:09:46]:
Literally yeah. I was surprised to you. Every single person on the survey and these are all, like, you know, data leaders from credible companies. A 100% of data leaders are currently building something that's joint of AI. Now the second question was, how many of them actually trust the data that they're gonna be using?

Jordan Wilson [00:10:06]:
Yeah. And that's that's that's interesting. It is a good point. And I was kinda shocked by that. Right? Because a lot of the studies I read say that even in enterprise companies, you know, I think the late the latest study is only 5% of companies, have generative AI solution top to bottom. Right? Fully implemented. Right? But I guess it's gotta start with the data first and trickle, trickle everywhere else. Is this maybe also right? Speaking of of shifts in generative AI, one thing that I'm personally always dorking out about is kind of this shift from the attention economy to the intention, economy.

Jordan Wilson [00:10:41]:
Right? Like, to be able to better understand, you know, what users on the Internet are going to do before they before they even know they're going to make a decision and that starts with data. Right? Is I I I think we've been hearing for 5, 10, 15 years, oh, data is the new goal. But is it even more in like, increasingly more and more important because of generative AI?

Barr Moses [00:11:03]:
Yeah. 100%. And so going back to that survey, only one out of 3 leaders feels confident in their data that's feeding their generative AI model. So, like, most of us, 2 thirds of us, don't have confidence in the data that we're using. And so, you know, to your question, why does it matter more in in in this new world or in generative AI? I'll explain that given one example and one sort of, more sort of, theoretical example. But the first sort of, you know, sort of real life example, if you will, and this was a couple months ago, someone sort of this went viral somewhere. Someone sort of Googled, you know, what should I do if the cheese is slipping off my pizza? And Google was like, oh, no problem. Just use organic super glue to, like I don't know if you saw that.

Barr Moses [00:11:47]:
It's like put it back on pizza. Like that went viral. And you're like, okay, well, if you're Google, maybe you can get away with that kind of answer. Right? Like, sure. I'm going to continue to use Google tomorrow. Right? But most of us can't afford that. We don't have the luxury of spitting out such falsely or, you know, you know, clearly, misinformed answers. And so it for most enterprises, the reliability of the data that you provide is actually intertwined with your brand and your reputation and the impact on the top line and the revenue that you're generating.

Barr Moses [00:12:19]:
And so that's sort of, you know, one example, to kinda bring that to life. More broadly, though, you know, if you think about what companies are sort of tasked to do now, we're seeing, you know, every single data leader needs to do something with generative AI now. How are they doing that? Because today, every single every single one of us has access to the latest and greatest LLM model. Right? Like, we can all switch between them. We can all use them. Like and so in a sense, we all have access to models built by, you know, thousands of amazing PhDs. Right? Like, we can all do that. So what's my competitive advantage? Right? How am I going to build a better data product than my competition? Or what's my long term moat? And what I believe in, what I'm hearing from my customers is that the moat is actually the data that you have.

Barr Moses [00:13:07]:
Because it's no longer simply, you know, connecting to an API and actually building a generative AI product. The power of building a highly personalized generative AI product is based on the ability to use first party sort of enterprise data. So I can actually build a way better product if I know something about you, Jordan. And if I know about your background and I know your habits, I can actually build something that's personalized for you. And the data that I have is something that, arguably, no one else has. And so I think in in you know, for for leaders thinking about what are they going to build or how are they going to, to use generative AI, the data that you have is the moat. That's actually how you gain competitive advantage and build data products. And so if you believe that's true, then the quality and the reliability of the data that you're using is of utmost important.

Barr Moses [00:13:56]:
Because if the data that you have is inaccurate, then what's the point of the moat that you have?

Jordan Wilson [00:14:01]:
And and I think this makes a lot more sense and is resonating for those that work at larger enterprises. Right? And they already, you know, have a data warehouse or, you know, data lake. They're using Amazon s 3. I don't know. Right? But for maybe those medium sized companies that don't have their data game strong. Right? But they have it in a lot of different places. Right? Maybe they have some, you know, floating around in in different places in Google or, you know, in their CRM, etcetera. How do these, you know, smaller, and medium sized organizations, how can they take advantage of that? Because what you said there is true.

Jordan Wilson [00:14:36]:
Data is the moat, But how do these smaller and medium sized organizations start to actually pull all of this data together so they can, you know, leverage generative AI with it?

Barr Moses [00:14:48]:
Yeah. I mean, I'd start by saying no data is better than bad data. So if you have shit data, I'm actually not convinced that you should be using it. I actually think it might be better to to to make sure that you have data that is reliable and and trusted. You know, I think just to give you an example on Monte Carlo. Monte Carlo as a company, we're about 200 or so employees. We build generative AI products, and it's of utmost important that the data that we have is accurate. And so I think even if you are a small organization, it you know, the the bar is not lower.

Barr Moses [00:15:20]:
In fact, I think it's higher. In fact, I find that enterprises, lark you know, large enterprises really struggle with getting their data together, really struggle with having a source of truth. Like, if I have multiple copies of the data, for example I mean, even answering questions like, you know, how many customers do we have? Or large organizations needing try to fig figuring out sales compensation. It's really freaking complicated to do that because the answer that I get from my finance from my finance team is different from what my sales team is saying, is different from what my marketing team is saying. And so every different team is looking at different sets of data, and so getting an answer is really, really difficult. So I actually think medium size and smaller organizations have an advantage. You have, you know, you actually in fact, you know, I think it's, I actually think smaller teams move faster. Like, there's probably some proof of that.

Barr Moses [00:16:12]:
And so as a smaller team, you know, you're probably small but mighty. So make use of the data that you have, and you have the advantage of being able to move faster and actually innovate faster because larger organizations now are are way, way slower and obviously way more risk averse. So I think smaller organizations have the benefit of being able to try a lot of things, experiment, move quickly, and and double down on, you know, some of the, the experiments that are working. And by the way, that's sort of the large majority of what we see companies do, both small and large, basically have, like, this mandate to go experiment in the organization and, like, have lots of teams, you know, try out different, different things and sort of diff build different applications, and and, you know, companies understanding that they will come up with a centralized strategy only later.

Jordan Wilson [00:17:04]:
Alright. So I do wanna talk about some of these use cases, but we're gonna take a quick 22nd break and have to shout out, one more time our partners, at Microsoft. So why should you listen to the WorkLab podcast from Microsoft? Because it tackles your burning questions about AI at work, like how can I guide my org's AI transformation? How can AI help maximize value and create new products and business models? What mindset shift do we have to make if we want to tap into its full potential? Find the answers on worklab. That's worklab. No spaces available wherever you get your podcast. Alright. So, Bart, I I I I do wanna jump into it a little bit here, because we've been talking about, like, some of the issues, with with getting good data, having reliable data that you can trust. So what happens when it does get together? Maybe could you walk us through a a a use case or 2 just for those that maybe are just getting their data feet wet, so to speak, so they can see, hey.

Jordan Wilson [00:18:02]:
How do these when when when good data and good GenAI comes together, here's good use cases.

Barr Moses [00:18:09]:
Yeah. Absolutely. And it's been a really fun sort of hearing kind of, you know, variety of use cases and sort of innovation. I'm really excited. And honestly, like, the hyperbound, this is so big, but I think even if it materializes only 10% of the way, that's enough to be such a disruption for us, and for future generations. So I'm I'm really excited. I'll I'll give a specific use case, actually one that, you know, we use at at Monte Carlo. So, you know, one of the challenges that we have is, oftentimes when this is sort of very meta.

Barr Moses [00:18:39]:
But, oftentimes when when we work with data teams, they actually don't know the state of their data, and they certainly don't know why their data might go wrong and what what might go wrong there. And so if you need to set up sort of coverage with data quality monitors, you don't always know how to get started, especially if you are a less technical user, that might be a lot harder. And so what we do, is actually sort of build, data quality monitor recommendations where we actually sort of profile the data that we that, you know, specific customer. We use Anthropic's cloud, 3.5 Sonnet. And one of sort of, kind of advantages of working with LLMs is that they have a really strong semantic understanding. And so we can actually use with a combination of profiling the data and the metadata, a bunch of other kind of contextual things that we bring together. We can use that to actually help define what monitors you should be setting up. So I'll give sort of a really, hopefully, an easily understandable example.

Barr Moses [00:19:42]:
You know, we work with, sports organizations, for example. And so if you take, like, a, you know, like, a baseball organization, for example, and you think about, like, pitch types, you know, and actually, like, baseball and in sports in general collect a ton of data about, you know, different athletes, different players, and a lot of sort of data about the games itself from a ton of statistics and analysis, for anyone who's seen Moneyball and others. One of, you know, one of the things and one of the types of data you might collect is, you know, the type of pit the type of pitch and, also the speed of the pitch. And so, for example, using analysis, you can actually learn that you can actually determine that if you have a fastball, that should always be over 80 miles per hour. And if it's under 80 miles per hour, then maybe there's a problem. It's not really a fastball. Right? And so that's a sort of a recommendation that we can make using generative AI, using LLMs to say, hey. You should set up this data quality monitor.

Barr Moses [00:20:40]:
And we can do a lot more, you know, that kinda help users actually, make sense of their data in order to drive what sort of data quality monitors they need.

Jordan Wilson [00:20:56]:
Why should you listen to the WorkLab Podcast from Microsoft? Because it tackles your burning questions about AI at work. How can I guide my org's AI transformation? How can AI help maximize value and create new products in business models? What mindset shift do we have to make if we want to tap into its full potential? Find the answers on WorkLab. That's worklab. No spaces available wherever you get your podcasts. And and that's that that's a good example, and I think it really illustrates, a a point. So, you know, because we can all, relate to, you know, oh, classifying a pitch. Right? And, you know, you never really know what it is until you, you know, you see it or, you know, you're watching on TV. But maybe could you walk us through 1 maybe one more, example of, you know, how good data in in you know, knowing that you can rely on it and how that can really make a difference.

Barr Moses [00:21:58]:
Yeah. For sure. So another example of, you know, I think a journey of AI use case that's really cool is something that, Credit Karma from Intuit, does. So, you know, for folks who don't know, you know, Credit Karma is sort of financial assistant, that's based on AI, and so can make recommendations for you on how to best manage your finances. And so like I mentioned before, you know, any organization has access to latest and greatest, you know, OpenAI, API, or others. What Credit Karma has that no other organization has is information specifically about their users. And, you know, they serve 100 of millions of users, and they can tell, you know, you have this specific credit score, and you've had this Honda car for the last 10 years, and you're gonna be selling it at this time, and you have this kind of, history. All that information can be used to help make specific recommendations for you about your specific financial situation.

Barr Moses [00:22:56]:
Now the downside is, you know, we don't we wanna make sure that we're not surfacing to you the wrong credit score. So you, Jordan, should be able to access only your credit score and not my credit score, for example. And, also, the financial recommendations that are being made to you should be based on your data and your data alone. And so I think the power you know, Credit Karma actually builds, rag, pipelines. And so they they, they use LLM and actually, like, enrich them with data that they have about their users in order to build these highly personalized, assistance, if you will. And so, you know, I think being able to actually build such a personalized, product that's also based on reliable accurate data results in a really, really strong outcome for customers. That's sort of one kind of example from the financial world. There's a lot more examples where companies make good use of LMM and GenAI actually for efficiency internally, which I've seen.

Barr Moses [00:23:56]:
So, you know, the Credit Karma Intune is more an example of an external data product that gives you a really strong ability to to make an impact on your customers externally. If you think about seeing value internally as well, lots of organizations the most basic example is where organizations see, you know, an increase in in engineering productivity. So where you have a sort of coding assistance, that's, like, the most basic, I think, that most organizations are seeing today. And, you know, I think that often helps for, you know, more sort of junior or inexperienced engineering, but also senior engineers as well. So if you have a largely sort of junior or new organization, you'll you'll realize even more benefit more benefits than that. But I think, you know, some numbers, like, you can increase significantly the number of, the ratio of sort of number of coders and code that's being reviewed with with with LLMs. Another example is, you know, in the in the, pharmaceutical or sort of medical space, but also in in, insurance. There's a lot of compliance reports that are being shared, and those oftentimes can take 6 to 12 to 18 months to generate.

Barr Moses [00:25:09]:
And those include, you know, a lot of internal data, but also sort of status and protocols and, you know, a lot of, sort of, like, wrote and sort of manual report writing. Generative AI can significantly reduce the time based on that. So if you sort of, use kind of the data that you have off the shelf and and actually train it on past reports, you can actually generate, really good examples or at least draft to start out with. So those are kind of examples of how folks are really gaining internal efficiencies. There's also some clever ways specifically around structured and unstructured data. So oftentimes, folks find that, it's it's I would say, in general, kind of the whole unstructured data stack is very new and is just emerging. Like, I think this is very early days for unstructured data. Mhmm.

Barr Moses [00:25:59]:
One of the things that are hard is how do you monitor and observe unstructured data to make sure that unstructured data is reliable. Like, that's really you know, I wanna say very, very early days of that. Something that Monte Carlo, obviously, we we think a lot about, but also our customers obviously think a lot about. One one good example of how you might use LLMs to better, observe unstructured data is, we work with an insurance company that has customer support chats. And if you think about customer support customer, customer support conversation, that's largely unstructured data. And you can use LMS to actually structure that specific support chat and give it a score based on, you know, what is the reading of, like, the tone and the conversation and the resolution and understanding whether, you know, that support conversation went well or not and basically assign it a score between 0 to 10 on how it went. And one of the sort of use cases is that is then you can observe that structured data. Like, let's say some let's say the LLMs LM gave it a score of a 12.

Barr Moses [00:27:07]:
Like, what does a score of 12 mean on, you know, a score between 0 to 10? Right? And so in those instances, you can actually make sure that, that data is reliable. So there's a lot of kind of clever ways in which folks are using LLMs to, structure the unstructured, if you will.

Jordan Wilson [00:27:23]:
Mhmm. Yeah. I yeah. I I love that. And that's something that, you know, I'm a I'm a huge proponent of, especially for smaller and medium sized businesses. It's like, yeah. Use LLMs to turn that unstructured data into structured data that you can actually use. So I would be remiss if I didn't ask you this though because, you know, it's been a growing growing trend at least of 2024 is, you know, using synthetic data.

Jordan Wilson [00:27:45]:
What's what's your thoughts on that? Right? Yeah. Like like, I love what you said. Like, good like, having no data is better than having bad data. Is using synthetic data better than using bad data? Like, where do you stand on that, and is this gonna be a big play in the future?

Barr Moses [00:28:02]:
Yeah. I mean, I think there's a I think I don't remember a little bit. I think there's a former scientist of OpenAI who said, like, we are at peak data right now. Like, there's so like, we, you know, sort of maxed out on on the data that there is, and, you know, we now need to turn to synthetic data in order to kind of make advancements. And so I think it's it's definitely an interesting time for synthetic data, and I think there's gonna be a rise for that in terms of, you know, how we're gonna train LLMs and, you know, how how we're gonna, sort of reach even better performance, if you will. But I do think that there's no, obviously, no replacement to, you know, real world data, that, that enterprises need to use. And I see most of the attention and the time spent there. So it's actually interesting.

Barr Moses [00:28:47]:
There's this return to some of the maybe more unsexy things, like data governance is sort of, like, rearing its head now. I haven't heard that, you know, that word in a couple of years. And now there's, like, a, you know, a return of data governance. And, so I actually think a lot of, like, you know, what what was old is new now, if you will, and perhaps synthetic data is is in that camp as well.

Jordan Wilson [00:29:10]:
Mhmm. Alright. So, we've covered a lot in today's conversation, Bar. We started with, you know, the reliability of data and how observability worked. We gave some examples, you know, and talked about the future of data as well. But as we wrap up, what's maybe the one most important thing that you think our listeners need to know, especially those making medium and long term decisions on how their companies play in the AI space? What's the one most important thing they need to know about reliability of their data?

Barr Moses [00:29:41]:
I would say, you know, your generative AI product is only as good as your data. And so excuse my language. But if your data is shit, then your generative AI is gonna be shit. And so getting that in order is the first thing. And, actually, that's a really tall order. It's actually really freaking hard to do that, but I think there's no other way. I do think that we are seeing more and more organizations actually seeing, ROI on their generative AI products. And so it's high time that, you know, any organization starts to if you haven't already invested, then you're already too late.

Jordan Wilson [00:30:15]:
Love to hear it. Yeah. That's a great, a a great warning call to all of you people that are somehow still sitting on the fence in 2025. I don't get it, but there's a lot of you out there. So, Bar, thank you so much, for joining the Everyday AI Show and taking time out of your day to help us all better understand data. We really appreciate it.

Barr Moses [00:30:34]:
It's been fun. Thanks, Jordan. Good luck to us all.

Jordan Wilson [00:30:36]:
Alright. And, hey, y'all. That was a lot to take in. Yeah. Just so much a data avalanche of great information. If you miss something, maybe you're on the elliptical, and and looked away, don't worry. We're gonna be recapping it all on our website, your everydayai.com. Sign up for the free daily newsletter where you will find, a lot more insights and complimentary supplementary info to go along with today's conversation as well as everything you need to know to be the smartest person in AI at your company.

Jordan Wilson [00:31:04]:
Thank you for joining us. Please join us tomorrow and every day for more everyday AI. Thanks, y'all.

Gain Extra Insights With Our Newsletter

Sign up for our newsletter to get more in-depth content on AI