Resources
Join the discussion: Ask Jordan questions on ChatGPT
Related Episodes: Ep 253: Custom GPTs in ChatGPT – A Beginner’s Guide
Ep 318: GPT-4o Mini: What you need to know and what no one’s talking about
Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Connect with Jordan Wilson: LinkedIn Profile
Try Our Free AI Prompting Course: Register for our free Prime, Prompt, and Polish AI Course!
Harnessing the Power of Tokens in Large Language Models
In today's technological world where Artificial Intelligence (AI) has become integrated into daily operations - ranging from customer service chatbots, content generation to search queries - understanding how language models work has become increasingly paramount.
Language models, such as ChatGPT, adopt tokens to breakdown complex languages into more manageable units. Tokens are essentially subwords or subsections assigned numerical values which may vary based on factors like capitalization, plurality, and adjacent words. As a result, it is the token domain that the models utilize in understanding, processing words and even in giving responses.
Tokenization: A Critical Factor in Processing Information
Tokenization helps AI models not only understand and handle language better but also retain context and process information much more efficiently. Essentially, tokenization processes involve breaking down extensive languages and queries into manageable tokens.
For instance, in the context of chatbots, the quality of chat interaction largely depends on how many tokens are processed within a given timeframe. This fundamental understanding not only helps in improving AI-powered customer experience but also in enhancing decision-making processes for businesses.
Token Limits and Memory Recall Techniques
When handling large language models, it's crucial to be aware of the token counts and memory limitations. Models tend to operate within a fixed range of tokens. For instance, the ChatGPT model remembers the most recent 32,000 tokens, and anything beyond that is forgotten.
To work around this limitation, innovative techniques such as memory recall can be utilized where important information is summarized and positioned as a "cliff notes version". This helps in continuing the same chat after maxing out the tokens.
Understand Differences in AI Models
There's enormous variety in AI models, and it's essential to note that different AI models have differing context windows and memory capabilities. Therefore, the proper utilization of large language models requires understanding the context window, which determines how much information a model retains before forgetting.
For example, a model like Google Gemini has an impressive 2,000,000 token context window for back-end development, which means an entire book can be pasted for interaction, compared to Chat GPT’s smaller context window.
Final Thoughts
Understanding how AI models work and adhering to the rules of these models can significantly improve interactions with them. Moreover, within language models, recognizing the importance of tokenization, its process and the concept of context windows are critical steps towards enhancing the effective use of AI in businesses.
Topics Covered in This Episode
1. Tokenization in ChatGPT
2. Comparison of Different AI Models
3. Importance of Tokenization and Memory in AI Models
4. Limitations of ChatGPT
5. Explanation of Tokenization Process
Podcast Transcript
Jordan Wilson [00:00:19]:
This small little detail about large language models is something that most people don't understand, and it's probably impacting your work. So we're gonna be talking today about a complete guide to tokens inside of chat gbt, well, and other models in helping you all understand what tokenization is, why it's important, and this might explain why maybe you're working with a large language model and things are going great, and then all of a sudden, they go off the rails. It's not because the model stinks, it's because you have to understand how they work, and that's what we're all about here at everyday AI. What's going on y'all? My name is Jordan, and I'm the host of Everyday AI, and this thing is for you. It is your daily livestream podcast and free daily news free daily newsletter, helping us all learn and leverage generative AI to grow our companies and our careers. So if that sounds like you, thank you so much for joining. If you're on the podcast, make sure to check out your show notes. You can always come back and, watch the video.
Jordan Wilson [00:01:25]:
This one might be one
Jordan Wilson [00:01:26]:
of those where visuals are going to help a little bit. I'm going to try to explain everything, that we have on screen here for our livestream audience, so make sure you check out your show notes. And if you are joining us live, thank you so much. But all of you, if you haven't already, also why haven't you gone to your everydayai.com yet? Sign up for our free daily newsletter as well as it is like a the world's largest, unbiased free library of generative AI education. Thousands of hours worth of content on there for free from myself and, leading experts in AI. Alright. So let's get into it, and make sure to check out the thanks a million giveaway while you're there. So let's go over what's going on in the world of AI news before we get into token talk.
Jordan Wilson [00:02:12]:
Alright. AI news, a lot going on today. A lot of the big names, new models, new features. Let's get into it. So first, Google Gemini has launched Gemini Live, its new AI assistant. So Google has launched Gemini Live, a new live AI chatbot feature promising more natural and emotionally expressive conversations on smartphones. So this was announced earlier at a Google event, but this new hands free AI assistant allows users to talk in real time and interrupt an AI assistant and ask questions with the chatbot adopting and, adapting in real time. So although it currently lacks the multimodal input feature, reports suggest they will it will soon be added along with broader language support and eventually iOS availability.
Jordan Wilson [00:03:01]:
So, yes, this is similar to this, more neural voice feature we've been waiting for from chat GPT and OpenAI as they go through some safety precautions, but, Google has beat them to the punch with Gemini live. Alright. Speaking of beating them to the punch, Google has also beat Apple to the punch. So Google has unveiled its new AI powered Pixel smartphone lineup. So Google has made a significant move to challenge Apple and Samsung by launching a new series of AI integrated devices with edge AI or on device large language models. So the array includes the Pixel 9 series of foldable phones, smartwatches, and earbuds. So, Google has introduced this new Pixel 9 series featuring their new advanced tensor g 4 chip, which allows for this on device edge AI. So the highlight of the launch is the Pixel 9 Pro Fold, which has one of the largest flexible screens in the market.
Jordan Wilson [00:03:56]:
So the new devices also include 2 smartwatches, the Pixel Watch 3 and the Pixel Buds Pro 3, both equipped with cutting edge AI technology. But the key feature, like we talked about, is Gemini Live, which will be available on that phone. Alright. Twitter. Yeah. Twitter or XAI, whatever you wanna call it. Yeah. Maybe I should just start calling it X.
Jordan Wilson [00:04:17]:
It just seems weird, but Elon Musk's XAI has launched Grok 2 and Grok 2 Mini. So Elon Musk AI company, XAI, has launched Grok 2 and Grok 2 Mini in beta, available exclusively to premium and premium plus users on the X social network. So Grok 2 features advanced capabilities in chat coding and reasoning, while Grok 2 mini is smaller but capable and great for developers. Early feedback, indicates Grok 2 can generate images without restrictions, including those of political figures, raising some concerns about potential misuse. So, the company aims to integrate grok 2 into x's features such as improved search post analytics and AI powered replies. Great. That can't go wrong. So also with the US presidential election approaching, x AI is definitely gonna face a lot of pressure to implement restrictions, especially on image generation to prevent misinformation.
Jordan Wilson [00:05:13]:
Alright. Last but not least, OpenAI's GPT 4 has reclaimed the top spot in the chatbot arena after a somewhat quiet, model update. So earlier this week, OpenAI quietly updated its GPT 4 o model and didn't even really give us a name until 24 hours after the update. So the latest model, chat, it's technically chat chat g p t four o, either the o eight version, which if you're a developer, you might see that, or, the GPT 4 o dash latest. Yeah. Great great naming there. Anyways, this new updated version of GPT 4 o has reclaimed the number one position in the chatbot arena, surpassing Google's Gemini 1.5 pro, with a score of 1314. So, Gemini, just so you know, Gemini 1.5 pro is available for developers only, but this newest version of, chat gbt gbt 4 o latest is available.
Jordan Wilson [00:06:12]:
That is the default mode now when you go to, chat gbt. So, the achievement comes after a week of being tested on the chatbot arena, with over 11,000 community votes after it was named the anonymous dat chatbot label. Alright. So, the new version of Chat gpt shows significant improvements in technical areas, particularly in coding with a 30 point increase over the previous models. Also, the model excels in instruction following and handling hard prompts showcasing its versatility. Wow. That's a lot going on from some 3 of the biggest, companies out there. Let me know.
Jordan Wilson [00:06:49]:
Livestream audience, you know, Jason and Denny and Kathleen, Chris, Woosie, Rolando, everyone else. Do you wanna hear more on any of these stories? Should should we be doing some, some reviews or tutorials today? Let me know. Alright. Let's get into it, though. Let's talk about tokens. Alright. I'm actually excited, for today's show. I'm gonna try to keep this one pretty short and factual, but, you know, I also wanna know from both our livestream audience and, you know, those on the podcast.
Jordan Wilson [00:07:18]:
I always put my, you know, LinkedIn, information in there. You can connect with me, email. But let me know. Do you know anything about tokens? Do you have questions? Do you know what they are, why they're important? At least for our livestream audience, let me know. I'm gonna make sure to answer your questions on tokens, but we're gonna do a complete look. And actually, y'all, if I'm being honest, not understanding how tokens work is one of the biggest reasons probably why, your chats go awry. So if you've ever had, an instant where you are working inside of a of a chatbot like chat gbt and things start out great. Right? And you're like, oh, this is great.
Jordan Wilson [00:08:01]:
You know, I'm getting fantastic, you you know, fantastic results. And then all of a sudden, like, all of a sudden it starts to just get dumb. And then you think, oh, it's just the model. Well, not really. What's happened is it's going past its context window, and that context window is measured in tokens. Okay? And all of the major kind of chatbots have, you know, have a different process of tokenization, but they all have kind of a a memory limit or a context window of a different number of tokens. Alright. So, we're gonna be going over the basics, and we're also we're gonna try to do a couple things live here.
Jordan Wilson [00:08:39]:
Alright? So, if you wanna see how all this works, you you know, this is one of those, if you are on the podcast, it might be it might be worth, you know, watching this one. But, let's let's get into it, and let's go over what you need to know. So here's here's what we're gonna go over. We're gonna talk about what a token actually is, why large language models even use tokens, the tokenization process, and then we're gonna talk about the context window, hopefully with some live demonstrations here. Alright. So let's dive into it. So what the heck is a token? Alright. A token, think of it as this.
Jordan Wilson [00:09:24]:
It is a a subset of a word. Alright? So, essentially, what large language models do is they break down words into smaller sub words or sections or tokens. So, we're gonna be going over examples, but let's say, I don't know. Let's pick a trending word if you're an AI dork like me. Strawberry. Right? Technically, large language models, if you, you know, ask a question about a strawberry, large language models don't see words. They technically don't even understand words, which I know might sound interesting. But, essentially, there is a process that goes on, right, because a lot of people might just not even understand.
Jordan Wilson [00:10:08]:
Oh my gosh. How does how do all these large language models how can they understand anything and do almost anything? Well, it breaks down words into these smaller subwords or subsections called tokens. Right? So as an example, the word strawberry, I know, it breaks it into 3 smaller tokens, and it assigns tokens essentially a numerical value. And then depending on the context that you might give a word, it's going to change that token value. Right? Even capitalizing something, turning something plural, having a space before or after a word, it is going to change its token value, and that is actually y'all. That is actually how large language models understand us. They understand our words by assigning token values to everything. So that is part of, an algorithm of sorts that helps all of these large language models be helpful assistance to us because in the end, they actually don't understand words.
Jordan Wilson [00:11:08]:
Right? And, technically, when they are spitting backwards, they're technically just spitting back tokens that get converted into words. Okay. So it's a common misconception. Right? And also, no. Models cannot count, as an example, the number of r's in the word strawberry or even the number of words it cannot consistently. Right? So if you say, you know, hey. Respond to this in 10 words or less, it might give you 13 words. It might give you 8 words.
Jordan Wilson [00:11:38]:
That's because it doesn't understand the concept of words. It understands the concept of tokens. That is how large language models think. That is how they process. That is how they respond back to you. Okay? One thing we're huge at at everyday a everyday AI is education. Right? And this is something, like I said, even the experts, quote, unquote, self anointed, self appointed experts are getting wrong. Right? Because they say, oh, large language models are dumb.
Jordan Wilson [00:12:05]:
They can't even count the number of r's in strawberry, or they can't even, you know, respond back when I say, you know, describe a strawberry in 10 words. It never gives me 10 words. Well, that's not how large language models work. Right? The key to getting the most out of generative AI is you first have to unearth and, you know, look at what's underneath. You have to understand things at a very basic level. Okay? So that is what a token is, and don't worry, we're gonna be going over some examples. So, you know, one example, dog might be one token, or the word running might be 2 tokens, run and then ning. Right? So it might split it up in different places.
Jordan Wilson [00:12:46]:
Alright. Let's keep going, and I see some of your questions y'all. I'm gonna get to them, at the end, but keep your questions coming in. Alright? So why do large language models even use tokens? Well, it helps them for context and, analysis. Right? So it helps them better understand what you are actually talking about. Alright? Also, it helps with language handling and process efficiency. Okay? So let's talk about that a little bit more. So with tokens, it helps large language models break our text, our sentences into smaller units, those words or sub words, and then it allows the model to better understand and retain context.
Jordan Wilson [00:13:32]:
Okay? Also, it allows this is how large language models can sometimes speak 100 of languages. Right? You're like, how? And how can it flawlessly, you know, switch between languages? And we've seen even with these new voice assistants. Right? We've seen, you know, demos of them where someone's saying, hey. You know, where it acts as a real time translator. That's how that's how it can handle different languages is through this process of tokenization. Also, this actually helps processing efficiency because using tokens technically reduces the computational load. Right? Think. It actually breaks down complex languages, complex queries into much more manageable tokens, into essentially an algorithm of stores of sorts.
Jordan Wilson [00:14:24]:
Right? When you break, complex languages and complex problems into tokens, and you'll see what that means. It actually helps the model be more efficient and process faster each and every time. Alright. So now let's talk about the tokenization process. Alright? And and and, hey, I agree. Right? Jason, y'all, like, if you're tuning in live, first of all, now you are the smartest person at your company unless you work at OpenAI and you're working with the people who are, you know, actually building these models. But otherwise, yeah, maybe you should get some sort of continuing education credit for this. But, you know, I guess now you're the smartest person in your company when it comes to tokenization and how large language models work.
Jordan Wilson [00:15:06]:
Alright. So let's talk about the tokenization process. Alright. And we're probably gonna jump into this later, but a lot of people don't know that OpenAI has a tokenizer. Right? You can go and play and learn how this works. So, I do have to say that this is not, the token the tokenizer for g p t four o is not yet available. So the tokenizer that we're gonna be going over both in the examples on my screen, and we'll probably do a little bit live, This is for the, GPT 3.5 and the GPT 4 version. The kind of like quote unquote old school.
Jordan Wilson [00:15:42]:
Right? So let's let's take a look. Let's take a look at this now. So what I'm showing here is different use cases of the words strawberries. Right? So, you know, tending to a strawberry patch, a heart shaped strawberry candy, blowing strawberries on a belly, strawberry birthmark on the face, scraped straw scraped strawberry knee wound. Right? Like, even the word strawberry technically has many different meanings. So you can go into the OpenAI tokenizer. Okay? And this is, technically gonna help explain 2 different things here. This is going to help explain how words and to show how words turn something into tokens, but also the context window.
Jordan Wilson [00:16:35]:
Okay. So now when we look at those different words, because the word strawberry is going to mean different things if it has different context. Right? The word strawberry by itself and then versus a strawberry patch or strawberry candy, that word strawberry is technically going to be broken into even different sub tokens depending on the context in which it is used. Right? That is how, large language models can understand nuances in language when it technically doesn't know language. It aside it looks at the context of how you are using words. Right? Is it capitalized? Is it is it singular? Is it plural? Are there words before or after? Do those words before or after change the meaning of said word. Right? So then after you put in, you know, and anyone can do this and, you know, we will have this link for the tokenizer in the newsletter today so you can go and play with it and get super smart. Right? So now I'm gonna show you the token values that this these are assigned.
Jordan Wilson [00:17:40]:
Alright. So, actually, this is just the, kind of the visual. Right? But you can even see, right, at the top here, I have the word strawberry. Even the word strawberry with a space versus not a space or a space before or it being capitalized or it being, you you know, singular versus plural, it is going to tokenize it a little differently. Okay? So we're gonna be going over, we're gonna be going over some actual examples. And no, we're not talking about tokens as a new currency. We're talking about how large language models work. Alright.
Jordan Wilson [00:18:22]:
So that is the tokenization
Jordan Wilson [00:18:24]:
process. And y'all, let's put this quickly into perspective here. We're gonna
Jordan Wilson [00:18:31]:
be talking about a context window. Okay? But, essentially, you see even these couple of sentences here, we're getting a total number of tokens as well. So tokens actually play a dual role. 1st and foremost, the tokens are important because that's how large language models can actually understand words. 2nd, there is something called a context window, or you can think of that as memory, and that is how much information chat gpt can retain before it starts to forget things. Okay. So this is the context window. Right?
Jordan Wilson [00:19:09]:
So it's different models. Also, this is important
Jordan Wilson [00:19:15]:
to know, but all models so so chat GPT, anthropic Claude, Google Gemini, they all have different context windows. Right? Technically, using chat g p t, inside of the, you know, quote unquote chat interface versus the API, it actually has a fairly small context window compared to competitors. Right? So, when you use Claude, Claude has the best and the largest context window or memory when using it inside of the chat product. Okay? If you're using it inside of, the development product or not the front end, it's actually Google Gemini with up to a 2,000,000 token context window, which is huge. Right? So that's as an example, you can paste an entire book into chat g or sorry, into, Google Gemini if you are using it on the back end in kind of the developer mode, and you can talk and ask questions of that book. Where if you're inside chat g p t, it has a much smaller context window. And we're gonna go ahead and probably describe this live. Alright.
Jordan Wilson [00:20:20]:
So bear with me y'all. We're gonna get a new chat going. Alright? And I want everyone to see and to understand this tokenization process and why it's so important. Alright. So, hey, livestream audience, let me know if you can, see my screen here. But for our podcast audience, this is what we're gonna do. Okay? We're giving chat gpt some basic information about myself. Alright? Also, I have inside of my chat GPT account, a tokenizer.
Jordan Wilson [00:21:00]:
Okay? So the one this is not by default, although I don't know why large language models do not have a token counter by default. They should. Okay? The one I'm using, I'm sharing it on my screen here. So I'm using Chrome, but I use what's called the chat GPT token counter. This is a Chrome extension, from amperly.com. Alright. So, I'll have that in the newsletter as well. Alright.
Jordan Wilson [00:21:32]:
So when I'm sharing my screen here and we're gonna be going over this process, bear with me. I promise you this is gonna help you understand tokens, and it's going to actually improve your outputs a ton. And even if you're listening on the podcast, I'm going to explain this to you. Alright. So as we go along, I'm going to put some information in. Okay. And you'll see the token count right here. So I have some information about myself.
Jordan Wilson [00:21:55]:
I'm saying my name is Jordan. My favorite color is Carolina blue. My favorite food is deep dish pizza. Oh, I just had some deep dish pizza last night. Shout out to Todorisis. And then I'm saying, I think the bears kind of stink, but they might be okay this year. Okay? So all of my words then when I hit enter are gonna be converted into tokens. I can't see that on the back end.
Jordan Wilson [00:22:16]:
Right? But we're gonna see also when chat gpt responds to me, we're gonna see my token count go up. So let me take it away. So right now, technically, this, this token counter is counting anything on the screen. Alright. So, you know, essentially, it's about 50 tokens, my query. So when I hit enter, ChatGPT is going to respond. It's probably gonna say, oh, great. Thanks so much for this information.
Jordan Wilson [00:22:39]:
And then you're gonna see my token count go up. Right? So there we go. So it's, you know, giving me some standard blah blah blah. Hey, Jordan. Nice to meet you. Alright. Let me say something else. Right now, ChatGPT has, different features that I can actually remember things by default in memory.
Jordan Wilson [00:22:55]:
I have that turned off, just so everyone knows. Because you might say, oh, when you tell chat gpt those things, it's gonna remember. So, of course, no. We're testing the context window here. Memory is off. Okay. So now what I'm going to do, I just have a bunch of random, transcripts here. Okay.
Jordan Wilson [00:23:16]:
So I can't, you you know, do, 30,000, all at once, but I'm gonna say, please I'm gonna say, please summarize this text. Alright. Give me a second here. I wanna prove something to everyone. Alright. So now you'll see I put in a ton of text. So we just jumped up already, to about 17,000 tokens. Alright? I'm gonna go grab some more, and I'm proving a point here y'all because a lot of people don't read between the lines.
Jordan Wilson [00:23:46]:
Right? And they take whatever a company says as truth, which you shouldn't do.
Jordan Wilson [00:23:52]:
Alright. Because OpenAI actually, yeah, they're a little
Jordan Wilson [00:23:57]:
I mean, it's technically on their website, but when they talk about GPT 4 0, you probably hear a context window of a 128,000 tokens. That's what most people think that the memory of chat GPT GPT 40 is. Alright. It's not. And I'm going to prove that to you. Alright. So now I'm going to put a bunch of other, information in here, and I'm keeping an eye on the tokens. Alright? So let me go ahead and grab a little more.
Jordan Wilson [00:24:26]:
Y'all, this is this is what we do. Right? We are investigative practitioners here at Everyday AI. We try to break things and test everything. Alright. So here we go. This should be a little bit better. Alright. There's a point here, y'all.
Jordan Wilson [00:24:42]:
Alright. So now you'll see after chat gbt responds, I'm probably gonna be at about 30,000 tokens. Okay? There's a point here, y'all. I am proving and showing you through an example why tokenization in the memory matters. Right? Because I talked earlier about, oh, maybe you're having a conversation with ChetGPT and you share some things and things start out great. And then after a while, it starts to lose its memory. Right? So now here's what I'm doing. We're at about 29,000 tokens, just under.
Jordan Wilson [00:25:15]:
Okay? So now I'm asking, what's my name? What's my favorite color? What's my favorite food? What do I think of the bears? Right? So why am I asking that? Right? I'm doing it to show you that even though there are there's about 30,000 of tokens of information between my question and where the answer is. Right? It should, in theory, get all of these things correct. The tokenizer is not a 100% correct. These are estimates. That's why it didn't go up to, like, 31,900. Okay? I wanted to show you, okay, we're at just about 30,000. Let me ask this question. And there we go.
Jordan Wilson [00:26:00]:
Right? I I asked the question, what's my name? What's my favorite color? What's my favorite food? What do I think of the bears? It says your name is Jordan. Your favorite color is Carolina blue. Your favorite food is deep dish pizza. You think the bears kinda stink, but they
Jordan Wilson [00:26:12]:
might be okay this year. Alright?
Jordan Wilson [00:26:15]:
Yeah. Monica Monica says love it. Investigative practitioners. That's what we are. You know what? And if you have taken our PPP course, our free prime prompt polish course, it's been taken by now, I think, 7,000 professionals. It is free. It is live. If you've taken our course, you probably know this, and you probably know kind of a a workaround for this called a memory recall.
Jordan Wilson [00:26:38]:
Not gonna go into this in this, in this session, but, you know, if you are tuning in live or on the podcast, you know, just put PPP, in there, and I'll send you a link where you can sign up for free. And I'm wondering, has has any of our, livestream audience taken our PPP course? Alright. So, anyways, there you just saw an example. Alright. We tested the context window. This is important because I say easily, this is a top three mistake that literally hundreds of millions of people are using large language model. And this is a top three mistake that people are making and they have no clue. And this is one of the reasons why people say, oh, we can't we can't use a large language model at our company.
Jordan Wilson [00:27:20]:
Look. It it it starts good, and then it goes off the rails. These models are broken.
Jordan Wilson [00:27:24]:
No. They're not. You have
Jordan Wilson [00:27:26]:
to understand how they work and play by the rules. I'm teaching you how. Alright. So now what we're doing is we are starting a new chat. I'm doing that so we don't have this same context window. So, essentially, we are starting over fresh. Okay? And I'm starting with the same thing. Right? The information about myself.
Jordan Wilson [00:27:45]:
My name is Jordan. My favorite color is Carolina blue. My favorite food is deep dish pizza. I think the Bears kinda stink, but they might be okay this way, this year. Any football fans in the house? Right? Any Bears fans? I've I've always been a realist Bear fan, to tell you the truth. All my friends are like, every year, they're like, oh, the Bears are gonna win the Super Bowl. And I'm like, no. They stink.
Jordan Wilson [00:28:03]:
This year, they might be okay. Anyways, alright. So now we're gonna do the same thing. We're gonna grab a ton of just information just to gobble up tokens, to gobble up the memory. Alright. So I'm in this transcript here from some of our hot take Tuesdays. Alright. So I'm just grabbing some information.
Jordan Wilson [00:28:22]:
I'm jumping back into chat gpt. I'm saying, please summarize this. And, again, remember, all of my words count toward that token count or the memory. Right? In all of chat g p t when it responds to me. That all counts on the memory as well. Alright, so what I'm going to try to do, I'm going to try to get probably like 3, 3, 3, 4,000, something like that. So let's go ahead and I'm grabbing some more, and I'm gonna put it in chat GPT. So I'm saying please summarize this.
Jordan Wilson [00:28:57]:
Okay. Again, you don't need to see this. It's just a bunch of random text just to eat through the context window. I wanna grab a little bit more here just to make sure that we go well over that 32,000 because, yes, it is not, it is not a 128,000. That is only if you're using the API people. So many people who are quote, unquote smart people, people who charge, you know, $500 an hour to talk to, they they tell you that they're wrong. And I'm proving it to you here live. Alright.
Jordan Wilson [00:29:31]:
So here we go. We're at 37,000. So we went over the 32,000 mark. That's what I wanted to do. Now in theory, and we do test this almost every week, because we do that live PPP course almost every week. Right? And when we tell people something, we test. Y'all have no clue how much work goes on behind the scenes at everyday AI to make sure you are the smartest person in AI at your company. This is what we do y'all.
Jordan Wilson [00:30:00]:
Alright. So now we are at 37,000 tokens, and I'm going to ask chat gbt what's my name, what's my favorite color, what's my favorite food, what do I think of the bears? And guess what we have? It forgot. I don't have access to personal data unless shared in the current conversation. Alright? But then you might be saying, wait. It is in the current conversation. You just told it. Okay? Think of this. That 32,000 tokens, that's roughly about 26,000 words.
Jordan Wilson [00:30:38]:
Okay? Give or take. Again, these are estimates. That's why I wanted to go well over 32,000 tokens. Alright? ChatGPT only remembers the most recent 32,000 tokens. Okay? So what that means now, because I'm at 37,000 tokens, it has forgotten
Jordan Wilson [00:31:00]:
the first everything
Jordan Wilson [00:31:02]:
in the first 5,000 tokens. Okay? So, essentially, people are like, oh, so when you get to 33,000, does it forget everything? No. It just remembers the most recent 32,000. So literally think, you know, in our PPP course, we go over some ways to kind of get around this. Right? Number 1 is you just always have to be, you always have to understand the tokenization process and where your memory or your context window is. Okay? So that y'all is a very brief and quick overview of how tokenization works. And this, like I said, is extremely important to understand, especially if you are using Cheggpt, because if I'm being honest, 32,000 tokens is not a lot. Obviously, I was working with you know, I was copying and pasting, like, 15 pages of text at a time, but you saw there, that was technically to hitting the enter key twice.
Jordan Wilson [00:32:01]:
Right? It was only 2 rounds of back and forth because it was a lot of information I was putting in there, obviously, but ChatGPT's memory, especially if
Jordan Wilson [00:32:10]:
you are working with long, long blocks of text, it can go quickly. Alright?
Jordan Wilson [00:32:17]:
So, yeah, Rolando says, always the receipts. You know it. You know we always bring the receipts here. Alright? So hey, Ira. Ira says she wants to check PPP. Chris says it's a great course. Alright. I'll send you I'll send you the information.
Jordan Wilson [00:32:37]:
Michael highly recommends it. Alright? And if Michael joins us from LinkedIn and YouTube, you know that that recommendation comes highly. Alright. So now just just with that y'all, you already know more about how large language models work and chat gpt works than, if I'm being honest, 99.5% of people out there. One of the key factors of getting the most out of large language models in generative AI is understanding it. Right? People always talk about this black box of generative AI. Right? Like, oh, no one knows what goes on under the scenes. Well, now you know a little bit.
Jordan Wilson [00:33:19]:
Right? You know the rules because, essentially, ChatGPT, OpenAI, and and Anthropic, you know, Google Gemini, They create a playing field, and there's rules, there's boundaries, and we all have to play within the confines of the rules that they set. However, because they update these models so often, new features, new functionality, new and improved tokenization processes, different context windows, all of these things are constantly changing. That's why you have to constantly tune in because we spend an insane amount of time every single week keeping up with everything that matters in artificial intelligence. Y'all, this is this is unedited, unscripted, the realest thing, but we research everything so you can be the smartest person in AI at your company. Alright. I think there's a couple questions here. If you have a question, live, please please get it in quickly because we're wrapping up here. Because, yeah, this there's a lot of receipts today, y'all.
Jordan Wilson [00:34:27]:
Kathleen is asking, can you ask it to respond within a token range? Yes and no. So we've tested that similarly to asking it to give you a certain number of words. In our experience, right, and this is not scientific fact, there's actually I don't know if there's actually been research papers on this, but, you know, as an example. And this is why you can't say, oh, how many r's are in strawberry. Right? You can say how many tokens, and it will get it a little closer to being correct.
Jordan Wilson [00:34:57]:
But in our experience, you
Jordan Wilson [00:35:00]:
you know, let's say, hey. Write me a a sentence about strawberry with 10 words, not super accurate. If you say write me a sentence about strawberries with 15 tokens, I it usually gets it a little closer. It is not a science. Again, because the intricacy of tokens, like I said, adding a period changes it. Putting a space after a word changes it. So it does get a little bit closer when you ask it to respond within a token range versus a word range, but it's still not going to be a 100% accurate. Alright, Monica.
Jordan Wilson [00:35:33]:
Good question here. How are you able to continue the same chat when you max out your tokens? Okay. Good question. So just because you get to, as an example, 33,000 tokens, it doesn't mean your chat is over. Okay? So we teach something, in our free prime prompt polish PPP course, which is live, and you can ask questions. We teach something called memory recall. Alright. Because usually what happens when you're using chat g p t, if you're using it correctly, We always tell people use chat like a consultant.
Jordan Wilson [00:36:07]:
Right. And what that usually entails, it's a lot of it's a lot of back and forth, a lot of back and forth conversation. Right? Everyone should actually go check, check out episode 310 if you haven't already. The one chat GPT mistake that we're all making that goes through the process of turning chat GPT into a consultant and more of this, take on augmented intelligence. But to get back to your question, Monica, how are you able to continue the same chat? Well, you can keep going. Right? You can keep, you know, a chat going to a 1000000 tokens. But like I said, it's only going to remember or recall the most recent 32,000 tokens. So what you can do, right, if you know, as an example, oh, I'm getting near the token limit.
Jordan Wilson [00:36:50]:
But a lot of times, there's quote, unquote wasted tokens. Right? Because you may be having a conversation with chat gpt. So So you can do something what's called a memory recall, which is essentially saying something along the lines of, hey. Please recall all important information, as if you are explaining the contents of our conversation to a large language model that has no information. Right? That's just, an an example of a memory recall prompt off the top of my head. And then what will happen, generally is chat gpt is going to recall, the most important information. And then what that does is it pushes the most important think of it like a cliff notes version, then it'll push that at the bottom of the context window. Right? And then for most intent and purposes, right, you're kind of resetting its memory because, think of it like pages as an example.
Jordan Wilson [00:37:41]:
Right? So instead of thinking 32,000 tokens, think 32 pages. Right? And maybe you have some important information on page 5 and some important information on page 10. And eventually, when you get to page 60, it's going to forget that information. So So you can essentially do a memory recall. Like I said, there's probably a lot of fluff in there in those quote unquote 32 pages. And then it's gonna put everything, let's just say as an example, on page 33. Right? All of the most important bullet pointed information. Now it's going to remember kind of that cliff note versions or the bullet point version, for the next 32,000 tokens.
Jordan Wilson [00:38:16]:
Right? So that's just a way that, in theory, you can try to surmise or distill a lot of the most important information and kind of reset it. So like I said, when you get to 33,000 tokens, it doesn't forget everything. Right? So let's just say that when you get to 33,000, it only forgets the first 1,000. Right? But if you're at 64,000 and you haven't done a kind of, quote, unquote, memory recall to get that important information, it's forgotten the top 32, 32,000 tokens. But, actually, what you can do alright. And I I I didn't plan on doing this, Monica, but this is actually a great example here.
Jordan Wilson [00:38:57]:
A little
Jordan Wilson [00:38:57]:
a little hack, y'all. Alright. We'll we'll we'll do some advanced things here. So, I'm going back into, my my chat here. Right? Because Monica had a great question that I think a lot of people are running into. Because in this chat now, we're at 37 1,000 tokens. Alright. So what I can do, because you might be saying, oh, well, if I get past it, how can I how can I ask chat gpt to recall that information? Because it technically can't, it technically can't even find it.
Jordan Wilson [00:39:24]:
Oh, no. I'm screwed. No. You're not. Here's a little trick y'all. So even when you do a memory recall, it you you can only do a memory recall on the most recent 32,000 tokens. Right? But one thing you can do, there's a nice little feature in here in ChatChippy Tea that not a lot of people know about. I, as a human, can go up, to that information.
Jordan Wilson [00:39:44]:
Right? That's past 32,000. Because inside ChatGPT, you get information that or you get access to all of it. So even if your conversation is a 100,000 tokens, you can still see as a human information that ChatGPT can't. So now as an example, I'm going up to this original information, right, that is 5,000 tokens past what ChatGPT can see. So I can as an example, I can highlight, you know, I can highlight this text. Right? Let's say this 5,000 tokens. Right? It's not. Alright.
Jordan Wilson [00:40:17]:
But when I highlight it, if I scroll up to the top here alright. This is gonna be a little tricky. I'm not getting the, I think I'm too far zoomed in. Let's try it again. There it was. Give me a second, y'all. Alright. There we go.
Jordan Wilson [00:40:37]:
So you can highlight certain information, inside of chat gpt. So it might be it might be kinda hard to see here, but there is this, this quotation button. Right? So even if something is outside of the context window, you could just copy and paste it and say summarize this or you can highlight it and then hover, and then you're gonna get this quotation mark inside of chat gpt, and then it says reply. So I can then do this, and I can say, you know, please summarize this, include everything. Right? And now as an example, it's going to recall that information, and then I can ask the exact same question. Right? Let me scroll down here to the bottom. Then I can say, alright. So this is this is interesting here.
Jordan Wilson [00:41:31]:
So it looks like because I didn't say it in full, it's not able to recall it. I've ever I've I've never actually run into that issue. Anyways, to try to accurately, kind of, answer that question, Monica, is at any point, you can go back up, copy and paste something that's, you know, out of the context window, put it back into the context window, and then ChatGPT should be able to then recall and remember that information. Alright. So I
Jordan Wilson [00:42:00]:
know this was a little bit
Jordan Wilson [00:42:03]:
kind of, technical and maybe a little bit dorky, but I hope this was helpful y'all. Because now if you understand and and and and, Denny, I I I think I kind of just answered your your question here about if there's a workaround to continue a conversation once you max out the tokens. Yeah. I just I just kinda showed you, and we go over that in our free prime prompt polish course as well. You essentially just need to put it back into the context window.
Jordan Wilson [00:42:31]:
So that's it y'all. I hope
Jordan Wilson [00:42:33]:
this was helpful. Now you know what tokens are. You know why they're important, and you know why now. Chat Jeepte might start out working for you very well and then not, perform very well. Also, this is why it's extremely important to have crisp language, a strong lexicon when working with chat gpt. Right? The other thing that we didn't dive too deeply into today, is the concept of tokens. Right? And how maybe if you're not super descriptive. Right? How the example that I showed, even the the same word could technically have different tokens because of different meanings.
Jordan Wilson [00:43:14]:
Right? Because as an example, you you might say, hey. You know, I want you to write along with blog post, but don't use the word just. Right? Guess what? The word just, if you say it like that, the word just has many different meanings. Right? I'm just talking. We need a fair and just system. Right? Justly this. Right? It's just the right amount. Right? The word just can have so many different meanings.
Jordan Wilson [00:43:41]:
So sometimes when you are trying to, tell chat GPT to include certain words or avoid certain words, if you're not giving it enough context, if you're not being crisp with your language, it it can struggle. Right? That's why one of the best kind of quote unquote prompt engineering skills that you can have is strong communication skills. Because the better you are talking to chat g p t, the stronger connection it is going to make with that context to better understand what you actually mean and assign the right tokens. Right? That's why sometimes a couple of words can mean all the difference. Alright. I hope this was helpful, y'all. Now you have a complete guide to tokens inside of ChatGPT. Alright.
Jordan Wilson [00:44:28]:
I hope this was helpful. Hey. If it was, go ahead. Sometimes we do this, I don't know, once or twice a month. You know, we charge, if I'm being honest, we charge a couple $100 if you wanna talk to us. We're gonna do it. We're gonna choose 1 person for free. Alright.
Jordan Wilson [00:44:40]:
So anyone, who repost this show. So repost this on, LinkedIn. You can, I guess, retweet this one on, Twitter? Reach out to me. Tell me you've reposted it, whatever, or tag us. Alright? Anyone who does this in, this week by Friday, I'm gonna do a little drawing. Alright? So and then, the winner, I'm gonna give them a little 45 minute session, where you can ask us anything. Right? Companies pay couple couple $100 to to be able to ask us questions. We're gonna do it, for 1 person for free.
Jordan Wilson [00:45:14]:
So if you find this helpful, even if you're listening on the podcast, don't worry. You have time, if you repost this by, Friday, August 16th. So, you know, these shows take us hours to plan and sometimes, hundreds of hours of experience to even have this information to distill it to you all. So if this is helpful, you wanna get any of your questions asked about generative AI, large language model, etcetera, go ahead, repost this to your network, here on LinkedIn or, Twitter, x, whatever. Tag me. Let me know you did it, and I will enter you into a drawing. We're gonna pick 1 person. Alright.
Jordan Wilson [00:45:52]:
So, we appreciate your support. Also, if you're listening on the podcast, make sure you subscribe. Leave us a rating on Apple or Spotify. Like I said, if you if you are a podcast listener, this one might be one to go check out, visually. So, you know, check out, the, LinkedIn post that we include in the show notes. Check out that thanks a million giveaway going through the end of the month, and check us out tomorrow and every day for more everyday AI. Thanks y'all.
