0:00 Music 0:06 So the first part, I'll start explaining very high level, no deep technical details, no prior experience, no understanding of computer science, just a very high level understanding of how 0:22 artificial intelligence works, how it's different, how we think and we work, and what are the limitations, and also how the most basic concepts like actually 0:34 working conceptualize. 0:36 So I'll try to focus on details that from my experience working with founders and like consulting in the AI adoption space, help companies to make better decisions 0:49 and try to kind of focus on those details that actually impact decision-making process. 0:57 So first of all, as I mentioned, basically all modern AI is built on mostly large language models. 1:06 It's a subset of machine learning and adopt the AI role. 1:11 We'll not go into details of that. 1:14 But in a nutshell, contrary to the name, large language model actually inside don't have anything about language themselves. 1:22 What they do and how they're different from kind of our brain operations, they tokenize 1:30 text and languages into figures, numbers. 1:34 So what happens, they break the input language into different tokens. 1:40 And one token can be a word or it can be a part of the word. 1:44 For English language, most of the cases, one token is one word. 1:49 In other languages, this is different because most of the large language models, native language is English. 1:56 And that's why foreign languages are typically consumed more tokens. 2:01 So when you use English to communicate with your LLM, they are usually more efficient because most of the internet is in English. 2:10 So there's not much we can do about it like today. 2:14 Maybe potential later someone will spend efforts and build LLMs intrinsically, natively, or educated and trained in other languages. 2:23 But as of today... 2:25 If you want to communicate with the lens most efficiently, you better use English initially and translate the outputs like for your, 2:36 for example, understanding or for customer facing questions or anything else. 2:42 But my recommendation is to have at least basic prompts as templates in English always. 2:49 And if you need translation at any point of time, just make it the last step. 2:54 Because at the end of the day, AI doesn't care what language you speak and talks. 3:00 But certainly because it was trained and the corpus of data that was trained, post-trained, and evaluated is mostly English, it's better to use it. 3:13 So coming back again, what is token? 3:16 For English, token is usually one word. 3:19 For more complex words, it can be two tokens per word, maybe in some particular more less frequent use cases, three tokens per word. 3:28 And every token has a representation of a particular number of figures. 3:33 So, please can be 3, 5, 6, 7, and lies can be 4, 8, 7, 7, 9, etc. 3:40 So, behind every word there are some numeric representations. 3:44 And these are numbers. 3:48 Are built into huge masses. 3:50 So you encode every piece of information that you input in the sequence of numbers. 3:56 So for a LAN for the machine, every text is just a sequence of numbers in particular order. 4:03 And how the training usually works, you provide the sequence of number on the input 4:08 So for this sequence of input numbers, I would expect this output. 4:14 And from these pairs of inputs and outputs, the training data set is built. 4:19 And it's pretty huge. 4:21 Thousands of actual humans work distributed all across the globe. 4:28 To create these pairs of inputs and outputs. 4:30 They provide text input, text output. 4:32 Machine encodes it into figures and numbers and then puts like into statistical prediction calculated. 4:39 So in a nutshell, LLM is a mass exercise. 4:44 What it does, it calculates the most probable and the most likely 4:53 statistical probable answer or number of sequence of numbers to the sequence of input tokens that you provide. 5:07 And that basically is just very, very complex, huge. 5:13 Sophisticated calculator. 5:15 So it doesn't have intrinsic understanding of the meaning, but it has internal intrinsic probability of the meaning outcome. 5:25 That's why it's so, because it trained on the huge data set, it excels at recognizing patterns. 5:32 So whenever your input and output tokens have some pattern variations in the huge data set, they are most likely to be repeated like in a similar manner. 5:43 So why would you care about tokens? 5:46 Because once you start diverging from using like subscription monthly based, 5:54 to like chat GPT or cloud where you just pay for the monthly and you eventually hit it. 6:00 to limits, but you don't care about the token consumption. 6:03 You just have a flat fee. 6:06 This kind of works to limit until you want to automate things. 6:10 Once you start thinking how to build agentic systems, that's where you have to think about connecting your inference as API. 6:21 API stands for Application Programming Interface. 6:23 It's basically a way to call some external service with your tokens, get the predictions, and pay for the amount of tokens that you provide. 6:36 And the difference in cost can be like really huge. 6:40 So just in this example, if you have like 50 page contract, 6:45 And there are three models from Anthropic with prices as of February 2026. 6:50 If you use Heiko model, you would pay 12 cents. 6:53 If you use most expensive Opus, you will pay 12. 6:57 And for example, if your request or your problem becomes like big and huge, you can buy a lot of information. 7:04 Opus currently charges $75 per million output tokens. 7:10 And because also they also made some upgrades and potentially models get larger and larger context window, it's not unimaginable to burn $75 in just one request if you are careless. 7:25 So, for example, Gemini has 3.1 has 2 million tokens window. 7:30 Sonnet 4.6 already has one million token window, and this is going to evolve. 7:35 So eventually models get larger context window, but the price is something you need to be aware of. 7:42 Because when you start automating things, you'll not get a flat fee. 7:47 You'll be charged for your consumption on a monthly basis. 7:51 That's why you need to be very conscious on what model to choose and how tokens are calculated and estimated. 7:58 The cost goes down, and I watch it very closely. 8:05 But if you look at it, for example, the cost for Opus tokens output, 8:11 hasn't went down for at least six months. 8:15 So, and like several good reasons for that. 8:19 It very much depends on the availability of the models, also the power consumption. 8:26 And yes, eventually it will go down. 8:29 And also, if you are cost sensitive, you can explore some cost saving approaches like hosting your own models or running it elsewhere. 8:39 Still, even if you don't care about the cost, eventually, say you have your own GPU infrastructure, the larger, more expensive compute expensive models. 8:49 Simply respond slower. 8:51 So it will take more time and energy to process the same request and often it might not make sense. 8:58 So in case of 9:00 token economics. 9:02 Right now, I would say that there were no dramatic changes in cost of token, at least from proprietary providers for at least six months. 9:13 It's been pretty expensive. 9:17 Prices sometimes go down a little bit. 9:20 But not dramatic for a while. 9:22 It might change anytime because like in this domain, like every week something goes up, new models get out in the market from China or from other competitors and this might change. 9:35 As of now, it's still... 9:38 Can go pretty costly if you don't think about your token consumption, like by default. 9:45 So you better watch it closely. 9:49 So this is kind of the concept that I was explaining of how the general intelligence works in... 9:58 How the sequence of one figures gets through complex calculations and converted into sequence of another figures. 10:08 And this is how like EI kind of basically LMA functioning. 10:12 And this all is like processed in tokens, which are basically words in English. 10:19 And parts of words in most of the other languages. 10:23 So the price difference can be up to hundreds or even more. 10:28 And one of the opportunities for you 10:30 Kind of to say, for example, usually what I tell kind of people who come to me for advice, if you see that your monthly 10:40 token consumption goes above $2,000 per month, it's probably a good time to think about having your own GPU infrastructure set up for yourself. 10:49 It can be your private server. 10:50 It can be a Neo Cloud provider with GPU hosting infra or something different. 10:57 Because right now, the level of intelligence that OpenWeights models provides is pretty amazing. 11:04 So it's already almost on par with commercial models. 11:07 And most of these providers that are competitive are actually from China. 11:13 So again, if your use case allows to use Chinese models and you really start to feeling the monthly token consumption, there's like a high chance that you can save a lot by switching to 11:28 Chinese model and have your own infrastructure set up and running. 11:33 So the thing about context, and that's where I probably want to... 11:43 Switch here. 11:48 What is context window? 11:51 Give you a high-level idea. 11:52 So context window is the amount of information that LLM can process 12:00 at a time. 12:01 At any single call that your system goes to AI, it puts everything in context window. 12:08 And context window is usually shared between input and output, depending on the model and architecture. 12:14 But the idea is that everything short term, everything kind of that you put into your request, it goes into model context window. 12:23 And it's limited to a certain extent. 12:26 So 200,000 tokens is like approximately 200,000 English words. 12:33 It seems a lot until you start feeding your information about your life, about your business, about your customers, your daily reports. 12:40 It gets filled up pretty quickly. 12:43 And especially when you work with the large amount of customer database, with your users'information, with reports, it fills out very, very fast. 12:53 And also one more important thing to remember about context window, that most of the training data sets 13:02 don't have the input or context at large. 13:06 So usually there are different figures and no proprietary company discloses how much, how long are the input requests in the training data, but they typically don't extend 50,000 tokens. 13:21 Which means that anything that goes beyond 50,000 on average is extrapolation. 13:28 It means that it's like artificial. 13:30 boosting and the quality of the outputs based on your input outside the first 50,000 will very much degrade. 13:38 And also another thing that you need to understand about your context window, and that's also how this knowledge can help you to build better prompts. 13:50 LLM, because of the architecture, they understand better and they keep more attention to what's in the beginning of the context window and what's in the end. 14:02 That means put everything that's important for you at the beginning of your prompt or in the end, and put everything else in the middle. 14:11 Because whatever is in the middle, it's like the cold zone. 14:14 It often are the attention span because how the architecture of the lane works get lost or at least get less attention that you would like it to. 14:26 So don't get surprised if you kind of provide a huge contract or a huge report and AI gets some figures that are in the middle of your contract wrong. 14:38 It's like actually how it works. 14:39 It's attention span is mostly focused on only on the beginning and on the end. 14:46 So how to solve the challenge of your information, your knowledge, not fitting into the context window? 14:57 And that's how... 15:00 RAG was implemented. 15:01 So has anyone of you heard about what RAG is? 15:06 Retrieval Augmented Generation. 15:10 Probably a couple of times. 15:11 So it's been quite wild in the market. 15:15 So RAG is a process of equipping your AI with a tool to request information from your knowledge base in the chunks that fit within your context window. 15:27 So it's a process of retrieving only relevant pieces of information from your knowledge so that they can fit within your context window. 15:34 They don't pollute it with extra necessary data information and get only pieces of the data that are relevant for your particular request. 15:46 And this... 15:48 Usually works with a large amount of data. 15:51 And I'll just give you like very high level understanding of how RAG and biddings work, just for your like overall knowledge. 15:59 But overall, the RAG architectures have become pretty sophisticated recently and building a good RAG. 16:07 Requires some efforts and some understanding of its fundamentals. 16:12 So let me give you... 16:16 Yeah. 16:17 So racks are typically built on the concept that is called embeddings. 16:23 So embedding is an interesting concept. 16:27 It's basically extracting meaning. 16:30 of your information into math and performing mathematical calculations on the meaning, which means you can do some 16:40 complex or like weird sometimes calculations like concatenation and also means that similar meanings will be close to each other. 16:52 So for example, like here in this example, anything that has to be about revenue forecast, projected income or expected earnings in your request, they will be 17:02 clustered together in the embedding values, and you can mathematically perform some calculations. 17:09 Hey, find me something that has to do with revenue. 17:12 And it will, because how embeddings work, they will pull all the information and semantic meaning that's related to revenue, financials, or similar, and rank them according to relevance. 17:26 So this way you can ask your RAC database that consists of embeddings, some semantic queries, and get relevant search results from your knowledge database. 17:40 And by providing these chunks of relevant information to your AI, 17:45 you can use your context window more efficiently, more reliably, and have it kind of relevant and have higher quality context. 18:00 feeling. 18:02 So in the end of the day, RAC is a process of converting your knowledge base into database of embeddings. 18:10 Embeddings, again, are just figures behind your semantic meaning and keywords, and then performing mathematical operations with the meaning. 18:19 So you can 18:20 are always retrieve something similar, some similar topics, get them clustered together, prioritized and provided in the input of your 18:33 AI inference call, for example. 18:41 So once you have that meaning extracted, it's being fed back into the context window, but only the parts that are most relevant. 18:51 And the complexity behind RAG architecture is to engineer this extraction part 18:59 so that the chunks of data information and semantic meanings that extract human knowledge is tailored for particular requests. 19:07 And only the pieces of data that are relevant to current workload or current execution for your agent actually gets into this context window.
0:00 Music 0:06 So the first part, I'll start explaining very high level, no deep technical details, no prior experience, no understanding of computer science, just a very high level understanding of how 0:22 artificial intelligence works, how it's different, how we think and we work, and what are the limitations, and also how the most basic concepts like actually 0:34 working conceptualize. 0:36 So I'll try to focus on details that from my experience working with founders and like consulting in the AI adoption space, help companies to make better decisions 0:49 and try to kind of focus on those details that actually impact decision-making process. 0:57 So first of all, as I mentioned, basically all modern AI is built on mostly large language models. 1:06 It's a subset of machine learning and adopt the AI role. 1:11 We'll not go into details of that. 1:14 But in a nutshell, contrary to the name, large language model actually inside don't have anything about language themselves. 1:22 What they do and how they're different from kind of our brain operations, they tokenize 1:30 text and languages into figures, numbers. 1:34 So what happens, they break the input language into different tokens. 1:40 And one token can be a word or it can be a part of the word. 1:44 For English language, most of the cases, one token is one word. 1:49 In other languages, this is different because most of the large language models, native language is English. 1:56 And that's why foreign languages are typically consumed more tokens. 2:01 So when you use English to communicate with your LLM, they are usually more efficient because most of the internet is in English. 2:10 So there's not much we can do about it like today. 2:14 Maybe potential later someone will spend efforts and build LLMs intrinsically, natively, or educated and trained in other languages. 2:23 But as of today... 2:25 If you want to communicate with the lens most efficiently, you better use English initially and translate the outputs like for your, 2:36 for example, understanding or for customer facing questions or anything else. 2:42 But my recommendation is to have at least basic prompts as templates in English always. 2:49 And if you need translation at any point of time, just make it the last step. 2:54 Because at the end of the day, AI doesn't care what language you speak and talks. 3:00 But certainly because it was trained and the corpus of data that was trained, post-trained, and evaluated is mostly English, it's better to use it. 3:13 So coming back again, what is token? 3:16 For English, token is usually one word. 3:19 For more complex words, it can be two tokens per word, maybe in some particular more less frequent use cases, three tokens per word. 3:28 And every token has a representation of a particular number of figures. 3:33 So, please can be 3, 5, 6, 7, and lies can be 4, 8, 7, 7, 9, etc. 3:40 So, behind every word there are some numeric representations. 3:44 And these are numbers. 3:48 Are built into huge masses. 3:50 So you encode every piece of information that you input in the sequence of numbers. 3:56 So for a LAN for the machine, every text is just a sequence of numbers in particular order. 4:03 And how the training usually works, you provide the sequence of number on the input 4:08 So for this sequence of input numbers, I would expect this output. 4:14 And from these pairs of inputs and outputs, the training data set is built. 4:19 And it's pretty huge. 4:21 Thousands of actual humans work distributed all across the globe. 4:28 To create these pairs of inputs and outputs. 4:30 They provide text input, text output. 4:32 Machine encodes it into figures and numbers and then puts like into statistical prediction calculated. 4:39 So in a nutshell, LLM is a mass exercise. 4:44 What it does, it calculates the most probable and the most likely 4:53 statistical probable answer or number of sequence of numbers to the sequence of input tokens that you provide. 5:07 And that basically is just very, very complex, huge. 5:13 Sophisticated calculator. 5:15 So it doesn't have intrinsic understanding of the meaning, but it has internal intrinsic probability of the meaning outcome. 5:25 That's why it's so, because it trained on the huge data set, it excels at recognizing patterns. 5:32 So whenever your input and output tokens have some pattern variations in the huge data set, they are most likely to be repeated like in a similar manner. 5:43 So why would you care about tokens? 5:46 Because once you start diverging from using like subscription monthly based, 5:54 to like chat GPT or cloud where you just pay for the monthly and you eventually hit it. 6:00 to limits, but you don't care about the token consumption. 6:03 You just have a flat fee. 6:06 This kind of works to limit until you want to automate things. 6:10 Once you start thinking how to build agentic systems, that's where you have to think about connecting your inference as API. 6:21 API stands for Application Programming Interface. 6:23 It's basically a way to call some external service with your tokens, get the predictions, and pay for the amount of tokens that you provide. 6:36 And the difference in cost can be like really huge. 6:40 So just in this example, if you have like 50 page contract, 6:45 And there are three models from Anthropic with prices as of February 2026. 6:50 If you use Heiko model, you would pay 12 cents. 6:53 If you use most expensive Opus, you will pay 12. 6:57 And for example, if your request or your problem becomes like big and huge, you can buy a lot of information. 7:04 Opus currently charges $75 per million output tokens. 7:10 And because also they also made some upgrades and potentially models get larger and larger context window, it's not unimaginable to burn $75 in just one request if you are careless. 7:25 So, for example, Gemini has 3.1 has 2 million tokens window. 7:30 Sonnet 4.6 already has one million token window, and this is going to evolve. 7:35 So eventually models get larger context window, but the price is something you need to be aware of. 7:42 Because when you start automating things, you'll not get a flat fee. 7:47 You'll be charged for your consumption on a monthly basis. 7:51 That's why you need to be very conscious on what model to choose and how tokens are calculated and estimated. 7:58 The cost goes down, and I watch it very closely. 8:05 But if you look at it, for example, the cost for Opus tokens output, 8:11 hasn't went down for at least six months. 8:15 So, and like several good reasons for that. 8:19 It very much depends on the availability of the models, also the power consumption. 8:26 And yes, eventually it will go down. 8:29 And also, if you are cost sensitive, you can explore some cost saving approaches like hosting your own models or running it elsewhere. 8:39 Still, even if you don't care about the cost, eventually, say you have your own GPU infrastructure, the larger, more expensive compute expensive models. 8:49 Simply respond slower. 8:51 So it will take more time and energy to process the same request and often it might not make sense. 8:58 So in case of 9:00 token economics. 9:02 Right now, I would say that there were no dramatic changes in cost of token, at least from proprietary providers for at least six months. 9:13 It's been pretty expensive. 9:17 Prices sometimes go down a little bit. 9:20 But not dramatic for a while. 9:22 It might change anytime because like in this domain, like every week something goes up, new models get out in the market from China or from other competitors and this might change. 9:35 As of now, it's still... 9:38 Can go pretty costly if you don't think about your token consumption, like by default. 9:45 So you better watch it closely. 9:49 So this is kind of the concept that I was explaining of how the general intelligence works in... 9:58 How the sequence of one figures gets through complex calculations and converted into sequence of another figures. 10:08 And this is how like EI kind of basically LMA functioning. 10:12 And this all is like processed in tokens, which are basically words in English. 10:19 And parts of words in most of the other languages. 10:23 So the price difference can be up to hundreds or even more. 10:28 And one of the opportunities for you 10:30 Kind of to say, for example, usually what I tell kind of people who come to me for advice, if you see that your monthly 10:40 token consumption goes above $2,000 per month, it's probably a good time to think about having your own GPU infrastructure set up for yourself. 10:49 It can be your private server. 10:50 It can be a Neo Cloud provider with GPU hosting infra or something different. 10:57 Because right now, the level of intelligence that OpenWeights models provides is pretty amazing. 11:04 So it's already almost on par with commercial models. 11:07 And most of these providers that are competitive are actually from China. 11:13 So again, if your use case allows to use Chinese models and you really start to feeling the monthly token consumption, there's like a high chance that you can save a lot by switching to 11:28 Chinese model and have your own infrastructure set up and running. 11:33 So the thing about context, and that's where I probably want to... 11:43 Switch here. 11:48 What is context window? 11:51 Give you a high-level idea. 11:52 So context window is the amount of information that LLM can process 12:00 at a time. 12:01 At any single call that your system goes to AI, it puts everything in context window. 12:08 And context window is usually shared between input and output, depending on the model and architecture. 12:14 But the idea is that everything short term, everything kind of that you put into your request, it goes into model context window. 12:23 And it's limited to a certain extent. 12:26 So 200,000 tokens is like approximately 200,000 English words. 12:33 It seems a lot until you start feeding your information about your life, about your business, about your customers, your daily reports. 12:40 It gets filled up pretty quickly. 12:43 And especially when you work with the large amount of customer database, with your users'information, with reports, it fills out very, very fast. 12:53 And also one more important thing to remember about context window, that most of the training data sets 13:02 don't have the input or context at large. 13:06 So usually there are different figures and no proprietary company discloses how much, how long are the input requests in the training data, but they typically don't extend 50,000 tokens. 13:21 Which means that anything that goes beyond 50,000 on average is extrapolation. 13:28 It means that it's like artificial. 13:30 boosting and the quality of the outputs based on your input outside the first 50,000 will very much degrade. 13:38 And also another thing that you need to understand about your context window, and that's also how this knowledge can help you to build better prompts. 13:50 LLM, because of the architecture, they understand better and they keep more attention to what's in the beginning of the context window and what's in the end. 14:02 That means put everything that's important for you at the beginning of your prompt or in the end, and put everything else in the middle. 14:11 Because whatever is in the middle, it's like the cold zone. 14:14 It often are the attention span because how the architecture of the lane works get lost or at least get less attention that you would like it to. 14:26 So don't get surprised if you kind of provide a huge contract or a huge report and AI gets some figures that are in the middle of your contract wrong. 14:38 It's like actually how it works. 14:39 It's attention span is mostly focused on only on the beginning and on the end. 14:46 So how to solve the challenge of your information, your knowledge, not fitting into the context window? 14:57 And that's how... 15:00 RAG was implemented. 15:01 So has anyone of you heard about what RAG is? 15:06 Retrieval Augmented Generation. 15:10 Probably a couple of times. 15:11 So it's been quite wild in the market. 15:15 So RAG is a process of equipping your AI with a tool to request information from your knowledge base in the chunks that fit within your context window. 15:27 So it's a process of retrieving only relevant pieces of information from your knowledge so that they can fit within your context window. 15:34 They don't pollute it with extra necessary data information and get only pieces of the data that are relevant for your particular request. 15:46 And this... 15:48 Usually works with a large amount of data. 15:51 And I'll just give you like very high level understanding of how RAG and biddings work, just for your like overall knowledge. 15:59 But overall, the RAG architectures have become pretty sophisticated recently and building a good RAG. 16:07 Requires some efforts and some understanding of its fundamentals. 16:12 So let me give you... 16:16 Yeah. 16:17 So racks are typically built on the concept that is called embeddings. 16:23 So embedding is an interesting concept. 16:27 It's basically extracting meaning. 16:30 of your information into math and performing mathematical calculations on the meaning, which means you can do some 16:40 complex or like weird sometimes calculations like concatenation and also means that similar meanings will be close to each other. 16:52 So for example, like here in this example, anything that has to be about revenue forecast, projected income or expected earnings in your request, they will be 17:02 clustered together in the embedding values, and you can mathematically perform some calculations. 17:09 Hey, find me something that has to do with revenue. 17:12 And it will, because how embeddings work, they will pull all the information and semantic meaning that's related to revenue, financials, or similar, and rank them according to relevance. 17:26 So this way you can ask your RAC database that consists of embeddings, some semantic queries, and get relevant search results from your knowledge database. 17:40 And by providing these chunks of relevant information to your AI, 17:45 you can use your context window more efficiently, more reliably, and have it kind of relevant and have higher quality context. 18:00 feeling. 18:02 So in the end of the day, RAC is a process of converting your knowledge base into database of embeddings. 18:10 Embeddings, again, are just figures behind your semantic meaning and keywords, and then performing mathematical operations with the meaning. 18:19 So you can 18:20 are always retrieve something similar, some similar topics, get them clustered together, prioritized and provided in the input of your 18:33 AI inference call, for example. 18:41 So once you have that meaning extracted, it's being fed back into the context window, but only the parts that are most relevant. 18:51 And the complexity behind RAG architecture is to engineer this extraction part 18:59 so that the chunks of data information and semantic meanings that extract human knowledge is tailored for particular requests. 19:07 And only the pieces of data that are relevant to current workload or current execution for your agent actually gets into this context window.