Published on September 19, 2026
What Is an LLM (Large Language Model) and How Does It Work: a Simple Guide
- artificial-intelligence
- llm
In the first article of this series I explained where a Large Language Model sits relative to artificial intelligence, machine learning, and deep learning: it's essentially a deep neural network built on the Transformer architecture, trained on enormous amounts of text to estimate the most probable continuation of a sequence of words. In this article I want to pop the hood and look more closely at the gears: what actually happens, step by step, between the moment you write a prompt and the moment a response appears on your screen.
You don't need any math to understand this at a serious level — but you do need a bit of patience, because the pieces are varied and they fit together. Let's go in order: first how text gets "digested" by the model (tokenization and embeddings), then how it gets processed (the Transformer architecture and attention), then how the model is trained across distinct stages, and finally what concretely happens when you write a prompt and hit enter.
First things first: a computer doesn't read words, it reads numbers
A language model is, at its core, a giant mathematical function. Mathematical functions work with numbers, not letters or words. So the very first problem to solve, before we even get to talk about "intelligence," is purely technical: how do you turn text into something a model can actually process?
Tokenization: breaking text into pieces
The first step is called tokenization: the input text is broken down into smaller units called tokens. A token, as OpenAI's own documentation explains, can be as short as a single character or as long as a whole word, depending on the language and context — spaces, punctuation, and partial words all contribute to the token count. In practice, for English, a common word often maps to a single token, while rare or compound words get split into several pieces (for example, "tokenization" might become "token" + "ization").
The most widely used technique for deciding how to split text is called Byte Pair Encoding (BPE), adapted for natural language processing in 2016 by a group of researchers at the University of Edinburgh in a paper that became foundational: "Neural Machine Translation of Rare Words with Subword Units" (Sennrich, Haddow, Birch, ACL 2016). The underlying idea is elegant: instead of having a fixed vocabulary of whole words (which could never contain every possible word, including made-up words, proper names, or typos), you build a vocabulary of "word pieces" starting from individual characters and progressively merging the pairs of symbols that appear together most often in the training text. The result is a vocabulary that can represent any word, even one it has never seen before, by breaking it down into already-known sub-units. This same approach, in various forms, underlies the tokenization used by GPT, BERT, and most modern language models.
Each token is finally converted into a number (an identifier pointing to an entry in the vocabulary), and this sequence of numbers is what actually enters the model. The vocabulary of tokens a model works with isn't infinite: it's a fixed set, built once during tokenizer preparation, that typically contains tens of thousands of distinct entries spanning individual characters, word fragments, and the most common whole words.
Uses the cl100k_base encoding (OpenAI, open-source js-tiktoken library). Claude's tokenizer isn't public in the same way, but the underlying mechanism — splitting text into frequent sub-units — is the same one described above.
Embeddings: from numbers to meaning
A plain identifier number, though, says nothing about a word's meaning: "cat" and "dog" might have IDs 4021 and 88, numbers with no relationship to each other despite being fairly close concepts (both domestic animals). So an additional step is needed: each token gets transformed into an embedding, that is, a vector — a list of hundreds or thousands of decimal numbers — representing that token in a multidimensional mathematical space.
Google's definition in its introductory machine learning course is precise: an embedding is "a vector representation of data in embedding space," obtained by projecting high-dimensional data into a lower-dimensional but more meaning-dense space (Google, Machine Learning Crash Course — Embeddings). The central idea is that, ideally, "an embedding captures some of the semantics of the input by placing semantically similar inputs close together in the embedding space" (same source). In practice, after training, the vectors for words with similar meanings or usage ("cat" and "dog," or "king" and "queen") end up close to each other in this numerical space, while unrelated words end up far apart. Nobody writes these vectors by hand: they emerge automatically during training, as a consequence of how words are actually used across the text.
The heart of the system: the Transformer architecture
At this point we have a sequence of numerical vectors representing the input text. The next step is the conceptually trickiest part: how does the model "understand" the relationships between words in a sentence, including ones far apart from each other?
Before 2017, the dominant approaches processed text word by word, in a rigid sequence — a bit like reading a sentence aloud without being able to easily go back. This made it hard to capture relationships between distant words in a sentence, and it was also slow to train, since each step had to wait for the previous one to finish.
The shift came with the paper "Attention Is All You Need," published in 2017 by a group of researchers at Google Brain (Vaswani et al., arXiv:1706.03762), which introduced the Transformer architecture. The key ingredient is a mechanism called self-attention. As described by Google's own research blog that introduced the paper, self-attention lets the model "aggregate information from all of the other words" in the sentence, generating a new representation for each word informed by the entire context, repeating this step multiple times, in parallel, for all the words at once (Google Research Blog, 2017).
What does that mean in practice, without the formulas? Take the sentence: "The bank refused the loan because it didn't have enough collateral." To understand what "it" refers to, the model needs to connect that word to "the bank," not to "the loan" — even though "loan" is the word immediately before it. The attention mechanism is exactly what makes this possible: for every word, the model computes how "relevant" every other word in the sentence is for interpreting it correctly, and uses that relevance score to build a more informed representation. This mutual-relevance computation is done for all pairs of words at once (hence the name "attention"), and it's repeated across multiple "layers" of the network, each of which further refines the representation.
One practical advantage, on top of the conceptual one, turned out to be decisive for the history of LLMs: unlike earlier sequential approaches, attention across all the words can be computed in parallel on modern hardware (GPUs), rather than word by word in sequence. That's what made it possible to train much larger models in far less time — it's one of the direct reasons why, from 2017 onward, language model size grew so quickly.
Training: not one pass, but several distinct stages
A common misconception is imagining that an LLM gets "trained" through a single process, all at once. In reality, training modern models typically unfolds across distinct stages, each with a different purpose.
Stage 1: pre-training
During pre-training, the model is exposed to enormous amounts of text — web pages, books, code, articles — with a task that's simple to describe but computationally massive to run: predict the next token, given everything that came before it. OpenAI's technical report on GPT-4 describes it this way: "GPT-4 is a Transformer-style model pre-trained to predict the next token in a document," using both publicly available data and data licensed from third-party providers (OpenAI, GPT-4 Technical Report, 2023). By repeating this task billions of times across wildly different texts, the model ends up absorbing, as a side effect, an enormous amount of linguistic, factual, and even implicit reasoning regularity present in the text — not because it "understands" those facts the way a human does, but because correctly predicting the next token statistically requires having picked up on those regularities.
Pre-training is by far the most expensive stage in terms of data and compute, and it's the one that determines most of the model's "raw" capabilities. OpenAI's own technical report notes an interesting point: the model's capabilities come primarily from pre-training, while the later stages serve more to steer its behavior than to add core competencies.
Stage 2: fine-tuning and alignment
A model that only knows how to predict "what word comes next" doesn't, by default, know how to behave like a helpful assistant: it might continue a question with more similar questions (because that's what appears most often in its training data, in a list of FAQs, say) instead of actually answering it. So a second stage is needed, generally called fine-tuning, in which the model is further trained on a smaller, curated set of examples — often conversations written by people demonstrating the desired behavior — to teach it to follow instructions and respond appropriately.
A technique that has become central to this stage is RLHF, Reinforcement Learning from Human Feedback. The idea, first described systematically at scale in OpenAI's paper on InstructGPT (Ouyang et al., 2022, arXiv:2203.02155), works like this: human reviewers are shown several responses the model generated for the same question and asked to indicate which one they prefer. Those preferences are used to train a second model (a "reward model") that learns to estimate how much a human evaluator would like a given response; the original language model is then further tuned to generate responses that the reward model judges more favorably. The GPT-4 technical report confirms that the model "is then fine-tuned using reinforcement learning with human feedback (RLHF) to align it with the user's intent" (OpenAI, GPT-4 Technical Report, 2023).
An alternative and extension to this approach, developed by Anthropic, is Constitutional AI: instead of relying solely on human reviewers to judge every single response, the model is guided by an explicit set of principles (a "constitution") and learns to critique and revise its own responses to better align with those principles, with much less direct human involvement (Bai et al., Constitutional AI: Harmlessness from AI Feedback, Anthropic, 2022). The stated goal of that work was an assistant that is simultaneously more helpful and harder to misuse for harmful purposes, easing a tension that traditional RLHF alone had struggled to resolve.
The context window: the model's "working memory"
Before getting to what happens when you write a prompt, it's worth introducing a very concrete, practical concept: the context window. Anthropic's documentation defines it this way: "the 'context window' refers to all the text a language model can reference when generating a response, including the response itself. This is different from the large corpus of data the language model was trained on, and instead represents a 'working memory' for the model" (Anthropic, Context windows).
This is a fundamental distinction worth keeping in mind: pre-training builds, once and for all, the neural network's "weights" (the model's internal parameters, the end result of all its training); the context window, on the other hand, is the limited space — measured in tokens — that the model can "see" within a single conversation, which includes the system prompt, every message exchanged so far, any attached documents, and the response the model is currently generating. A model doesn't "remember" previous conversations once they're closed (unless they're explicitly fed back into the context): absent additional mechanisms, every new request starts with no memory of anything that happened outside the current window.
An interesting detail from the same documentation: making the context window bigger isn't automatically a pure win. As token count grows, accuracy and the ability to recall specific information tend to get worse, a phenomenon Anthropic calls "context rot" (Anthropic, Context windows). Curating what goes into the context, in other words, matters just as much as how much space is available.
What really happens when you write a prompt
We finally get to the practical moment: you type a question and hit enter. Here, in sequence, is what happens behind the scenes, pulling together everything covered so far:
- Tokenization. The prompt's text (plus the entire prior conversation still inside the context window) gets split into tokens, using the same Byte Pair Encoding scheme described earlier.
- Embedding. Each token is converted into its corresponding numerical vector.
- Processing through the Transformer. The vectors pass through the network's various layers, where self-attention mechanisms repeatedly recompute each token's representation in light of the context provided by every other token.
- Probability computation. At the output of the last layer, the model produces a probability distribution over every possible token in its vocabulary: essentially, a score for each candidate token representing how plausible it is as the next word, given everything that came before. As a simplified example: after "The cat sat on the...," the model might assign a high probability to "mat" or "couch," a lower but non-zero probability to "roof" or "ledge," and a near-zero probability to words that wouldn't make sense at that point in the sentence, like an incorrectly conjugated verb.
- Token sampling. The system picks a token from this distribution — not always the single most probable one: depending on the settings (often called "temperature"), it can be chosen with some controlled randomness among the more likely options, rather than always taking whichever one sits at the top of the list. Lower temperature makes responses more predictable and repetitive; higher temperature makes them more varied, but also riskier in terms of coherence.
- Repeat, one token at a time. The newly generated token gets appended to the sequence, and the entire process (tokenization already done, embedding, passing through the Transformer, computing probabilities, sampling) repeats from scratch to generate the next token. This continues until the model generates a special "end of response" token or a length limit is reached.
Step (6) is probably the most counterintuitive part for anyone who's never had it explained before: when we read an LLM's response that appears "all at once" (or nearly so, if we're not watching it stream word by word), it was actually built, quite literally, one word — or rather one word-piece — at a time, each one conditioned on everything generated so far, including the original prompt. This process is called inference, to distinguish it from training: during inference, the network's weights no longer change; the model simply uses what it learned to generate, one step at a time, the most plausible output sequence.
Why this mechanism matters
Knowing that an LLM generates text token by token, based on probabilities learned during training rather than direct access to a database of facts, isn't a detail only for specialists — it's the key to understanding both its strengths and its most serious limitations. It's exactly this mechanism — a statistical model estimating the most probable continuation, not an archive of verified truths — that gives rise to the "hallucination" phenomenon, which I cover in the third and final article of this series. Understanding how the machine works, in other words, is the first step toward using it well.