
Image: Olkeri
By Olkeri.space
What Is a Large Language Model? How LLMs Actually Work, Explained Simply
A clear, non-technical guide to what large language models are, how they are trained, why they hallucinate, and what they can and cannot do.
Read this story in: Français · Deutsch · Español
A large language model, usually shortened to LLM, is a computer program trained to predict text. That single sentence explains far more about modern artificial intelligence than most people expect, because nearly every AI product in the news today, from chatbots to coding assistants to customer service agents, is built on this one idea.
This guide explains what an LLM is, how it is built, why it behaves the way it does, and where its real limits lie. No mathematics is required.
What a large language model actually is:
An LLM is a very large statistical model of language. During training it is shown enormous quantities of text and given one repetitive task: given everything so far, predict what comes next. Over billions of repetitions, the model adjusts internal numbers, called parameters, until its predictions get very good.
Modern models contain hundreds of billions of these parameters. They are not facts stored in a database. They are weights that encode patterns: which words tend to follow which, how questions relate to answers, how a legal clause differs in tone from a recipe, how code is structured.
When you type a question, the model is not looking anything up. It is generating a plausible continuation of your text, one token at a time, where a token is roughly a word fragment. Everything an LLM does, including reasoning, translating and writing code, emerges from that prediction process.
How an LLM is trained:
Training happens in stages, and understanding them explains most model behaviour.
The first stage is pre-training. The model reads a very large body of text: books, websites, code repositories, reference works and more. It learns grammar, facts, styles, reasoning patterns and structure. This stage is by far the most expensive, requiring thousands of specialised chips running for weeks or months, which is why only well funded organisations build frontier models from scratch.
The second stage is fine-tuning. A pre-trained model will happily continue any text, including unhelpful or harmful text. Fine-tuning teaches it to behave like an assistant, using curated examples of good responses.
The third stage is alignment, often using human feedback. People compare model outputs and indicate which are better. The model is then adjusted to produce responses people prefer: more helpful, more honest, less harmful. This is where a raw text predictor becomes something you can usefully talk to.
Why the transformer mattered:
Almost all current models use an architecture called the transformer, introduced by Google researchers in 2017. Its key mechanism is attention, which lets the model weigh how relevant every part of the input is to every other part.
Earlier approaches processed text strictly in order and struggled to connect distant ideas. Attention lets a model link a pronoun to a name mentioned three paragraphs earlier, or connect a function definition to its use later in a file. Just as importantly, transformers train efficiently on modern parallel hardware, which is what made scaling to today's model sizes practical.
The context window, and why it matters:
The context window is how much text a model can consider at once, measured in tokens. Everything in the window, your question, the conversation so far, any documents you paste, is what the model can actually see.
Early models handled a few thousand tokens, roughly a long article. Current models handle hundreds of thousands, equivalent to entire books or large codebases. A larger window means you can hand a model a full contract or repository instead of fragments.
Anything outside the window does not exist for the model. This is why long conversations lose earlier details, and why systems that need to reference large archives must retrieve relevant pieces and place them into the window rather than relying on memory.
Why models hallucinate:
Hallucination, where a model states something false with complete confidence, is the most consequential limitation in practice.
The cause is structural, not a bug that will simply be patched away. The model is trained to produce plausible text, and a fluent, confident, incorrect answer is often more statistically plausible than an admission of uncertainty. The model has no built-in mechanism for checking a claim against reality. It has patterns, not knowledge in the human sense.
Hallucination risk rises sharply with obscure facts, specific numbers, citations, quotations and recent events. It falls when the model is given the source material directly, which is the entire logic behind retrieval systems.
The practical rule: an LLM is reliable for reasoning over text you supply, and unreliable as a substitute for a database or a search engine.
Training data has a cut-off:
A model's knowledge stops at the point its training data ends. Ask about something after that date and it will either say it does not know or, worse, guess. Products that appear to know current events are usually retrieving live search results and placing them into the context window, not recalling them.
What LLMs are genuinely good at:
The pattern in successful deployments is consistent. LLMs excel at transforming text: summarising, rewriting, translating, extracting structured data from unstructured documents, converting requirements into code, drafting and reformatting.
They are strong at tasks where a competent draft saves substantial time and a human reviews the result. They are weakest where accuracy is non-negotiable, verification is impossible, or the cost of a confident error is high.
What comes next:
The frontier is moving in three directions. Models that reason for longer before answering, trading time for accuracy on hard problems. Multimodal models that handle images, audio and video as naturally as text. And agents that take actions in software rather than only producing text.
None of these changes the fundamental picture. Underneath, the system is still predicting what comes next, extremely well, at enormous scale. Understanding that is the difference between using these tools effectively and being surprised by them.