A chatbot can draft a cover letter, untangle a bug in your code, and explain a tax form in the same minute. Under the hood, it is doing one narrow thing over and over: predicting what small piece of text should come next. Once that clicks, how large language models work stops feeling like magic and starts feeling like very large-scale pattern matching, with all the strengths and blind spots that implies.
Step One: Chopping Text Into Tokens
Models do not read words the way you do. They split text into tokens, which Google’s machine learning course describes as a word, a subword, or even a single character. OpenAI’s help center adds that a token can be a character, part of a word, a whole word, or punctuation, and offers a rough English rule of thumb: about four characters, or three-quarters of a word, per token. Other languages split differently, so the ratio shifts.
This matters in practice. Usage limits, pricing, and how much text a model can hold in mind at once are all counted in tokens, not words. A long document that feels short to you may be bigger than you think from the model’s side.
Step Two: Learning by Filling in Blanks
The training recipe sounds almost too simple. Developers feed a model enormous amounts of text with pieces hidden, and the model learns by guessing the missing tokens. Google’s course explains that models are trained by trying to predict those missing tokens across massive datasets, and that this takes enormous computing resources and electricity. That hardware story is its own topic, covered in our look at AI chips and what CPUs, GPUs, and NPUs do.
No one hand-writes rules about grammar, facts, or coding style. Those patterns emerge because predicting the next token well requires picking up on them. A model that has seen countless recipes, support emails, and code snippets learns what usually follows what.
Step Three: Attention Decides What Matters
Most modern models use an architecture called the Transformer. Its key trick is self-attention, a mechanism that asks, for each token, how much every other token in the input affects its meaning. Take the sentence “The trophy didn’t fit in the suitcase because it was too big.” Attention helps the model link “it” to the trophy rather than the suitcase, because the surrounding words point that way.
Stack many layers of this, add billions of adjustable numbers called parameters, and the model can track context across paragraphs rather than a few words. That is a big reason these systems outperform older language tools.
From Raw Predictor to Useful Assistant
A model trained only to continue text is not yet a good assistant. Developers typically refine it afterward with additional training on examples of helpful answers and with human feedback, so it follows instructions instead of rambling onward. The details differ by company and are rarely published in full, so treat any single description of that stage as a general outline rather than a recipe.
Why Fluent Does Not Mean Correct
Because the goal is a plausible next token, not a verified fact, a model can produce a confident sentence that is simply wrong. Google’s course notes that these systems can hallucinate, generating factually incorrect content. We unpack that failure mode in why chatbots get things wrong.
A few other limits follow from the design:
- Limited memory per conversation. The model only sees what fits in its context window, so details from far back can drop out.
- A training cutoff. Knowledge stops at a point in time unless the product adds search or other tools.
- Sensitivity to wording. Small changes in a request can change the answer, which is why learning to write better AI prompts pays off.
What This Means When You Use One
Treat a language model as a fast, well-read collaborator that needs checking, not an oracle. Give it context, ask it to show its reasoning on anything important, and verify names, numbers, and quotes against primary sources before you rely on them. For creative drafting, summaries, and brainstorming, the prediction approach shines. For legal, medical, or financial specifics, it is a starting point, never the final word.
FAQs
Do large language models understand what they write?
They model statistical relationships between tokens very well, which often looks like understanding, but they have no built-in way to check a claim against the real world. That is why they can sound certain while being wrong.
What is a parameter in a language model?
A parameter is an adjustable number inside the neural network that training tunes. Models have vastly more of them than older approaches, which is part of why they gather broader context and perform better.
Why do AI chatbots sometimes give different answers to the same question?
Generation involves choosing among likely next tokens, and that choice can vary from run to run. Wording changes in your question also shift which tokens look most likely.
Is a token the same as a word?
No. A token can be a whole word, part of a word, a character, or punctuation. OpenAI’s rule of thumb for English is roughly four characters, or about three-quarters of a word, per token.












Discussion about this post