You type a question, and something on the other end writes back like it gets you. It’s tempting to imagine a tiny mind in there, thinking it over. It’s not that — and honestly, the real answer is stranger and more interesting than “robot brain.” Let’s build it up piece by piece.
An LLM is a giant autocomplete
A Large Language Model (that’s the “LLM” in the name) has one job: predict what word comes next. That’s genuinely it. Your phone’s keyboard does a tiny version of this when it suggests the next word as you type — an LLM is that same idea, scaled up by a factor of a few billion.
It read a mind-boggling amount of text — books, websites, forums, articles — and picked up the patterns in how words follow other words. Ask it something, and it’s not “looking up” an answer. It’s writing, one predicted piece at a time, the response that its training says fits best.
It doesn’t read words — it reads “tokens”
Before it reads anything, your sentence gets chopped into little chunks called tokens — sometimes a whole word, sometimes just a piece of one. Here’s roughly how “unbelievable results” gets split up:
un believ able results
This is why an LLM can handle words it’s never technically seen before — it just reassembles them from familiar pieces, the same way you can sound out a word you’ve never read.
Guessing the next piece, one at a time
Once it has your tokens, it plays a very fast, very informed guessing game. Say the sentence so far is “The capital of France is” — it doesn’t “know” Paris the way you do. It’s weighing likely next words based on everything it’s read:
“The capital of France is ___” Paris91% a4% banana0.001%
It picks the highest-odds word, adds it to the sentence, and repeats the whole guessing game for the next word — and the next, and the next — until the reply is done. That’s the entire trick behind entire essays, poems, and code.
how it got so good at guessing
Where it learned all this
During training, the model was shown an enormous stack of text and played a simple game over and over: hide the next word, guess it, check the real answer, adjust slightly if wrong. Multiply that by hundreds of billions of guesses, and the small adjustments add up into something that’s shockingly good at “what word probably comes next” — across nearly any topic you throw at it.
Nothing in there is memorized like a textbook. It’s closer to muscle memory for language — built from repetition, not lookup.
Why it “forgets” your earlier messages
An LLM only sees what fits in its context window — think of it as short-term memory with a hard size limit. Once a conversation runs long enough, the earliest bits quietly fall out of view, which is why it might lose track of something you said a while back. It’s not being forgetful on purpose — it genuinely can’t see that far anymore.
What it's great at
Patterns it's seen a lot: writing, summarizing, explaining, brainstorming, code it's seen similar versions of before.
This entire article is a "seen this pattern before" job.
Where it slips up
When no good pattern exists, it still has to guess something — confidently. That's a "hallucination": a fluent, wrong answer with zero warning label.
Always worth double-checking facts, dates, and citations.
where you've already met one
LLMs you’ve probably already used
- Chatbots — ChatGPT, Claude, Gemini — the obvious one
- Autocomplete — Your keyboard, email, and search bar finishing your thought
- Writing tools — Grammar fixes, tone rewrites, “make this shorter” buttons
- Customer support — Those chat widgets answering your questions at 2am
- Coding help — Autocomplete for code, explaining errors, writing boilerplate
- Translation — Reading the “shape” of one language and rewriting it in another