What is an LLM, in one sentence? A model that has one job: given the text so far, predict the next token, append it, and repeat. That’s it. Everything else — chat interfaces, coding assistants, agents that call tools — is scaffolding built around that one loop running thousands of times per response.
Most explanations either stop at “it’s trained on the internet” or jump straight into transformer math you’ll never touch. Neither one helps when you’re deciding how much VRAM to buy, why a long conversation suddenly gets dumber, or why your token bill doubled overnight. This is the model you actually need.
What is an LLM?
A large language model is a neural network trained to predict the next token in a sequence. Repeat that over enough text, and it picks up grammar, facts, and reasoning patterns as a side effect of getting good at that one task. Hugging Face’s own generation docs put it plainly: an LLM “is trained to generate the next word (token) given some initial text (prompt) along with its own generated outputs up to a predefined length or when it reaches an end-of-sequence (EOS) token.”
That’s the whole mechanism. There’s no separate “understanding” step. The model has learned, from its training data, which token is statistically likely to come next given everything before it — and at large enough scale, “statistically likely” starts looking a lot like reasoning.
Size is measured in parameters — the numeric weights the network adjusts during training. A 7B model has about 7 billion of them. More parameters generally means better output, but it also means more memory to hold the weights and more compute per token generated. That trade-off is the entire reason quantized, smaller models exist.
What Is an LLM Actually Doing When It Generates Text?
One token at a time, in a loop, with no ability to revise what it already wrote. The model looks at every token so far — your prompt plus whatever it has generated in this response. Then it outputs a probability distribution over its entire vocabulary for what comes next.
Then a decoding strategy picks the actual token from that distribution. Hugging Face’s transformers library documents two defaults worth knowing:
- Greedy search — always pick the single most likely next token. This is the library’s default decoding strategy. Deterministic, and prone to repetitive output.
- Sampling — pick randomly from the distribution, weighted by probability, controlled by a
temperatureparameter. Low temperature (under 0.4) stays close to greedy; high temperature (above 0.8) gets more varied and more prone to nonsense.
That picked token gets appended to the sequence, and the whole thing runs again to pick the next one. This continues until the model emits an end-of-sequence token or hits a length limit you set. There is no lookahead and no backtracking — once a token is chosen, it’s part of the context for every token that follows, mistakes included.
This is also why LLMs hallucinate confidently instead of saying “I don’t know.” The model isn’t checking facts; it’s continuing a pattern. If the highest-probability continuation of a plausible-sounding sentence is a wrong fact, the model produces the wrong fact with the same fluency as a right one.
What is a token, and why does the context window matter?
A token is the model’s actual unit of work — usually a word, part of a word, or a punctuation mark, not a full sentence. “Deployment” might be one token; “unconfigured” often splits into two or three. Everything you feed the model and everything it generates gets converted to and from tokens.
The context window is the maximum number of tokens the model can hold in memory at once. That includes your system prompt, the conversation history, and the response it’s generating — all counted together. Run out of room, and the oldest tokens get dropped, or the request fails outright, depending on the client.
⚠️ Note: this is a hard ceiling, not a soft one. A model with a 4,096-token context window doesn’t get slower as you approach the limit. It silently loses the beginning of the conversation, or errors out, depending on the client. If your chatbot “forgets” something you said 20 messages ago, this is why.
Context windows vary by model and by how you run it — and Ollama’s own docs aren’t even fully consistent on the default. Its FAQ quotes a flat “4096 tokens,” but the newer context-length docs say the real default scales with your GPU: 4k below 24 GiB VRAM, 32k from 24–48 GiB, 256k at 48 GiB and up. Either way, don’t assume — check num_ctx or ollama show <model> before relying on a specific window size. Every extra token of context also costs GPU memory for the KV cache — more on how much RAM and VRAM that adds per model size below.
How to measure your LLM’s token throughput yourself
Don’t take a vendor’s speed claims on faith — Ollama exposes the real numbers on every request. Run any local model with the verbose flag:
ollama run llama3.1 --verbose "Explain quicksort in one paragraph"
The output includes a full timing breakdown:
total duration: 4.21s
load duration: 0.31s
prompt eval count: 14 token(s)
prompt eval duration: 0.18s
prompt eval rate: 77.78 tokens/s
eval count: 142 token(s)
eval duration: 3.72s
eval rate: 38.17 tokens/s
prompt eval rate tells you how fast the model reads your input. eval rate tells you how fast it generates a response — the number that actually determines whether a chat feels instant or sluggish. Ollama’s API returns the same data as raw fields on every call: prompt_eval_count, eval_count, prompt_eval_duration, and eval_duration, all in nanoseconds. Calculate tokens per second yourself with eval_count / eval_duration * 10^9. That’s useful when you’re comparing two GPUs, or two quantization levels, and want a number instead of a feeling.
For programmatic access to these fields (building a dashboard, logging latency, whatever), the Ollama API guide covers the request format for /api/generate and /api/chat in full.
Why model size and quantization determine your hardware
Bigger models are not automatically better for your use case — they’re just bigger. A 70B model outperforms a 7B model on complex reasoning. But it also needs roughly 8x the memory, and runs a fraction of the tokens per second on the same GPU. For a lot of production tasks — classification, extraction, simple chat — a well-chosen 7B or 12B model gets you most of the quality at a fraction of the latency and cost. That “most” varies by task and isn’t a benchmarked number here — run your own comparison before committing to a size.
Quantization is how you shrink that memory footprint without retraining anything. It reduces the precision of each weight, typically from 16-bit floating point down to 4-bit integers. That cuts a model’s size by roughly 4x, with a modest quality loss. This is why an 8B model tagged Q4_K_M downloads at around 5 GB instead of the 16+ GB its full-precision weights would take.
The practical upshot: your hardware budget and your context-window budget are the same budget. A bigger context window means a bigger KV cache means less headroom for a bigger model. The full breakdown by model size — RAM, VRAM, and which GPU actually handles which tier — is in the Ollama hardware requirements guide; it’s worth reading before you buy anything.
What Is an LLM Bad At? Common Mistakes Engineers Make
Treating a bigger model as always the right call. Say your task is narrow — summarizing tickets, extracting fields from a form. A smaller model tuned for that job will often match a frontier model’s accuracy at a tenth of the latency and cost. Benchmark on your actual task before defaulting to the biggest model available. See how current frontier models actually stack up against each other before assuming bigger wins.
Not budgeting for context growth. A tool-calling agent reads a file, calls an API, and reasons over the result. That can burn through thousands of tokens of context before it produces a single user-facing answer. If you’re wiring an LLM into an MCP server for tool access, every tool description and every tool response counts against that same context window.
Assuming greedy decoding is always the safe choice. It’s deterministic, which is nice for testing. But it also makes models more repetitive, and more likely to loop on themselves in longer generations. Most production chat use cases want some sampling temperature, not zero.
Skipping the auth and permission model on tool-calling setups. An LLM that can call tools is only as safe as what those tools are allowed to do. That gap bit real deployments when the MCP spec changed how servers authenticate tool calls and older integrations broke silently.
Frequently Asked Questions
Q: What is an LLM in simple terms?
A: A large language model is software trained to predict the next word in a piece of text, one word at a time, based on everything written before it. Run that prediction in a loop and you get full sentences, code, and conversations.
Q: Is an LLM the same thing as AI?
A: No. An LLM is one specific type of AI model, built for generating and understanding text. AI is the broader field that also includes image models, recommendation systems, robotics, and other approaches that have nothing to do with predicting tokens.
Q: What happens when an LLM runs out of context window?
A: The oldest tokens in the conversation get dropped, or the request fails, depending on how the client handles it. The model doesn’t slow down or warn you — it just stops “remembering” the earlier parts of the exchange.
Q: Does a bigger LLM always give better answers?
A: Not for every task. Bigger models generally handle complex reasoning better, but for narrow, well-defined tasks a smaller model is often just as accurate and considerably faster and cheaper to run.
Q: Can I run an LLM without a GPU?
A: Yes, on CPU alone, though generation is much slower — often a few tokens per second instead of tens or hundreds. It’s usable for scripts and batch jobs, less so for interactive chat.
Quick Summary — what is an LLM, in practice:
– An LLM predicts one next token at a time based on everything before it, then repeats — there’s no separate “understanding” step
– Greedy decoding always picks the most likely token; sampling (controlled by temperature) introduces variation and is usually the better default for chat
– A token is the model’s real unit of text; the context window is the hard cap on how many tokens (prompt + response) it can hold at once — Ollama’s default scales with VRAM (4k/32k/256k tiers), not one fixed number
– Quantization cuts model size roughly 4x by lowering weight precision, which is why an 8B Q4_K_M model downloads at ~5 GB instead of 16+ GB
– Measure real throughput with ollama run <model> --verbose or the eval_count/eval_duration fields from the API — don’t guess
Run ollama run <model> --verbose on whatever model you already have installed and look at the eval rate line before you assume you need bigger hardware — you might already have enough.