How LLMs Work
You don't need a PhD. You need six intuitions. That's genuinely all it takes to use LLMs well.
▶ Watch this reelWhat you'll learn
- What is GenAI
- Transformer & attention intuition
- Tokens & tokenization
- Context window
- Sampling parameters
- Training stages
Remember this
- Generative = predictive AI with a huge token vocabulary — steer a prediction engine
- Tokens are the meter; context window is the desk — plan its budget, mind lost-in-the-middle
- temperature 0 = deterministic; training = pretrain → fine-tune → align
Generative vs predictive
- Predictive: input → label/score (fraud? 94%).
- Generative: input → new text, one token at a time.
Architect's take: it's a prediction engine with a ~100k-token output alphabet — steering (schemas, sampling) works because of that.
Transformer & attention
- 2017: Attention Is All You Need.
- Self-attention: every token weighs every other token ('it' ← 'animal').
- Nothing is programmed; grammar/facts emerge from next-token prediction at scale.
Tokens
- ~4 chars each (~¾ word). Prompt + output are billed per token.
- Generation = probability list → sample → append → repeat.
Context window
- Everything (system + history + docs + question) shares one token budget.
- Lost in the middle: attention fades mid-context → key instructions first AND last.
Sampling
| Param | Effect |
|---|---|
| temperature=0 | deterministic (extraction, code) |
| ~0.7–1 | varied phrasing (drafting) |
| top_p | trims low-probability tail |
| max_tokens | hard cost/length ceiling |
Training stages
1. Pretraining — web-scale next-token prediction (the expensive one). 2. Fine-tuning — curated datasets for domains/formats/tools. 3. Alignment (RLHF) — preference ranking → helpful/safe behavior.
Code: Your first LLM call — with eyes open
from openai import OpenAI
client = OpenAI()
# The three dials from Chapter 5 — set them ON PURPOSE
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You extract keywords. Reply with JSON only."},
{"role": "user", "content": "Extract keywords from: 'RAG pipelines need chunking'"},
],
temperature=0, # deterministic: same input → same JSON
max_tokens=60, # cost ceiling
top_p=0.9, # trim the nonsense tail
)
print(resp.usage) # prompt_tokens / completion_tokens — THE METER
print(resp.choices[0].message.content)
# Architect's take: log usage on every call in production.
# Cost = (prompt_tokens × input price) + (completion_tokens × output price),
# both quoted per 1M tokens — and output is usually 3-5x pricier.