How LLMs Work

You don't need a PhD. You need six intuitions. That's genuinely all it takes to use LLMs well.

▶ Watch this reel

What you'll learn

  1. What is GenAI
  2. Transformer & attention intuition
  3. Tokens & tokenization
  4. Context window
  5. Sampling parameters
  6. Training stages

Remember this

Generative vs predictive

Architect's take: it's a prediction engine with a ~100k-token output alphabet — steering (schemas, sampling) works because of that.

Transformer & attention

Tokens

Context window

Sampling

ParamEffect
temperature=0deterministic (extraction, code)
~0.7–1varied phrasing (drafting)
top_ptrims low-probability tail
max_tokenshard cost/length ceiling

Training stages

1. Pretraining — web-scale next-token prediction (the expensive one). 2. Fine-tuning — curated datasets for domains/formats/tools. 3. Alignment (RLHF) — preference ranking → helpful/safe behavior.

Code: Your first LLM call — with eyes open

from openai import OpenAI

client = OpenAI()

# The three dials from Chapter 5 — set them ON PURPOSE
resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": "You extract keywords. Reply with JSON only."},
        {"role": "user", "content": "Extract keywords from: 'RAG pipelines need chunking'"},
    ],
    temperature=0,        # deterministic: same input → same JSON
    max_tokens=60,        # cost ceiling
    top_p=0.9,            # trim the nonsense tail
)

print(resp.usage)         # prompt_tokens / completion_tokens — THE METER
print(resp.choices[0].message.content)

# Architect's take: log usage on every call in production.
# Cost = (prompt_tokens × input price) + (completion_tokens × output price),
# both quoted per 1M tokens — and output is usually 3-5x pricier.