ML Basics

You won't train models — but you'll evaluate them, daily. Five concepts make that possible.

▶ Watch this reel

What you'll learn

  1. Supervised vs unsupervised
  2. Train/test split & overfitting
  3. Metrics: accuracy, precision, recall, F1
  4. scikit-learn basics
  5. Neural network intuition

Remember this

Supervised vs unsupervised

Your role: mostly evaluate and consume models.

Train/test split & overfitting

Metrics

MetricAnswers
Accuracyoverall correct — misleading with rare classes
Precisionof flagged, how many right? (trust)
Recallof real ones, how many caught? (coverage)
F1balance of both

scikit-learn

Neural network intuition

Code: Eval vocabulary + a real baseline, runnable

from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.metrics import precision_score, recall_score, f1_score

# --- the honest split ------------------------------------------
(train, test) = (
    fetch_20newsgroups(subset="train", categories=["sci.space", "rec.sport.baseball"]),
    fetch_20newsgroups(subset="test",  categories=["sci.space", "rec.sport.baseball"]),
)

# --- baseline before any LLM ------------------------------------
pipe = Pipeline([
    ("tfidf", TfidfVectorizer()),
    ("clf",   LogisticRegression(max_iter=1000)),
])
pipe.fit(train.data, train.target)
pred = pipe.predict(test.data)

# --- the metrics that matter ------------------------------------
print(f"precision {precision_score(test.target, pred):.3f}")
print(f"recall    {recall_score(test.target, pred):.3f}")
print(f"f1        {f1_score(test.target, pred):.3f}")

# 2 categories, ~1,900 docs — trains in seconds.
# When THIS baseline fails your eval set, an LLM is justified.
# Until then, you ship this.