ML Basics
You won't train models — but you'll evaluate them, daily. Five concepts make that possible.
▶ Watch this reelWhat you'll learn
- Supervised vs unsupervised
- Train/test split & overfitting
- Metrics: accuracy, precision, recall, F1
- scikit-learn basics
- Neural network intuition
Remember this
- Supervised = labeled prediction; unsupervised = structure discovery — you'll mostly evaluate, not train
- Overfitting = memorized noise; sacred test set, tune on dev — same law as LLM evals
- Precision = trust, recall = coverage; sklearn's fit/predict/score builds cheap baselines before LLMs
Supervised vs unsupervised
- Supervised: labeled data → classification / regression.
- Unsupervised: clustering, embeddings, topics.
Your role: mostly evaluate and consume models.
Train/test split & overfitting
- Overfitting = memorized noise → fails on new data.
- Train / dev (tuning) / test (report once, sacred).
- LLM echo: benchmark contamination = same disease.
Metrics
| Metric | Answers |
|---|---|
| Accuracy | overall correct — misleading with rare classes |
| Precision | of flagged, how many right? (trust) |
| Recall | of real ones, how many caught? (coverage) |
| F1 | balance of both |
- Optimize by business cost: false alarm vs miss.
scikit-learn
- Universal API:
fit/predict/score. - Baseline first: tf-idf + logistic regression ≈ strong text classifier.
Pipeline(vectorizer, model)— no preprocessing leakage.
Neural network intuition
- Layers of weighted sums + nonlinearities.
- Loop: predict → loss → backprop nudges weights downhill → repeat × millions.
- Transformers are this loop at trillion-token scale.
Code: Eval vocabulary + a real baseline, runnable
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.metrics import precision_score, recall_score, f1_score
# --- the honest split ------------------------------------------
(train, test) = (
fetch_20newsgroups(subset="train", categories=["sci.space", "rec.sport.baseball"]),
fetch_20newsgroups(subset="test", categories=["sci.space", "rec.sport.baseball"]),
)
# --- baseline before any LLM ------------------------------------
pipe = Pipeline([
("tfidf", TfidfVectorizer()),
("clf", LogisticRegression(max_iter=1000)),
])
pipe.fit(train.data, train.target)
pred = pipe.predict(test.data)
# --- the metrics that matter ------------------------------------
print(f"precision {precision_score(test.target, pred):.3f}")
print(f"recall {recall_score(test.target, pred):.3f}")
print(f"f1 {f1_score(test.target, pred):.3f}")
# 2 categories, ~1,900 docs — trains in seconds.
# When THIS baseline fails your eval set, an LLM is justified.
# Until then, you ship this.