Embedding Concepts

Meaning, as math. Five concepts and you'll understand the technology half of Stage 4.

▶ Watch this reel

What you'll learn

  1. What is an embedding
  2. Dimensions & trade-offs
  3. Similarity metrics
  4. Semantic vs keyword search
  5. Visualizing embeddings

Remember this

What is an embedding

Dimensions

Similarity metrics

MetricWhen
Cosinetext default (direction only)
Dotnormalized vectors → == cosine, faster
Euclideanwhen magnitude carries signal (some image/recsys)

Semantic vs keyword

Visualizing

Code: Embeddings, from text to plot

import numpy as np
from openai import OpenAI
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt

client = OpenAI()

def embed(texts: list[str]) -> np.ndarray:
    r = client.embeddings.create(model="text-embedding-3-small", input=texts)
    return np.array([d.embedding for d in r.data])

docs = [
    "Quarterly revenue rose 12% to $4.1M",
    "Q1 earnings: total income $4.1 million, up 12%",
    "The sofa arrives in grey fabric on Friday",
    "Error AX-2041: printer queue timeout",
    "Policy updated: remote work now 3 days per week",
    "नई नीति: दूरस्थ कार्य अब सप्ताह में 3 दिन",
]
vecs = embed(docs)                     # (6, 1536)
print(vecs.shape)

# similarity sanity check (unit-normalized → dot = cosine)
u = vecs / np.linalg.norm(vecs, axis=1, keepdims=True)
sim = u @ u.T
print("revenue pair:", round(sim[0, 1], 3))   # high — paraphrase
print("revenue vs sofa:", round(sim[0, 2], 3)) # near 0 — unrelated

# 1536 → 2 for a look
xy = PCA(n_components=2).fit_transform(vecs)
plt.scatter(xy[:, 0], xy[:, 1])
for (x, y), d in zip(xy, docs):
    plt.annotate(d[:20], (x, y))
plt.show()                              # clusters = healthy corpus