Embedding Concepts
Meaning, as math. Five concepts and you'll understand the technology half of Stage 4.
▶ Watch this reelWhat you'll learn
- What is an embedding
- Dimensions & trade-offs
- Similarity metrics
- Semantic vs keyword search
- Visualizing embeddings
Remember this
- Embeddings = meaning as vectors; similar direction = similar meaning
- Dimensions trade quality for memory; cosine is the default metric (dot if normalized)
- Semantic search finds meaning, keyword finds literals — visualize with PCA/t-SNE to debug your corpus
What is an embedding
- Text → fixed-length vector of floats; similar direction = similar meaning.
- Meaning becomes measurable geometry.
Dimensions
- 384 / 768 / 1536 / 3072… richer ≠ always better.
- Memory: 1M docs × 1536 × 4B ≈ 6 GB (+index). Pick the smallest size that passes your eval.
- Dimensions are model-locked → model change = full re-embedding (GA-08).
Similarity metrics
| Metric | When |
|---|---|
| Cosine | text default (direction only) |
| Dot | normalized vectors → == cosine, faster |
| Euclidean | when magnitude carries signal (some image/recsys) |
- Normalize at write time → query with cheap dot product.
Semantic vs keyword
- Keyword: exact tokens — precise on codes/IDs, deaf to synonyms.
- Semantic: meaning — synonym-robust, blurry on exact strings.
- Production answer: hybrid (GA-15).
Visualizing
- PCA: fast linear projection — first look.
- t-SNE: cluster-preserving, nonlinear — pretty, distorts global distances.
- Healthy corpus → topic clusters + cross-lingual neighbors. Blob → fix chunking/model upstream.
Code: Embeddings, from text to plot
import numpy as np
from openai import OpenAI
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
client = OpenAI()
def embed(texts: list[str]) -> np.ndarray:
r = client.embeddings.create(model="text-embedding-3-small", input=texts)
return np.array([d.embedding for d in r.data])
docs = [
"Quarterly revenue rose 12% to $4.1M",
"Q1 earnings: total income $4.1 million, up 12%",
"The sofa arrives in grey fabric on Friday",
"Error AX-2041: printer queue timeout",
"Policy updated: remote work now 3 days per week",
"नई नीति: दूरस्थ कार्य अब सप्ताह में 3 दिन",
]
vecs = embed(docs) # (6, 1536)
print(vecs.shape)
# similarity sanity check (unit-normalized → dot = cosine)
u = vecs / np.linalg.norm(vecs, axis=1, keepdims=True)
sim = u @ u.T
print("revenue pair:", round(sim[0, 1], 3)) # high — paraphrase
print("revenue vs sofa:", round(sim[0, 2], 3)) # near 0 — unrelated
# 1536 → 2 for a look
xy = PCA(n_components=2).fit_transform(vecs)
plt.scatter(xy[:, 0], xy[:, 1])
for (x, y), d in zip(xy, docs):
plt.annotate(d[:20], (x, y))
plt.show() # clusters = healthy corpus