NumPy & Vectors
Embeddings are vectors. Vectors are NumPy. This reel is the math your AI career runs on.
▶ Watch this reelWhat you'll learn
- NumPy arrays
- Dot product & norm
- Cosine similarity
- Broadcasting
Remember this
- ndarray = typed, packed, vectorized — 50-100x faster than lists; embeddings live here
- Cosine similarity = dot ÷ norms; 1 = same meaning — the retrieval computation in one line
- Normalize to unit vectors, then query @ corpus.T = all similarities at once
Arrays
ndarray: one dtype, contiguous memory, vectorized C ops → 50-100× lists.a.shape,a.dtype—(1536,)float32 = one embedding.- 2D: rows = documents;
a[0]row ·a[:, 0]column.
Dot product & norm
- Dot:
np.dot(a, b)= Σ aᵢbᵢ — agreement measure. - Norm:
np.linalg.norm(a)= √(Σ aᵢ²) — vector length.
Cosine similarity
cos(θ) = dot(a, b) / (norm(a) · norm(b))
- 1.0 = same direction · 0 = unrelated · −1 = opposite.
- Similarity = angle (direction), not distance.
Broadcasting
- Align dims right-to-left; must match or be 1.
- Pattern: normalize all vectors →
sims = corpus_u @ query_u(one matmul = all cosine scores). - Vector DBs optimize exactly this — you now know what they do.
Code: Hand-rolled similarity search — 15 lines
import numpy as np
# Simulate: 10,000 chunks, 1,536-dim embeddings (like OpenAI's)
rng = np.random.default_rng(42)
corpus = rng.normal(size=(10_000, 1536)).astype(np.float32)
query = rng.normal(size=(1536,)).astype(np.float32)
# --- 1. normalize to unit length (norm = 1) --------------------
def unit(v: np.ndarray) -> np.ndarray:
return v / np.linalg.norm(v, axis=-1, keepdims=True)
corpus_u = unit(corpus) # (10000, 1536), norms now 1
query_u = unit(query) # (1536,)
# --- 2. ALL cosine similarities in ONE matmul -------------------
sims = corpus_u @ query_u # (10000,) — dot = cosine (unit vectors)
# --- 3. top-5 most similar chunks -------------------------------
top5 = np.argsort(sims)[::-1][:5]
for rank, idx in enumerate(top5, 1):
print(f"#{rank} chunk={idx} sim={sims[idx]:.4f}")
# This 3-step pattern is the beating heart of every vector DB.
# FAISS/HNSW just make step 2 faster — the math is identical.