NumPy & Vectors

Embeddings are vectors. Vectors are NumPy. This reel is the math your AI career runs on.

▶ Watch this reel

What you'll learn

  1. NumPy arrays
  2. Dot product & norm
  3. Cosine similarity
  4. Broadcasting

Remember this

Arrays

Dot product & norm

Cosine similarity

cos(θ) = dot(a, b) / (norm(a) · norm(b))

Broadcasting

Code: Hand-rolled similarity search — 15 lines

import numpy as np

# Simulate: 10,000 chunks, 1,536-dim embeddings (like OpenAI's)
rng = np.random.default_rng(42)
corpus = rng.normal(size=(10_000, 1536)).astype(np.float32)
query  = rng.normal(size=(1536,)).astype(np.float32)

# --- 1. normalize to unit length (norm = 1) --------------------
def unit(v: np.ndarray) -> np.ndarray:
    return v / np.linalg.norm(v, axis=-1, keepdims=True)

corpus_u = unit(corpus)          # (10000, 1536), norms now 1
query_u  = unit(query)           # (1536,)

# --- 2. ALL cosine similarities in ONE matmul -------------------
sims = corpus_u @ query_u        # (10000,) — dot = cosine (unit vectors)

# --- 3. top-5 most similar chunks -------------------------------
top5 = np.argsort(sims)[::-1][:5]
for rank, idx in enumerate(top5, 1):
    print(f"#{rank} chunk={idx}  sim={sims[idx]:.4f}")

# This 3-step pattern is the beating heart of every vector DB.
# FAISS/HNSW just make step 2 faster — the math is identical.