Model Landscape

Hundreds of models. Five questions to pick the right one — every single time.

▶ Watch this reel

What you'll learn

  1. Model families
  2. Reasoning vs standard
  3. Multimodal models
  4. Selection criteria
  5. Hallucination & limits

Remember this

Three families

FamilyExamplesStrengthCost shape
Frontier APIGPT/Claude/Gemini flagshipsMax capabilityper-token
Open-weightLlama, Mistral, Qwenprivacy, fine-tune, self-runyour GPU
Small-mini / -nanocheap, fast, bulk workper-token, tiny
Hybrid (frontier for hard + small for bulk) is a legitimate architecture.

Reasoning vs standard

Multimodal

Selection: quality · cost · latency · privacy

Hallucination — a structural property

Code: A pragmatic model-selection harness

import time, json
from openai import OpenAI

client = OpenAI()

# Your real examples — 50 is enough to rank candidates
EVAL_SET = [
    {"input": "Extract the total from: 'Invoice total: $1,240.00'", "expect": "1240.00"},
    {"input": "Is this refund request abusive? 'give me my money back now!!!'", "expect": "no"},
]

CANDIDATES = ["gpt-4o-mini", "gpt-4o"]   # + your open-weight endpoint

def run(model: str, prompt: str) -> str:
    t0 = time.time()
    r = client.chat.completions.create(
        model=model, temperature=0, max_tokens=80,
        messages=[{"role": "user", "content": prompt}],
    )
    ms = (time.time() - t0) * 1000
    usage = r.usage
    return r.choices[0].message.content, ms, usage

for model in CANDIDATES:
    correct, cost, lat = 0, 0.0, []
    for ex in EVAL_SET:
        out, ms, u = run(model, ex["input"])
        correct += ex["expect"] in out
        lat.append(ms)
        # cost += tokens × model price  (fill from the current pricing page)
    print(f"{model}: {correct}/{len(EVAL_SET)}  "
          f"avg {sum(lat)/len(lat):.0f}ms")

# Architect's take: keep this harness. Re-run it whenever you
# consider switching models — decisions become data, not debates.