Model Landscape
Hundreds of models. Five questions to pick the right one — every single time.
▶ Watch this reelWhat you'll learn
- Model families
- Reasoning vs standard
- Multimodal models
- Selection criteria
- Hallucination & limits
Remember this
- Three families: frontier API / open-weight self-run / small cheap — hybrid architectures are normal
- Reasoning models for hard logic, standard for bulk language work; route and escalate
- Hallucination is structural — anchor (retrieval), constrain (schema), judge (humans)
Three families
| Family | Examples | Strength | Cost shape |
|---|---|---|---|
| Frontier API | GPT/Claude/Gemini flagships | Max capability | per-token |
| Open-weight | Llama, Mistral, Qwen | privacy, fine-tune, self-run | your GPU |
| Small | -mini / -nano | cheap, fast, bulk work | per-token, tiny |
Hybrid (frontier for hard + small for bulk) is a legitimate architecture.
Reasoning vs standard
- Reasoning: visible chain-of-thought → math, code, multi-step planning. Slower + pricier.
- Standard: direct answers → classification, extraction, chat.
- Pattern: route — cheap first, escalate on failure.
Multimodal
- Vision (images/screenshots/charts), PDF understanding, audio.
- Collapses point solutions (OCR, chart parsers) into one call.
Selection: quality · cost · latency · privacy
- Eval on your 50-100 examples · cost math × volume · latency budget per UX type.
- Write the decision + reason in an architecture note.
Hallucination — a structural property
- Models generate plausible text; they don't look things up.
- Defense in depth: anchor (RAG) → constrain (schema/validators) → judge (human-in-loop).
Code: A pragmatic model-selection harness
import time, json
from openai import OpenAI
client = OpenAI()
# Your real examples — 50 is enough to rank candidates
EVAL_SET = [
{"input": "Extract the total from: 'Invoice total: $1,240.00'", "expect": "1240.00"},
{"input": "Is this refund request abusive? 'give me my money back now!!!'", "expect": "no"},
]
CANDIDATES = ["gpt-4o-mini", "gpt-4o"] # + your open-weight endpoint
def run(model: str, prompt: str) -> str:
t0 = time.time()
r = client.chat.completions.create(
model=model, temperature=0, max_tokens=80,
messages=[{"role": "user", "content": prompt}],
)
ms = (time.time() - t0) * 1000
usage = r.usage
return r.choices[0].message.content, ms, usage
for model in CANDIDATES:
correct, cost, lat = 0, 0.0, []
for ex in EVAL_SET:
out, ms, u = run(model, ex["input"])
correct += ex["expect"] in out
lat.append(ms)
# cost += tokens × model price (fill from the current pricing page)
print(f"{model}: {correct}/{len(EVAL_SET)} "
f"avg {sum(lat)/len(lat):.0f}ms")
# Architect's take: keep this harness. Re-run it whenever you
# consider switching models — decisions become data, not debates.