API Basics
Every LLM app is a loop: build messages, call, stream, count, retry. Master the loop — skip the pain.
▶ Watch this reelWhat you'll learn
- Chat API anatomy
- Streaming responses
- Token counting & cost
- Rate limits & retries
- Official SDKs
- Batch API
Remember this
- The API is stateless — YOUR message list is the conversation; system sets behavior
- Stream for UX; count tokens before the call; log usage per feature
- Backoff+jitter on 429s; official SDKs; batch API halves bulk costs
Chat API anatomy
- Stateless:
messages=[...]in → one assistant message out. Your list = the conversation. - Roles:
system(operating manual) ·user·assistant(prior replies, appended by you).
Streaming
stream=True→ iterate chunks;delta.content= new text.- Time-to-first-token is the UX metric.
stream_options={"include_usage": True}→ token counts in the final chunk.
Tokens & cost
- cost = (prompt × input price) + (completion × output price — usually higher).
- Count before the call (tiktoken), trim context, cap
max_tokens, log per feature.
Rate limits & retries
- 429 = too many requests → exponential backoff + jitter, honor
Retry-After, cap attempts. - Add timeouts + circuit breakers — this is network code.
SDKs
- OpenAI SDK covers OpenAI + Azure OpenAI; Anthropic SDK mirrors the shape.
- Abstract behind your own interface only when multi-provider is real.
Batch API
- JSONL of requests → results ≤24h at ~half price.
- Route by deadline: interactive for users, batch for bulk.
Code: The production call loop — streaming, usage, retry
import random, time
from openai import OpenAI, RateLimitError
client = OpenAI()
def chat(messages, model="gpt-4o-mini", max_retries=4, **kw):
"""Call with retry/backoff + usage logging. The only wrapper you need."""
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model=model, messages=messages,
stream_options={"include_usage": True},
**kw,
)
except RateLimitError as e:
if attempt == max_retries - 1:
raise
wait = (2 ** attempt) + random.random() # backoff + jitter
ra = e.response.headers.get("Retry-After")
if ra: wait = float(ra)
print(f"429 — retry {attempt + 1} in {wait:.1f}s")
time.sleep(wait)
# --- streaming render ------------------------------------------------
stream = chat(
[{"role": "user", "content": "Explain chunking in 2 sentences"}],
stream=True, max_tokens=120,
)
full = ""
for chunk in stream:
piece = chunk.choices[0].delta.content or ""
print(piece, end="", flush=True) # render live
full += piece
if chunk.usage:
print(f"\n[usage] {chunk.usage}") # tokens for cost logging
# --- batch for bulk (sketch) -----------------------------------------
# requests.jsonl: {"custom_id":"row-1","method":"POST","url":"/v1/chat/completions","body":{...}}
# batch = client.batches.create(input_file_id=uploaded.id, endpoint="/v1/chat/completions", completion_window="24h")