HTTP Clients
Every LLM call you make travels over HTTP. Auth headers, retries on 429s, pagination, streaming — the client layer decides whether your app is robust or fragile.
▶ Watch this reelWhat you'll learn
- requests vs httpx
- Auth & headers
- Retries & backoff
- Pagination & streaming
Remember this
- httpx for everything new: one API for sync and async, HTTP/2, enforced timeouts — requests blocks the event loop
- Auth dialects: x-api-key, Bearer tokens, full OAuth for user-delegation; secrets from env, never code
- Retry only transient statuses with exponential backoff + jitter; paginate via opaque cursors; stream LLM tokens with client.stream for instant UX
requests vs httpx
- requests: sync-only → blocks event loops. Legacy scripts only.
- httpx: same API, sync + async, HTTP/2, strict default timeouts.
- Shared client instance: connection pooling; never client-per-request.
Auth
- Dialects:
x-api-keyheader ·Authorization: Bearer· OAuth2 (user-delegation flows). - Service-to-service: env-var key/bearer. Never hardcode; secrets manager in prod.
raise_for_status()everywhere — silent failures are worse than crashes.
Retries
- Transient only: 429, 500/502/503. 4xx fails fast.
- Exponential backoff + jitter; respect Retry-After; cap attempts; tenacity.
Pagination & streaming
- Cursor-based (opaque next_cursor) > offset for changing data.
- SSE via
client.stream+aiter_lines(); parsedata:prefix; blank lines are heartbeats. - Streaming = sub-second time-to-first-token UX.
Code: The resilient LLM client, assembled
import os, httpx, tenacity
client = httpx.AsyncClient(
base_url="https://api.llm-provider.com/v1",
headers={"Authorization": f"Bearer {os.environ['LLM_KEY']}"},
timeout=httpx.Timeout(60.0, connect=5.0),
limits=httpx.Limits(max_connections=100), # pool cap
)
def transient(e: BaseException) -> bool:
return isinstance(e, httpx.HTTPStatusError) \
and e.response.status_code in (429, 500, 502, 503)
@tenacity.retry(retry=tenacity.retry_if_exception(transient),
wait=tenacity.wait_exponential(1, max=30)
+ tenacity.wait_random(0, 1),
stop=tenacity.stop_after_attempt(5), reraise=True)
async def complete(prompt: str) -> str:
r = await client.post("/chat", json={"prompt": prompt})
r.raise_for_status()
return r.json()["text"]
async def complete_stream(prompt: str):
async with client.stream("POST", "/chat",
json={"prompt": prompt}) as r:
async for line in r.aiter_lines():
if line.startswith("data: ") and line != "data: [DONE]":
yield json.loads(line[6:])["token"]
# One client, two personalities: buffered complete + live stream