asyncio

One LLM call takes 2 seconds. Fifty in a row take 100. Fifty concurrently take 4. asyncio is the single highest-leverage Python skill for GenAI apps.

▶ Watch this reel

What you'll learn

  1. async / await
  2. The event loop
  3. gather & semaphores
  4. Timeouts & cancellation

Remember this

async / await

Event loop

gather & Semaphore

Timeouts & cancellation

Code: Production batch pattern: gather + semaphore + timeout + retry

import asyncio

sem = asyncio.Semaphore(10)

async def call_llm(prompt: str) -> str:
    async with sem:
        for attempt in range(3):
            try:
                resp = await asyncio.wait_for(
                    client.chat.completions.create(
                        model="gpt-4o",
                        messages=[{"role": "user", "content": prompt}],
                    ),
                    timeout=30,          # per-call deadline
                )
                return resp.choices[0].message.content
            except (asyncio.TimeoutError, RateLimitError):
                if attempt == 2:
                    raise               # permanent failure after 3 tries
                await asyncio.sleep(2 ** attempt)  # 1s, 2s backoff

async def main(prompts: list[str]) -> list[str]:
    # tolerate individual failures; don't lose 49 for 1
    return await asyncio.gather(
        *[call_llm(p) for p in prompts],
        return_exceptions=True,
    )

results = asyncio.run(asyncio.wait_for(main(prompts), timeout=600))
# failures arrive as exception objects in-place — inspect & handle