asyncio
One LLM call takes 2 seconds. Fifty in a row take 100. Fifty concurrently take 4. asyncio is the single highest-leverage Python skill for GenAI apps.
▶ Watch this reelWhat you'll learn
- async / await
- The event loop
- gather & semaphores
- Timeouts & cancellation
Remember this
- async def returns a coroutine, nothing runs until awaited; the event loop juggles thousands of I/O waits on one thread
- gather runs calls concurrently (results stay ordered); Semaphore caps in-flight calls to respect rate limits
- Ship reliability: per-call timeouts, retry with backoff, and a budget on the whole batch
async / await
async def→ calling returns a coroutine object; nothing runs until awaited.asyncio.run(main())— entry point, called once.- Mental model: C# syntax, single-threaded cooperative runtime.
Event loop
- One thread; queue of ready coroutines; runs each to its next
await. - While a coroutine waits on I/O, others run — waiting becomes doing.
- CPU-heavy work inside a coroutine blocks the whole loop (→ PY-11).
- Beware sync/blocking calls inside async code — freezes the loop; use httpx (PY-12).
gather & Semaphore
asyncio.gather(*coros)— concurrent execution, ordered results.Semaphore(N)caps in-flight calls — rate-limit politeness.return_exceptions=True— batch survives individual failures.
Timeouts & cancellation
asyncio.wait_for(coro, timeout)→ CancelledError at the await point.- Retry transient (429/5xx) with exponential backoff; fail fast on 4xx.
- Budget the whole batch — per-call timeouts × retries ≠ bounded runtime.
Code: Production batch pattern: gather + semaphore + timeout + retry
import asyncio
sem = asyncio.Semaphore(10)
async def call_llm(prompt: str) -> str:
async with sem:
for attempt in range(3):
try:
resp = await asyncio.wait_for(
client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
),
timeout=30, # per-call deadline
)
return resp.choices[0].message.content
except (asyncio.TimeoutError, RateLimitError):
if attempt == 2:
raise # permanent failure after 3 tries
await asyncio.sleep(2 ** attempt) # 1s, 2s backoff
async def main(prompts: list[str]) -> list[str]:
# tolerate individual failures; don't lose 49 for 1
return await asyncio.gather(
*[call_llm(p) for p in prompts],
return_exceptions=True,
)
results = asyncio.run(asyncio.wait_for(main(prompts), timeout=600))
# failures arrive as exception objects in-place — inspect & handle