Threads vs Processes
asyncio, threads, processes — three ways to do "at the same time", each right for a different enemy: waiting, latency, and the GIL.
▶ Watch this reelWhat you'll learn
- The GIL
- Threads for I/O
- Processes for CPU
- Choosing & executors
Remember this
- GIL: one thread executes Python bytecode at a time — threads never parallelize pure-Python CPU work, but the lock releases on I/O and C extensions
- Threads: shared memory, good for I/O waits and wrapping blocking libraries; still need locks for shared mutable state
- Processes: separate interpreters = true multi-core for CPU-bound work; pay in memory duplication and pickling overhead — chunk tasks, not data
The GIL
- Global Interpreter Lock: one thread runs Python bytecode at a time.
- Released during I/O and C-extension calls (NumPy, crypto, DB drivers).
- Threads ≠ CPU parallelism in Python; processes = own GIL per interpreter.
Threads — I/O-bound
- Overlap waits; simple; GIL released during blocking I/O.
- Use to wrap blocking libs inside async apps (
run_in_executor). - Shared mutable state still needs
threading.Lock.
Processes — CPU-bound
- True multi-core; costs: RAM per worker, pickling IPC.
ProcessPoolExecutor.map(fn, items, chunksize=64)— batch tasks.- Mapped fn must be module-level (picklable by reference).
Decision
- Bottleneck = waiting → asyncio; = blocking lib → threads; = computing → processes.
- Hybrid is normal: async app + thread/process adapters per pipeline stage.
Code: The hybrid reality: async app + thread/process adapters
import asyncio
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
# blocking sync SDK you CANNOT change
def legacy_translate(text: str) -> str:
return old_sdk.translate(text) # blocks, no async version
def cpu_heavy_score(text: str) -> float:
return sum(ord(c) for c in text) % 100 / 100 # pretend: pure-Python math
async def pipeline(texts: list[str]) -> list[str]:
loop = asyncio.get_running_loop()
with ThreadPoolExecutor(8) as tpool, ProcessPoolExecutor(4) as ppool:
# 1 · blocking SDK off the event loop (threads)
translations = await loop.run_in_executor(
tpool, lambda: list(map(legacy_translate, texts)))
# 2 · CPU math on real cores (processes)
scores = await loop.run_in_executor(
ppool, lambda: list(ppool.map(cpu_heavy_score, texts)))
# 3 · back on the loop: async I/O continues
return await asyncio.gather(
*[notify_api(t, s) for t, s in zip(translations, scores)])