Polars: pandas at Warp Speed
Pandas made Python the language of data — and then showed its age. Polars is the rewrite: multi-threaded, query-planned, and 10-50x faster on the workloads agents actually generate.
▶ Watch this reelWhat you'll learn
- Why Polars wins
- The expression API
- Lazy queries
- When to switch
Remember this
- Polars = Rust + Arrow + multi-threading: columnar vectorized ops, strict dtypes, no index — 10-50x faster than pandas on real workloads
- Expressions (pl.col) replace loops and apply: compose when/then/otherwise, string ops, and JSON-into-structs as data
- Lazy evaluation builds a plan, then pushes filters and projections into the file read — files bigger than RAM become routine
- Switch for new and large pipelines; keep pandas where the ecosystem demands it; bridge with to_pandas() at the boundary
Why it wins
- Rust + Arrow: columnar, vectorized, strict dtypes, null bitmaps.
- Multi-threaded by default — no GIL workarounds.
- No index — no alignment surprises.
Expressions
- pl.col composes: arithmetic, when/then/otherwise, str.*, struct.field.
- JSON → struct → columns in two moves.
- .apply() → native expressions. That's the migration.
Lazy
- scan_* builds a plan; collect() executes.
- Predicate + projection pushdown at the file layer; sink_parquet streams.
- explain() shows the optimized plan — read it.
When to switch
- Switch: new pipelines, >1M rows, agent logs/evals/JSONL.
- Stay: ecosystem-locked tooling, large legacy codebase.
- Hybrid: to_pandas()/from_pandas at the boundary.
Cross-links
AG-19 (lazy → collect_async in async services), AG-21 (eval datasets), AG-23 (token accounting feeds cost model), GA-20 (SLO analysis), GA-25/30 (cost pipelines).
Code: The agent-data one-liner set
import polars as pl
# 80% of GenAI data work is these four:
logs = pl.scan_ndjson("runs/*.jsonl")
costs = logs.group_by("model").agg(
(pl.col("in") + pl.col("out")).sum().alias("tokens")
) # 1 · token accounting (AG-23)
quality = logs.with_columns(
pl.when(pl.col("score") >= 0.92).then("pass")
.otherwise("fail").alias("v")
).group_by(["model", "v"]).agg(pl.len()) # 2 · SLO verdicts (GA-20)
slow = logs.filter(pl.col("ms") > 3000) # 3 · latency outliers (GA-07)
query = logs.select(pl.col("prompt").str.extract(r"Goal: (.*?)")
.alias("goal")).unique() # 4 · prompt mining
# .collect() when ready. No loops were harmed.