File Formats

Formats are contracts. CSV lies, JSON flexes, Parquet tells the truth — choose per workload.

▶ Watch this reel

What you'll learn

  1. CSV pitfalls
  2. JSON & JSONL
  3. Parquet
  4. Parquet partitioning
  5. Arrow / PyArrow
  6. Parquet vs DB vs Vector

Remember this

CSV pitfalls

JSON & JSONL

Parquet

Partitioning

Arrow / PyArrow

Parquet vs DB vs Vector

JobTool
Mutating transactional truthPostgres / SQL Server
Bulk immutable analyticsParquet (+ DuckDB)
Semantic similarity at scaleVector store (Stage 4)

Code: Formats in practice — CSV in, Parquet out, partitioned

import pandas as pd
import pyarrow as pa
import pyarrow.parquet as pq

# --- CSV: parse explicitly (no silent inference in prod) ----------
dtypes = {"amount": "float64", "status": "category"}
df = pd.read_csv("events.csv", encoding="utf-8", dtype=dtypes,
                 parse_dates=["created"])

# --- Parquet: typed + compressed -----------------------------------
df.to_parquet("events.parquet", index=False, compression="snappy")
# Typical result: 10x smaller than CSV, schema embedded, 10-50x faster reads

# --- Partitioned dataset: folders as an index -----------------------
table = pa.Table.from_pandas(df)
pq.write_to_dataset(
    table,
    root_path="events_lake",
    partition_cols=["status"],       # year/month for time series
)
# → events_lake/status=P1/0000.parquet  events_lake/status=P2/...
# DuckDB/pandas read only the partitions a filter touches

# --- JSONL: the pipeline interchange --------------------------------
with open("embeddings.jsonl", "w", encoding="utf-8") as f:
    for row in df.itertuples():
        f.write(json.dumps({"id": row.id, "text": row.text}) + "\n")
# Read back with constant memory:
# for line in open("embeddings.jsonl"): rec = json.loads(line)