Project: Document Processing Agent

Capstone three: PDFs in, structured fields out, low-confidence cases to humans, results into your systems — extraction at scale with honesty built in.

▶ Watch this reel

What you'll learn

  1. Field extraction
  2. Validation & exception queue
  3. Integration writes
  4. The whole loop

Remember this

Extraction

Exception queue

Integration

Template

Code: The extraction loop, one function

async def process_document(pdf_path: str) -> Route:
    pages = parse_layer.extract(pdf_path)          # text + quality flags
    record, issues = await extract_and_verify(pages)
    conf = confidence(record_valid=record is not None,
                      verify_issues=issues,
                      parse_quality=pages.quality)

    if conf >= AUTO_ACCEPT and record:
        result = await systems.upsert(record, key=record.natural_key())
        audit.log(hash(pdf_path), record, conf, result)
        return Route("written", result)

    package = review_package(pdf=pdf_path, record=record,
                             issues=issues, confidence=conf)
    await exception_queue.push(package)            # human decides
    audit.log(hash(pdf_path), record, conf, "queued")
    return Route("review", package.id)

# Invariants: idempotent (natural keys) · complete (reconciled counts) ·
# provenance (audit from byte to row) · learning (corrections → gold).