Beyond Text
Language was the doorway — not the house. Vision, speech, and images make GenAI a full-sense system, and every pipeline you built gains a pair of eyes and a voice.
▶ Watch this reelWhat you'll learn
- Vision & OCR
- Speech in and out
- Realtime voice agents
- Image generation
Remember this
- Vision is an input-format change: image_url content blocks in the same chat API — manage token cost with detail levels and cropping
- Speech is a solved commodity — STT/TTS wrap your existing text pipeline; sentence-level TTS streaming hides latency behind playback
- Realtime speech-to-speech replaces request/response with a persistent bidirectional stream — keep it as the conversational front door, route structured work through tool calls
- Image generation renders but doesn't reason — pipeline text-first, run generation async, and label provenance at birth
Vision & OCR
- image_url content blocks; detail low/high is the token-cost dial.
- Clean structure from Document Intelligence (GA-29) + judgment from the model.
- Ask 'from the image only' — groundedness discipline.
Speech
- STT: commodity. Chunk long audio, timestamp, LLM post-pass for domain terms.
- TTS: per-sentence streaming hides latency behind playback.
- Voice is a product choice — offer options, default unobtrusive.
Realtime voice agents
- Speech-to-speech, persistent bidirectional stream (WebSocket — PY-14).
- ~1s loop budget; interruption = cancellation event (AG-19).
- Pattern: realtime model as front door; structured work via tool calls (AG-16).
Image generation
- Render, not reason — pipeline text-first.
- Async job pattern (AG-19); editing beats raw generation for products.
- Provenance at birth (C2PA, watermark); generated media out of RAG corpora until reviewed (AG-22); transparency duties under EU AI Act.
Cross-links
GA-03 (chat API shape), GA-07 (latency), GA-10/29 (retrieval, ingestion), AG-09 (guardrails), AG-13 (context budget), AG-16 (tool contracts), AG-17 (session memory), AG-19 (async/abort), AG-22 (injection), PY-14 (websockets for realtime).
Code: The multimodal pipeline, end to end
# ONE pipeline, every sense:
async def handle(user_msg):
# eyes: screenshots/charts/doc scans (vision blocks)
# ears: audio → STT → text
# brain: text + retrieved facts → answer [UNCHANGED CORE]
# voice: answer → sentence-TTS → stream
# hands: tool calls for structured work
# guardrails: screen every surface (AG-09)
return pipeline.run(user_msg)
# The core you built in 76 reels didn't change.
# It just learned to see, hear, and speak.