Generation & Citations
Great retrieval, sloppy generation = a confident liar with footnotes. This reel: assembling context, forcing citations, and teaching the model to say “I don’t know.”
▶ Watch this reelWhat you'll learn
- Context assembly
- Grounded answers & citations
- The courage to not know
- Conversation-aware RAG
Remember this
- Assembly is design: strong chunks first AND last, delimited and source-labeled, curated — never dumped
- Citations ([Handbook §7.1]) turn an oracle into a verifiable assistant — and make debugging instant
- Abstention is engineered: prompt contract + score threshold + unanswerable questions in the eval; follow-ups need standalone-query rewriting over history
Context assembly
- Attention is U-shaped → strongest chunk first AND last.
- Delimit each chunk:
<doc src="label">…→ enables precise citation. - Curate ~5 chunks; padding = retrieval problem, fix GA-15.
Grounded answers & citations
- Every claim carries a source marker; UI renders it as a link.
- Citations = trust + debuggability (retrieval-fail vs read-fail) + compliance.
“I don't know”
- Prompt contract + score-threshold gate + unanswerable questions in the golden set.
- Good UX: abstain + offer a next step.
Conversation-aware RAG
- Rewrite follow-up → standalone query (history-aware), retrieve on it.
- Generation sees history + fresh context → consistent answers.
- Pitfalls: unbounded history (trim/summarize), stale topics (rewriter must pivot).
Code: Generation stage: assembly, citations, abstention, follow-ups
SYSTEM = """You are a helpful assistant. Rules:
1. Answer ONLY from the <context> below.
2. Cite every claim as [source-label].
3. If the context lacks the answer, say exactly:
"I don't have that information."
Never use outside knowledge."""
def ask(history: list[dict], question: str) -> str:
# 1 · follow-up → standalone query
standalone = llm(
"Rewrite as a self-contained search query.",
history=history, question=question)
# 2 · retrieve (GA-15 stack)
chunks = retrieve(standalone, k=5)
# 3 · mechanical abstention — before any LLM call
if not chunks or chunks[0].score < 0.62:
return "I couldn't find that in the handbook. Search the wider wiki?"
# 4 · assemble: best first AND last (U-shaped attention)
ordered = [chunks[0], *chunks[1:], chunks[0]] if len(chunks) > 2 else chunks
context = "\n".join(
f'<doc src="{c.source}">{c.text}</doc>' for c in ordered)
# 5 · generate with history + grounded context
return chat(SYSTEM, history + [
{"role": "user",
"content": f"<context>\n{context}\n</context>\n\n{standalone}"}])