Generation & Citations

Great retrieval, sloppy generation = a confident liar with footnotes. This reel: assembling context, forcing citations, and teaching the model to say “I don’t know.”

▶ Watch this reel

What you'll learn

  1. Context assembly
  2. Grounded answers & citations
  3. The courage to not know
  4. Conversation-aware RAG

Remember this

Context assembly

Grounded answers & citations

“I don't know”

Conversation-aware RAG

Code: Generation stage: assembly, citations, abstention, follow-ups

SYSTEM = """You are a helpful assistant. Rules:
1. Answer ONLY from the <context> below.
2. Cite every claim as [source-label].
3. If the context lacks the answer, say exactly:
   "I don't have that information."
Never use outside knowledge."""

def ask(history: list[dict], question: str) -> str:
    # 1 · follow-up → standalone query
    standalone = llm(
        "Rewrite as a self-contained search query.",
        history=history, question=question)

    # 2 · retrieve (GA-15 stack)
    chunks = retrieve(standalone, k=5)

    # 3 · mechanical abstention — before any LLM call
    if not chunks or chunks[0].score < 0.62:
        return "I couldn't find that in the handbook. Search the wider wiki?"

    # 4 · assemble: best first AND last (U-shaped attention)
    ordered = [chunks[0], *chunks[1:], chunks[0]] if len(chunks) > 2 else chunks
    context = "\n".join(
        f'<doc src="{c.source}">{c.text}</doc>' for c in ordered)

    # 5 · generate with history + grounded context
    return chat(SYSTEM, history + [
        {"role": "user",
         "content": f"<context>\n{context}\n</context>\n\n{standalone}"}])