Agent Reasoning
Four reasoning strategies — ReAct, plan-and-execute, reflection, decomposition — the difference between an agent that lunges and an agent that thinks.
▶ Watch this reelWhat you'll learn
- ReAct
- Plan-and-execute
- Reflection & self-critique
- Task decomposition
Remember this
- ReAct interleaves reasoning and acting — grounded but myopic and costly per step; plan-and-execute sees the whole board and needs budgeted re-planning
- Reflection improves drafts but can't see its own blind spots — pair with external verification; process reflection seeds memory
- Decomposition: verifiable artifacts, explicit dependency DAGs, one-context steps, riskiest-first ordering
ReAct
- Reason ↔ act interleaved; grounded in fresh observations.
- Weaknesses: per-step cost, short-horizon myopia, reasoning drift.
Plan-and-execute
- Plan (ordered, dependency-aware, reviewable) → execute cheaply → re-plan on failure (budgeted).
- Long horizons; hybrid: plan milestones, ReAct within.
Reflection
- Structured critique (dimensions, artifact-not-intention) → revise.
- Blind spots survive self-review → external verification.
- Process reflection → durable notes → memory (AG-12).
Task decomposition
- Verifiable artifacts · explicit dependency DAG · one-context steps.
- MECE; reformulate the goal; riskiest step first.
Code: Hybrid reasoning: plan milestones, ReAct within, reflect at the end
async def run_hybrid(goal: str):
# 1 · PLAN — global view, machine-readable steps
plan = await llm(PLAN_PROMPT, goal=goal,
schema="list[{step, depends_on, verify_by}]")
await human_review(plan) # checkpoint
for step in topological_order(plan):
# 2 · EXECUTE — ReAct within the milestone
artifact = await react_loop(
f"{goal}\nCurrent step: {step.step}",
tools=tools_for(step), max_iters=8)
# 3 · VERIFY per step — the plan's own verification hook
if not verify(step.verify_by, artifact):
plan = await replan(plan, failed=step, state=artifact)
# 4 · MACRO-REFLECTION — durable lessons into memory
lessons = await llm("What worked? What failed? Rules for next time?",
transcript=full_history())
memory.store(goal_domain(goal), lessons)
return collect_artifacts(plan)