topic: RAG Retrieval, Step by Step
Parent-Child Retrieval
Retrieve precise child chunks and return larger parent context to the LLM, with evaluation guidance, trade-offs, and a LangChain example.
Table of contents
Introduction
Parent-Child Retrieval is a retrieval strategy used when a small chunk is good for finding the correct information, but does not contain enough surrounding context for the LLM to answer correctly.
The main idea is:
Child Chunk → used for retrieval
Parent Chunk → returned to the LLM
This is useful when production evaluation shows that the correct chunk is already appearing in the Top-K results, but the final answer is still incomplete because an exception, condition, or related rule exists nearby.
How It Works
Imagine we have API documentation:
POST /payments
Transactions above $10,000 require manual verification.
Enterprise accounts using batch settlement cannot use this endpoint.
They must use /batch-payments instead.
We can first create a larger parent chunk, then split it into smaller child chunks:
Parent: POST /payments
│
├── Child 1: endpoint description
├── Child 2: $10,000 verification rule
└── Child 3: enterprise batch settlement restriction
Suppose a user asks:
Can an enterprise batch-settlement account send a $15,000 payment?
Semantic search may correctly retrieve:
Child 2:
Transactions above $10,000 require manual verification.
If only this child is passed to the LLM, it may incorrectly answer:
Yes, but manual verification is required.
The problem is not that retrieval completely failed. It found the relevant location, but the retrieved context was too small.
With Parent-Child Retrieval:
Query
↓
Search small Child chunks
↓
Child 2 matched
↓
Find parent_id
↓
Retrieve POST /payments Parent
↓
Send larger Parent to LLM
Now the LLM also sees the enterprise restriction and can produce a more complete answer.
In production, I would test this when I repeatedly observe:
Correct information appears in Top-K
↓
Final answer is still incomplete
↓
Required information exists nearby
I could then A/B test:
Version A
Search 300-token children
→ Return children
Version B
Search the same 300-token children
→ Return their 1,200-token parents
Then compare answer correctness and completeness while monitoring token usage and latency.