Short answer
Retrieval-augmented generation (RAG) is a method where an AI system first retrieves relevant information from a set of documents or data, then gives that information to a language model to generate the answer. It lets AI answer from your own policies, manuals or product data, reduces made-up answers, and keeps responses current without retraining the model.
By Chandan Maheshwari, AI mentor & consultant, founder of School of AI · Last updated
Key takeaways
- RAG = retrieve relevant information, then generate the answer from it.
- Common pipeline: documents → chunks → embeddings → retrieval → LLM → answer.
- Used for company knowledge assistants and support bots.
- Answer quality depends on document quality.
How RAG works
The typical pipeline.
- Collect documents — policies, manuals, FAQs, product data
- Split them into chunks
- Convert chunks into embeddings and store them
- On each question, retrieve the most relevant chunks
- Give those chunks to the LLM as context
- Generate an answer, ideally citing the source
Where businesses use RAG
Internal knowledge assistants, customer support answers, sales and product Q&A, and searching contracts or SOPs.
Learning this with School of AI
AI & Automation Program
Live, mentor-led and online, open to learners anywhere in India.