RAG Prompt Engineering: Keeping the Model Grounded in Retrieved Documents
RAG prompts that do not ground the model in retrieved documents defeat retrieval. Here is the structure that keeps answers grounded. Retrieval-augmented generation, or RAG, combines a retrieval step that finds relevant documents with a generation step that answers using those documents. The prompt is what connects the two: it tells the model to answer from the retrieved documents, to cite them, and to avoid using information that is not in them. A weak RAG prompt produces answers that ignore the retrieved documents and fall back on the model training data, which defeats the purpose of retrieval. After building several RAG features, I have a prompt structure that keeps the model grounded in the retrieved context. This guide covers it. The RAG Prompt Skeleton My RAG prompts follow a fixed skeleton: a grounding instruction, the retrieved documents, the user question, and the answer rules. The grounding instruction tells the model to answer only from the provided documents. The retrieved documents are delimited clearly so the model treats them as the source of truth. The question comes after the documents. The answer rules specify citation, length, and refusal behavior when the documents do not contain the answer. // RAG prompt skeleton Answer the question using only the documents below. If the documents do not contain the answer, say you do not know. Do not use outside knowledge. Cite each claim with the document number in brackets.