RAG retrieves passages related to a question and sends those passages to the model together with the question.
The problem it addresses
A model does not contain your private documents, or a page that changed this morning. Retrieval supplies context. It is not a retraining run.
A minimal pipeline
Split documents, build an index, retrieve a few passages, then generate. Too many passages bury the useful part. Too few leave the model guessing.
contexts = index.search(question, limit=4)
answer = model.generate(question, contexts)
Where it fails
A chunk cuts a definition in half, the query and the document use different words, or ranking prefers a similar but irrelevant passage. Read the retrieved text before blaming the model.
Tip
If the retrieved text is wrong, a larger model usually will not repair the answer.
Discussion