Imagine a support assistant that answers a question accurately but cites another customer's contract. Better ranking will not fix that failure. A retrieval system needs an access model that travels with the documents, their chunks and the requests that search them. Start there before comparing embedding models.
Give each source an owner
Build an inventory containing the source system, document identifier, tenant, access groups, revision and deletion policy. Decide which system is authoritative for each field. If a shared drive controls membership, copying its permissions once during ingestion leaves a gap when a person changes teams. Define how changes reach the index and what the assistant does while an update is pending.
Microsoft's RAG overview describes retrieval as the step that supplies grounding material to a language model. In your design, draw the permission check before that material enters the model context. Asking the model to ignore a confidential passage is not an access boundary.
Test with identities, not just questions
Use a small collection of synthetic documents with deliberately different access rules. Ask the same question as a support agent, a manager and a user from another tenant. Verify the returned passages and citations as well as the final answer. Include renamed documents, removed group membership and a deleted source. A correct refusal is a successful result when no authorised evidence exists.
Make the answer traceable
Keep a stable connection between each passage and the source revision it came from. A citation should lead a permitted reader to useful evidence, not merely to a document title. If the source is withdrawn, stop retrieving it and decide how previously cached answers expire. Apply the same access rules to logs and evaluation exports, which can otherwise become another copy of restricted text.
For a RAG development project, the initial deliverable should include this access map alongside retrieval examples. The accompanying data pipeline needs a repair path for missed updates, so the team can reconcile the index against the source instead of relying on a successful ingestion message.











