Executive Summary & Key Takeaways
Key Insights- Derive tenant scope from authenticated identity, not user-controlled fields.
- Filter search and recheck document permissions before fetching text.
- Authorize citations, cached responses and background ingestion independently.
- Test revocation, cache isolation and cross-tenant candidate injection.
- Fail closed when required authorization services are unavailable.
Quick Definition / Direct Answer
Direct SummarySecure multi-tenant RAG requires authenticated tenant context, tenant-constrained retrieval, authoritative document-level permission checks before protected content is fetched, and independent authorization for citations and caches. Prompt instructions cannot enforce access control; unauthorized text must never reach the model.
Direct answer: Prevent cross-tenant data leaks in retrieval-augmented generation by deriving tenant scope from authenticated identity, applying authorized filters during search, verifying every document before fetching its text, and independently protecting caches and citations. Semantic similarity does not confer permission, and model prompts cannot replace server-side authorization.
The Enterprise Problem: Retrieval Can Bypass Authorization
A retrieval-augmented generation application may authenticate the user correctly yet still return a private passage belonging to another customer. This happens when authorization is enforced at the chat endpoint but not at the retrieval boundary. A vector database returns semantically similar content; similarity does not imply permission. The resulting leak may occur before the model writes an answer, because snippets can already be present in logs, reranker inputs or prompt context. Treat retrieved passages as protected data. A secure system must enforce tenant and document access at every read, including search, citation expansion, caching and background workflows. This guide focuses on implementation details rather than repeating general RAG architecture.
Threat Model and Security Objectives
Identify the principals, resources and trust boundaries. A principal is an authenticated user or service identity. Resources include documents, chunks, embeddings, search results, citations and cached responses. Trust boundaries exist between the API, identity provider, metadata database, vector index, reranker, model and telemetry pipeline. The attacker may be an authenticated user in tenant A attempting to retrieve tenant B's content, a user with revoked document access, or an external actor who injects instructions into indexed text. Define the core invariant: no unauthorized document content may enter the context, output, citation service or cache for a request. This is stronger than asking the language model to avoid mentioning private information.
Architecture Overview: Identity, Policy and Retrieval
A reference architecture has an authenticated API gateway, a server-side authorization service, a tenant-aware document registry, a vector search backend, an authorization-aware retrieval orchestrator, a reranker and an LLM. The gateway resolves the user's identity from a validated token. The authorization service determines which tenant and documents the principal can access. The orchestrator passes trusted scope constraints to search and verifies every candidate against authoritative permissions before releasing text to downstream components. The model receives only approved passages. Logs store opaque IDs and aggregate counts rather than raw sensitive content. The system must handle both ingestion-time indexing and query-time access changes; indexing a document once does not grant permanent read access.
Define a Trustworthy Tenant Identity
Never accept tenant_id from an untrusted request body as the sole basis for retrieval. Derive the active tenant from an authenticated, authorized session or token, and verify membership and requested tenant switching server-side. A service-to-service identity must also carry scoped authority rather than unrestricted access by default. For users belonging to multiple organizations, require an explicit authorized active-tenant selection. Propagate the verified tenant context through asynchronous tasks and background workers. Do not let a model-generated tool argument override that context. Use opaque tenant identifiers in logs where practical and maintain clear audit records of authorization decisions.
Document and Chunk Metadata Design
Every indexed chunk should reference an immutable document identifier and tenant identifier, plus a revision or index generation. Store access-control policy in an authoritative registry; avoid treating vector metadata as the only source of truth when permissions change independently. The vector index may include tenant and document fields for efficient filtering, but the application must validate their integrity. Record source version, classification, retention status and deletion tombstones. When a document is reindexed, ensure that old chunks cannot remain searchable indefinitely. Prefer an explicit ingestion state machine that tracks pending, indexed, revoked and deleted resources rather than silently relying on eventual consistency.
Reference Python Authorization Model
The following executable Python 3.11 example demonstrates the security invariant using a small in-memory dataset. It is intentionally not a production identity provider or vector database. The caller supplies an authenticated Principal object created by trusted middleware; user input must not construct that object directly. A document belongs to exactly one tenant, and its allowed_users set represents an illustrative document ACL. Real systems should use a centralized policy service and durable authorization records. Save as secure_retrieval.py.
from dataclasses import dataclass
@dataclass(frozen=True)
class Principal:
user_id: str
tenant_id: str
@dataclass(frozen=True)
class Document:
id: str
tenant_id: str
text: str
allowed_users: frozenset[str]
DOCUMENTS = [
Document("a-1", "tenant-a", "Alpha roadmap", frozenset({"alice"})),
Document("a-2", "tenant-a", "Alpha pricing", frozenset({"bob"})),
Document("b-1", "tenant-b", "Beta roadmap", frozenset({"carol"})),
]
def can_read(principal: Principal, document: Document) -> bool:
return (
document.tenant_id == principal.tenant_id
and principal.user_id in document.allowed_users
)
def retrieve(principal: Principal, query: str) -> list[Document]:
terms = set(query.lower().split())
if not terms:
return []
return [
document for document in DOCUMENTS
if can_read(principal, document)
and terms.intersection(document.text.lower().split())
]
if __name__ == "__main__":
alice = Principal("alice", "tenant-a")
assert [d.id for d in retrieve(alice, "roadmap")] == ["a-1"]
assert retrieve(alice, "pricing") == []Enforce Access Before Sensitive Content Leaves Storage
The example filters before returning document text. In a real vector system, push tenant constraints and authorized document filters into the search operation wherever the backend supports them. This reduces unnecessary exposure and prevents unauthorized candidates from entering reranking or generation. A second server-side authorization check should validate each returned document against the authoritative policy before its content is fetched or passed onward. A post-filter alone is not always sufficient: if the search engine first ranks globally and then filters, authorized results may be omitted, and unauthorized content may already have crossed an internal boundary. Design both filtering and verification into the retrieval API.
Separate Metadata Lookup from Content Fetch
A strong design searches over constrained candidate identifiers, verifies permission and only then fetches protected text. This reduces the risk that unauthorized snippets reach logs or external reranking providers. Where a vector backend requires storing plaintext snippets, enforce strict tenant isolation at that service boundary and evaluate provider data handling. Treat embeddings as potentially sensitive derived data, not automatically anonymous information. Do not expose raw vector search responses directly to clients. A citation endpoint must independently authorize its document and passage identifiers rather than trusting that a previous chat response was authorized.
A Safe Retrieval Boundary with Defensive Checks
The next function demonstrates defense in depth when an upstream search system returns candidate document IDs. It does not trust the search results to be authorized. The function fetches records from an authoritative mapping, checks access and returns only permitted text. Unknown identifiers are ignored rather than producing a cross-tenant error message that reveals document existence. The example uses the same Principal, Document and can_read definitions above.
def authorized_context(
principal: Principal,
candidate_ids: list[str],
registry: dict[str, Document],
) -> list[str]:
passages: list[str] = []
for document_id in candidate_ids:
document = registry.get(document_id)
if document is None:
continue
if can_read(principal, document):
passages.append(document.text)
return passages
registry = {d.id: d for d in DOCUMENTS}
alice = Principal("alice", "tenant-a")
assert authorized_context(
alice, ["b-1", "a-2", "a-1"], registry
) == ["Alpha roadmap"]Prevent Permission TOCTOU Races
Time-of-check to time-of-use races occur when permissions change between search, authorization and content delivery. A user may lose access while a long retrieval or generation request is running. Define an explicit consistency policy based on application risk. For sensitive systems, validate authorization at content fetch and again before delivering a cached citation or downloadable source. Version access-control decisions or use short-lived authorization capabilities when appropriate. Invalidate relevant caches after revocation. Do not promise instantaneous revocation if the architecture cannot enforce it; measure and document the maximum propagation delay and design compensating controls.
Cache Keys Must Include Authorization Scope
A cache keyed only by question text can leak one tenant's answer to another tenant asking the same question. Include verified tenant, user or authorization-scope identity, policy version, model or prompt version and source-index generation as relevant to the cached artifact. Prefer caching non-sensitive computations where possible. Never treat a cached answer as authorized solely because it was safe when first generated. On permission revocation, invalidate dependent entries or reauthorize before serving. Avoid placing raw credentials or personally identifying information into cache keys that appear in operational logs. Test cache isolation independently of the retrieval implementation.
Secure Ingestion and Background Jobs
Ingestion workers often have elevated access to document stores and vector indexes. Scope service identities by tenant and operation, validate document ownership before indexing, and attach immutable tenant identifiers from trusted metadata rather than untrusted file content. Enforce malware scanning, file-type validation and size limits before parsing uploads. Use idempotent ingestion jobs, versioned indexes and deletion tombstones. When access is revoked, ensure that cached chunks, search indexes and backups follow the documented lifecycle. Prevent an indexing job from accidentally writing chunks into a shared namespace with missing tenant filters. Audit both successful and rejected ingestion attempts.
Prompt Injection Does Not Replace Access Control
Retrieved documents can contain adversarial instructions such as requests to reveal hidden data or call unauthorized tools. The model must treat those instructions as untrusted content. However, prompt-injection defenses are not a substitute for retrieval authorization: once secret text enters the prompt, the boundary has already failed. Restrict tool permissions independently of the model, validate every tool argument against the principal's scope and use explicit output controls for sensitive fields. Avoid including secrets in system prompts. Separate the authorization decision from any model-generated justification. Evaluate prompt injection and cross-tenant retrieval as different security test categories.
Automated Cross-Tenant Isolation Tests
Tests should prove both positive access and denied access. A useful suite creates two tenants, users with different document ACLs and overlapping document vocabulary. It then asserts that a user cannot retrieve another tenant's document or a same-tenant document without permission. Include unknown IDs, revoked permissions, stale cache entries and malicious candidate lists. The following pytest example uses the executable Python module above. Save as test_secure_retrieval.py and run pytest -q after installing pytest.
from secure_retrieval import (
DOCUMENTS, Principal, authorized_context,
retrieve,
)
def test_cross_tenant_search_is_denied():
alice = Principal("alice", "tenant-a")
ids = [d.id for d in retrieve(alice, "roadmap")]
assert ids == ["a-1"]
def test_same_tenant_acl_is_denied():
alice = Principal("alice", "tenant-a")
assert retrieve(alice, "pricing") == []
def test_untrusted_candidate_ids_are_rechecked():
registry = {d.id: d for d in DOCUMENTS}
alice = Principal("alice", "tenant-a")
assert authorized_context(
alice, ["b-1", "a-2", "a-1"], registry
) == ["Alpha roadmap"]Observe Authorization Without Logging Secrets
Production monitoring should count denied retrievals, missing tenant filters, invalid policy versions, unauthorized citation requests, cache invalidation lag and ingestion failures. Track incidents by opaque document and request IDs. Do not log protected snippets or full prompts by default. Use structured audit events with a decision outcome, principal reference, tenant scope, policy version and trace ID. Keep retention and access rules proportionate to the data classification. Alert on unexpected cross-tenant candidates even when a downstream check successfully blocks them, because that may indicate a vector index misconfiguration or policy integration defect.
Operational Failure Modes and Recovery
If the authorization service is unavailable, fail closed for protected retrieval rather than returning unfiltered search results. If a vector query cannot enforce required tenant constraints, reject the request or use a separately validated safe path. If policy changes lag behind the index, consult the authoritative registry before fetching text. If an incident exposes private content, stop affected retrieval paths, preserve appropriate audit evidence, revoke relevant access, identify impacted tenants and follow the incident-response policy. Avoid destructive log cleanup that would remove forensic evidence. Rebuild contaminated indexes and invalidate caches after the root cause is corrected.
Performance and Architecture Trade-Offs
Permission checks add latency and can reduce the candidate pool for semantic ranking. Optimize safely through batched authorization queries, indexed tenant and document identifiers, policy decision caching with explicit invalidation, and partitioned vector namespaces when appropriate. Do not trade authorization correctness for an unmeasured latency improvement. Dedicated tenant indexes provide stronger operational separation but increase index management cost. Shared indexes with metadata filters may be economical, but demand rigorous policy enforcement and leakage testing. Choose based on tenant count, data sensitivity, retrieval volume, regulatory requirements and operational capabilities rather than generic benchmarks.
Deployment and Security Review Checklist
- Derive tenant identity from verified authentication and membership.
- Enforce authorization before content enters rerankers, prompts or responses.
- Verify every candidate against authoritative document permissions.
- Authorize citation and source-download endpoints independently.
- Include authorization scope and policy version in cache design.
- Test revocation races and index propagation delays.
- Scope ingestion workers and validate tenant ownership.
- Fail closed when mandatory policy checks are unavailable.
- Monitor rejected cross-tenant candidates without logging protected text.
- Rehearse isolation tests and incident-response procedures before launch.
Frequently Asked Questions
Does a tenant filter in a vector database guarantee isolation?
No. It helps constrain search, but trusted tenant resolution, document-level authorization, cache isolation and citation checks are still required.
Can a model prompt prevent cross-tenant leaks?
No. Prompts are not authorization boundaries. Unauthorized passages must be blocked before they reach the model.
Should each tenant have a separate vector index?
Separate indexes can simplify some isolation boundaries but increase operational overhead. Shared indexes require strong filtering, authoritative rechecks and tests.
How should permission revocation affect existing answers?
Reauthorize sensitive cached results and citations, invalidate affected caches and define explicit consistency guarantees for in-flight requests.
Related Acadify Engineering Guides
For broader retrieval design, read Enterprise RAG Architecture: Hybrid Search, Reranking, Security & Observability. For testing retrieval quality and grounding, see Enterprise RAG Evaluation. For graph-based retrieval, see Agentic GraphRAG Production Guide. These articles complement the authorization-first implementation focus of this guide.
Conclusion
Secure RAG begins with a simple rule: relevance never grants access. Authenticate principals, enforce tenant and document permissions before protected content is released, recheck candidates and citations, and invalidate caches when permissions change. Use automated cross-tenant tests and operational monitoring to demonstrate the isolation boundary. Production confidence comes from verified authorization and recovery procedures, not from model instructions promising confidentiality.
Glossary & Key Architecture Definitions
- • Multi-tenant RAG: Retrieval-augmented generation serving multiple isolated organizations.
- • Document-level authorization: Policy determining whether a principal may read a document.
- • Tenant filter: Search constraint limiting candidates to an authorized tenant.
- • TOCTOU race: Permission change between an access check and resource use.
- • Fail closed: Deny access when a mandatory security decision cannot be completed.
Engineering Research & Citations
- [1] OWASP Top 10 for LLM Applications: https://genai.owasp.org/llm-top-10/
- [2] OWASP Authorization Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Authorization_Cheat_Sheet.html
- [3] OWASP API Security Top 10: https://owasp.org/API-Security/editions/2023/en/0x11-t10/
- [4] NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework
No perspectives submitted yet. Be the first to start the discussion.