Table of contents
The single requirement that separates enterprise RAG from a basic implementation is this: the retriever must not return information the current user isn't authorized to access. Everything else — hybrid search, reranking, evaluation, observability — matters for quality. Permission-aware retrieval is what makes the difference between a system you can deploy across an organization and one you can only run as an internal demo with a handful of trusted users.
This guide covers the full reference architecture for production-grade enterprise RAG: connectors and ingestion, document processing, permission-aware retrieval, the LLM and guardrail layers, evaluation, observability, and scaling — with a production-readiness checklist at the end. It assumes the reader is already familiar with basic RAG concepts; if not, start with our [RAG development guide](/blog/rag-development-guide) first.
- Non-negotiable requirement
- Permission-aware retrieval
- Most overlooked layer
- Ingestion — change detection & deletion propagation
- Security model
- RBAC + document-level ABAC, not app-level auth alone
- Before production
- A defined evaluation set, not manual spot-checks
Full reference architecture
This is the shape a production-grade enterprise RAG system converges toward, whether it's built incrementally or designed upfront.
Enterprise data sources
Document stores, wikis, databases, CRM/ERP, file systems, APIs
Connector layer
Source-specific connectors for each system's access pattern
Ingestion pipeline
Scheduled and event-driven pulls, with change and deletion detection
Document processing
Parsing, OCR where needed, structure-aware chunking, PII handling
Embedding service
Converts processed chunks into vector representations
Vector / search layer
Vector search combined with keyword (BM25) indexing
Permission-aware retrieval
Filters results by the requesting user's actual access rights
Reranking
Re-scores permitted candidates for relevance before the LLM sees them
Prompt / context builder
Assembles the final context window with citations metadata
LLM gateway
Model selection, fallback, token control across one or more providers
Guardrails
Groundedness and citation validation before the answer is returned
Application / API
The interface or API surfacing the grounded answer
Observability & audit
Tracing, logging and audit trail across every step above
Data sources and connectors
Enterprise data typically lives across document stores, wikis, structured databases, data warehouses, file systems, internal APIs, CRM and ERP systems, and internal documentation tools. Each source needs its own connector, respecting that source's own access model and rate limits — a connector isn't just "read the API," it's understanding what that system's permission model looks like so it can be preserved downstream. (This article describes architecture patterns generally — it doesn't imply CodeSurge integrates with any specific named platform unless stated elsewhere on this site.)
Ingestion layer
Beyond initial ingestion, production systems need: scheduled ingestion for sources without change notifications, event-driven ingestion for sources that support webhooks or change feeds, change detection so updated documents are re-indexed rather than duplicated, versioning so retrieval reflects the current version of a document, and — frequently missed — deletion propagation, so a document removed from the source system is actually removed from the retrieval index, not left retrievable indefinitely.
Processing layer
Parsing handles the actual extraction of text and structure from source documents — PDFs, Word documents, HTML, presentations — including OCR where documents are scanned images rather than text. Document structure (headings, sections, tables) should inform chunking rather than being discarded. Metadata enrichment attaches source, owner, sensitivity classification, and access-relevant attributes to each chunk. Deduplication prevents near-identical content from multiple sources cluttering retrieval results. PII handling — detecting and appropriately redacting or flagging personal data — is a requirement in most regulated environments, not an optional hardening step.
Retrieval layer
Production retrieval combines embeddings-based vector search with keyword/BM25 search in a hybrid approach, applies metadata filtering before or alongside similarity search, uses reranking to improve precision on the final candidate set, and increasingly applies query expansion — generating related query variants to improve recall on ambiguous questions. The specific retrieval strategy (how these are combined and weighted) should be tuned against a real evaluation set, not assumed from a generic best practice.
Security
This is the section that determines whether an enterprise RAG system is actually safe to deploy.
The retriever must not return information the current user is not authorized to access. This single principle drives most of what follows.
Application-level auth is not the same as permission-aware retrieval
Authenticating a user at the application layer tells you who they are. It does not, by itself, prevent the retriever from pulling a document the user shouldn't see into the LLM's context. Permission checks need to happen at the retrieval layer itself — filtering candidates by the user's actual access rights before they ever reach the model — not only at the application's outer edge.
Security controls that matter specifically for RAG, beyond standard application security:
- Authentication and authorization — verifying identity, then checking what that identity is permitted to access.
- RBAC and ABAC — role-based access for broad permission tiers; attribute-based access control where permissions depend on document or user attributes (department, classification level, project membership) rather than role alone.
- Document-level permissions — access control enforced per document or chunk, not just per data source.
- Tenant isolation — for multi-customer systems, one tenant's data must never be retrievable by another tenant's queries, enforced structurally rather than by convention.
- Row/document security propagation — if the source system (a database, a CRM) has row-level security, that model needs to be preserved through ingestion into the retrieval layer, not flattened away.
- Permission-aware retrieval — the actual filtering step described above.
- Encryption and secrets management — data encrypted at rest and in transit; API keys and credentials in a secrets manager, never in code or prompts.
- Audit logging — every query, retrieved document, and generated answer logged with enough detail to reconstruct what happened.
- Data residency — where regulation requires it, ensuring data and processing stay within required jurisdictions.
- Prompt injection defenses — treating retrieved document content as untrusted input that could contain instructions aimed at manipulating the model, not just the user's own query.
- Protection against data exfiltration and malicious documents — a compromised or adversarial document in the corpus shouldn't be able to cause the system to leak unrelated data.
- Access-token propagation — where retrieval calls out to live systems (rather than a pre-indexed copy), the user's own access token — not a shared service credential — should scope what's returned.
These controls are individually necessary but not sufficient on their own — they need to sit inside a broader policy for who can deploy what, on what data, with what oversight. See our AI governance framework guide for how that fits together at an organizational level.
LLM layer
A model gateway abstracts which specific LLM provider or model version handles a request, enabling fallback and controlled model selection without changing application code. Prompt templates standardize how context and instructions are assembled. Context management controls what fits in the context window and in what order. Citations should be structurally tied to the specific retrieved chunk, not just mentioned in free text. Structured output (JSON schemas, defined formats) makes answers reliably consumable by downstream systems. Fallback logic handles provider outages or rate limits gracefully. Token control manages cost and latency by bounding how much context is actually sent.
Guardrails
Beyond the LLM layer's own controls: content controls filter inappropriate or out-of-policy content; groundedness checks verify the answer is actually supported by retrieved context rather than the model's own assumptions; answer validation and citation validation confirm claims are structurally traceable to sources; tool restrictions limit what actions a RAG-backed assistant can trigger if it has any action-taking capability; and sensitive data handling ensures PII or classified content surfaced in an answer is handled per policy, not just per the retriever's confidence score.
Evaluation
Production RAG evaluation is a discipline, not a one-time check before launch:
| Metric | What it tells you |
|---|---|
| Retrieval precision | How much of what's retrieved is actually relevant |
| Retrieval recall | How much of the relevant content is actually being retrieved |
| Answer relevance | Whether the generated answer actually addresses the question |
| Faithfulness / groundedness | Whether the answer is supported by the retrieved context, not fabricated |
| Citation correctness | Whether cited sources actually support the specific claim made |
| Latency | End-to-end response time under real load |
| Cost | Token usage and infrastructure cost per query |
| User feedback | Direct signal from real usage — thumbs up/down, corrections, escalations |
Observability
Production systems need tracing across the full request path (query → retrieval → rerank → generation → response), token usage tracking for cost control, retrieval logs showing exactly what was retrieved for each query, latency breakdowns by stage, error tracking, logged model responses for debugging and audit, a channel for user feedback, and dashboards on the quality metrics above — not just uptime. Our LLM observability guide goes deeper on tracing strategy and the difference between observability and evaluation.
Scaling
As usage grows: caching for repeated or similar queries reduces redundant retrieval and generation cost; asynchronous ingestion and queues prevent large re-indexing jobs from blocking normal operation; horizontal scaling of the retrieval and generation layers handles increased query volume; vector database scaling (sharding, replica strategies) becomes relevant well before most teams expect it to; rate limiting protects both cost and downstream system stability; and cost control — capping token usage, choosing cheaper models for lower-stakes queries — should be a deliberate architectural decision, not an afterthought when a bill arrives.
Basic RAG vs. enterprise RAG
| Dimension | Basic RAG | Enterprise RAG |
|---|---|---|
| Retrieval | Vector search over one source | Hybrid, multi-source, permission-filtered |
| Security | None or app-level auth only | RBAC/ABAC, document-level, tenant-isolated |
| Ingestion | One-time or manual | Scheduled + event-driven, with deletion propagation |
| Evaluation | Ad hoc manual review | Maintained evaluation set, tracked over time |
| Observability | Basic logs | Full tracing, audit trail, quality dashboards |
| Scaling | Single instance | Horizontally scaled, cached, queued ingestion |
| Governance | None | Data classification, retention policy, access review |
- Permission-aware retrieval enforced at the retrieval layer, not just app-level auth
- Document-level access control and tenant isolation verified with real test cases
- Deletion and update propagation confirmed — removed source documents are actually unretrievable
- PII detection and handling policy defined and implemented
- Hybrid retrieval (vector + keyword) with reranking, tuned against real queries
- Groundedness and citation validation guardrails in place
- Prompt injection defenses tested against adversarial document content
- Evaluation set covering real, representative questions — re-run on every pipeline change
- Full tracing and audit logging across the request path
- Rate limiting and cost controls configured before wide rollout
- Data residency and retention requirements confirmed against applicable regulation
- Incident response plan defined for a data exposure or hallucinated-answer incident
Treat this as a minimum bar before deploying to a broad internal audience, not just before an external launch.
Planning an enterprise RAG deployment?
Talk to our engineering team about permission-aware retrieval, security architecture and evaluation before you commit to a build.
Frequently asked questions
What makes enterprise RAG different from a basic RAG implementation?+
Permission-aware retrieval, document-level access control and tenant isolation, continuous evaluation, full observability, and governance around data classification and retention — none of which a basic demo needs to handle.
How do you prevent a RAG system from leaking data a user shouldn't see?+
By enforcing permission checks at the retrieval layer itself — filtering candidate documents by the requesting user's actual access rights before they reach the LLM — not relying on application-level authentication alone.
What is permission-aware retrieval?+
A retrieval design where every search result is filtered against the requesting user's real access permissions (role, document-level grants, tenant membership) before being included in the context sent to the LLM.
How do you evaluate a production RAG system?+
Against a maintained set of real, representative questions, measuring retrieval precision and recall, answer faithfulness/groundedness, citation correctness, latency and cost — re-run whenever the ingestion or retrieval pipeline changes, not just once before launch.
What is RBAC vs. ABAC in the context of RAG?+
RBAC (role-based access control) grants access by role — e.g., "HR team." ABAC (attribute-based access control) grants access based on attributes of the user, document or context — e.g., department, classification level or project membership — and is often needed for finer-grained document-level permissions.
How do you handle multi-tenant RAG securely?+
Tenant isolation must be enforced structurally — at the data storage, indexing and retrieval layers — so one tenant's queries can never retrieve another tenant's data, rather than relying on application logic alone to keep tenants separate.
What is prompt injection, and does it apply to RAG?+
Prompt injection is when text (often from an untrusted source) contains instructions designed to manipulate the model's behavior. In RAG systems, retrieved documents are a real injection vector — a malicious or compromised document in the corpus should be treated as untrusted input, not implicitly trusted content.
Written by
CodeSurge AI Engineering Team
The CodeSurge AI team designs and builds AI systems, SaaS products and enterprise integrations for clients in India, the UAE and beyond — this section shares the architecture patterns, cost drivers and implementation tradeoffs we work through on real projects.