RAG

Enterprise RAG Architecture: How to Build Secure Production-Grade AI Search

Permission-aware retrieval, tenant isolation, evaluation and observability — the parts of a RAG system that only matter once real users, real data sensitivity and real scale are involved.

CodeSurge AI Engineering TeamPublished 11 September 202610 min read
Table of contents

The single requirement that separates enterprise RAG from a basic implementation is this: the retriever must not return information the current user isn't authorized to access. Everything else — hybrid search, reranking, evaluation, observability — matters for quality. Permission-aware retrieval is what makes the difference between a system you can deploy across an organization and one you can only run as an internal demo with a handful of trusted users.

This guide covers the full reference architecture for production-grade enterprise RAG: connectors and ingestion, document processing, permission-aware retrieval, the LLM and guardrail layers, evaluation, observability, and scaling — with a production-readiness checklist at the end. It assumes the reader is already familiar with basic RAG concepts; if not, start with our [RAG development guide](/blog/rag-development-guide) first.

Quick answer
Non-negotiable requirement
Permission-aware retrieval
Most overlooked layer
Ingestion — change detection & deletion propagation
Security model
RBAC + document-level ABAC, not app-level auth alone
Before production
A defined evaluation set, not manual spot-checks

Full reference architecture

This is the shape a production-grade enterprise RAG system converges toward, whether it's built incrementally or designed upfront.

Enterprise RAG reference architecture

Enterprise data sources

Document stores, wikis, databases, CRM/ERP, file systems, APIs

Connector layer

Source-specific connectors for each system's access pattern

Ingestion pipeline

Scheduled and event-driven pulls, with change and deletion detection

Document processing

Parsing, OCR where needed, structure-aware chunking, PII handling

Embedding service

Converts processed chunks into vector representations

Vector / search layer

Vector search combined with keyword (BM25) indexing

Permission-aware retrieval

Filters results by the requesting user's actual access rights

Reranking

Re-scores permitted candidates for relevance before the LLM sees them

Prompt / context builder

Assembles the final context window with citations metadata

LLM gateway

Model selection, fallback, token control across one or more providers

Guardrails

Groundedness and citation validation before the answer is returned

Application / API

The interface or API surfacing the grounded answer

Observability & audit

Tracing, logging and audit trail across every step above

Cross-cutting concerns — identity, security, tenant isolation, monitoring and governance — apply to every layer in this diagram, not just one step. They're covered as their own sections below.

Data sources and connectors

Enterprise data typically lives across document stores, wikis, structured databases, data warehouses, file systems, internal APIs, CRM and ERP systems, and internal documentation tools. Each source needs its own connector, respecting that source's own access model and rate limits — a connector isn't just "read the API," it's understanding what that system's permission model looks like so it can be preserved downstream. (This article describes architecture patterns generally — it doesn't imply CodeSurge integrates with any specific named platform unless stated elsewhere on this site.)

Ingestion layer

Beyond initial ingestion, production systems need: scheduled ingestion for sources without change notifications, event-driven ingestion for sources that support webhooks or change feeds, change detection so updated documents are re-indexed rather than duplicated, versioning so retrieval reflects the current version of a document, and — frequently missed — deletion propagation, so a document removed from the source system is actually removed from the retrieval index, not left retrievable indefinitely.

Processing layer

Parsing handles the actual extraction of text and structure from source documents — PDFs, Word documents, HTML, presentations — including OCR where documents are scanned images rather than text. Document structure (headings, sections, tables) should inform chunking rather than being discarded. Metadata enrichment attaches source, owner, sensitivity classification, and access-relevant attributes to each chunk. Deduplication prevents near-identical content from multiple sources cluttering retrieval results. PII handling — detecting and appropriately redacting or flagging personal data — is a requirement in most regulated environments, not an optional hardening step.

Retrieval layer

Production retrieval combines embeddings-based vector search with keyword/BM25 search in a hybrid approach, applies metadata filtering before or alongside similarity search, uses reranking to improve precision on the final candidate set, and increasingly applies query expansion — generating related query variants to improve recall on ambiguous questions. The specific retrieval strategy (how these are combined and weighted) should be tuned against a real evaluation set, not assumed from a generic best practice.

Security

This is the section that determines whether an enterprise RAG system is actually safe to deploy.

The retriever must not return information the current user is not authorized to access. This single principle drives most of what follows.

Application-level auth is not the same as permission-aware retrieval

Authenticating a user at the application layer tells you who they are. It does not, by itself, prevent the retriever from pulling a document the user shouldn't see into the LLM's context. Permission checks need to happen at the retrieval layer itself — filtering candidates by the user's actual access rights before they ever reach the model — not only at the application's outer edge.

Security controls that matter specifically for RAG, beyond standard application security:

  • Authentication and authorization — verifying identity, then checking what that identity is permitted to access.
  • RBAC and ABAC — role-based access for broad permission tiers; attribute-based access control where permissions depend on document or user attributes (department, classification level, project membership) rather than role alone.
  • Document-level permissions — access control enforced per document or chunk, not just per data source.
  • Tenant isolation — for multi-customer systems, one tenant's data must never be retrievable by another tenant's queries, enforced structurally rather than by convention.
  • Row/document security propagation — if the source system (a database, a CRM) has row-level security, that model needs to be preserved through ingestion into the retrieval layer, not flattened away.
  • Permission-aware retrieval — the actual filtering step described above.
  • Encryption and secrets management — data encrypted at rest and in transit; API keys and credentials in a secrets manager, never in code or prompts.
  • Audit logging — every query, retrieved document, and generated answer logged with enough detail to reconstruct what happened.
  • Data residency — where regulation requires it, ensuring data and processing stay within required jurisdictions.
  • Prompt injection defenses — treating retrieved document content as untrusted input that could contain instructions aimed at manipulating the model, not just the user's own query.
  • Protection against data exfiltration and malicious documents — a compromised or adversarial document in the corpus shouldn't be able to cause the system to leak unrelated data.
  • Access-token propagation — where retrieval calls out to live systems (rather than a pre-indexed copy), the user's own access token — not a shared service credential — should scope what's returned.

These controls are individually necessary but not sufficient on their own — they need to sit inside a broader policy for who can deploy what, on what data, with what oversight. See our AI governance framework guide for how that fits together at an organizational level.

LLM layer

A model gateway abstracts which specific LLM provider or model version handles a request, enabling fallback and controlled model selection without changing application code. Prompt templates standardize how context and instructions are assembled. Context management controls what fits in the context window and in what order. Citations should be structurally tied to the specific retrieved chunk, not just mentioned in free text. Structured output (JSON schemas, defined formats) makes answers reliably consumable by downstream systems. Fallback logic handles provider outages or rate limits gracefully. Token control manages cost and latency by bounding how much context is actually sent.

Guardrails

Beyond the LLM layer's own controls: content controls filter inappropriate or out-of-policy content; groundedness checks verify the answer is actually supported by retrieved context rather than the model's own assumptions; answer validation and citation validation confirm claims are structurally traceable to sources; tool restrictions limit what actions a RAG-backed assistant can trigger if it has any action-taking capability; and sensitive data handling ensures PII or classified content surfaced in an answer is handled per policy, not just per the retriever's confidence score.

Evaluation

Production RAG evaluation is a discipline, not a one-time check before launch:

What production RAG evaluation actually measures
MetricWhat it tells you
Retrieval precisionHow much of what's retrieved is actually relevant
Retrieval recallHow much of the relevant content is actually being retrieved
Answer relevanceWhether the generated answer actually addresses the question
Faithfulness / groundednessWhether the answer is supported by the retrieved context, not fabricated
Citation correctnessWhether cited sources actually support the specific claim made
LatencyEnd-to-end response time under real load
CostToken usage and infrastructure cost per query
User feedbackDirect signal from real usage — thumbs up/down, corrections, escalations
A maintained evaluation set of real questions, re-run whenever the pipeline changes, catches regressions that manual spot-checking reliably misses.

Observability

Production systems need tracing across the full request path (query → retrieval → rerank → generation → response), token usage tracking for cost control, retrieval logs showing exactly what was retrieved for each query, latency breakdowns by stage, error tracking, logged model responses for debugging and audit, a channel for user feedback, and dashboards on the quality metrics above — not just uptime. Our LLM observability guide goes deeper on tracing strategy and the difference between observability and evaluation.

Scaling

As usage grows: caching for repeated or similar queries reduces redundant retrieval and generation cost; asynchronous ingestion and queues prevent large re-indexing jobs from blocking normal operation; horizontal scaling of the retrieval and generation layers handles increased query volume; vector database scaling (sharding, replica strategies) becomes relevant well before most teams expect it to; rate limiting protects both cost and downstream system stability; and cost control — capping token usage, choosing cheaper models for lower-stakes queries — should be a deliberate architectural decision, not an afterthought when a bill arrives.

Basic RAG vs. enterprise RAG

Basic RAG vs. enterprise RAG
DimensionBasic RAGEnterprise RAG
RetrievalVector search over one sourceHybrid, multi-source, permission-filtered
SecurityNone or app-level auth onlyRBAC/ABAC, document-level, tenant-isolated
IngestionOne-time or manualScheduled + event-driven, with deletion propagation
EvaluationAd hoc manual reviewMaintained evaluation set, tracked over time
ObservabilityBasic logsFull tracing, audit trail, quality dashboards
ScalingSingle instanceHorizontally scaled, cached, queued ingestion
GovernanceNoneData classification, retention policy, access review
Production-readiness checklist
  • Permission-aware retrieval enforced at the retrieval layer, not just app-level auth
  • Document-level access control and tenant isolation verified with real test cases
  • Deletion and update propagation confirmed — removed source documents are actually unretrievable
  • PII detection and handling policy defined and implemented
  • Hybrid retrieval (vector + keyword) with reranking, tuned against real queries
  • Groundedness and citation validation guardrails in place
  • Prompt injection defenses tested against adversarial document content
  • Evaluation set covering real, representative questions — re-run on every pipeline change
  • Full tracing and audit logging across the request path
  • Rate limiting and cost controls configured before wide rollout
  • Data residency and retention requirements confirmed against applicable regulation
  • Incident response plan defined for a data exposure or hallucinated-answer incident

Treat this as a minimum bar before deploying to a broad internal audience, not just before an external launch.

Planning an enterprise RAG deployment?

Talk to our engineering team about permission-aware retrieval, security architecture and evaluation before you commit to a build.

Frequently asked questions

What makes enterprise RAG different from a basic RAG implementation?+

Permission-aware retrieval, document-level access control and tenant isolation, continuous evaluation, full observability, and governance around data classification and retention — none of which a basic demo needs to handle.

How do you prevent a RAG system from leaking data a user shouldn't see?+

By enforcing permission checks at the retrieval layer itself — filtering candidate documents by the requesting user's actual access rights before they reach the LLM — not relying on application-level authentication alone.

What is permission-aware retrieval?+

A retrieval design where every search result is filtered against the requesting user's real access permissions (role, document-level grants, tenant membership) before being included in the context sent to the LLM.

How do you evaluate a production RAG system?+

Against a maintained set of real, representative questions, measuring retrieval precision and recall, answer faithfulness/groundedness, citation correctness, latency and cost — re-run whenever the ingestion or retrieval pipeline changes, not just once before launch.

What is RBAC vs. ABAC in the context of RAG?+

RBAC (role-based access control) grants access by role — e.g., "HR team." ABAC (attribute-based access control) grants access based on attributes of the user, document or context — e.g., department, classification level or project membership — and is often needed for finer-grained document-level permissions.

How do you handle multi-tenant RAG securely?+

Tenant isolation must be enforced structurally — at the data storage, indexing and retrieval layers — so one tenant's queries can never retrieve another tenant's data, rather than relying on application logic alone to keep tenants separate.

What is prompt injection, and does it apply to RAG?+

Prompt injection is when text (often from an untrusted source) contains instructions designed to manipulate the model's behavior. In RAG systems, retrieved documents are a real injection vector — a malicious or compromised document in the corpus should be treated as untrusted input, not implicitly trusted content.

Written by

CodeSurge AI Engineering Team

The CodeSurge AI team designs and builds AI systems, SaaS products and enterprise integrations for clients in India, the UAE and beyond — this section shares the architecture patterns, cost drivers and implementation tradeoffs we work through on real projects.

AI EngineeringEnterprise ArchitectureSaaSCloudSoftware Development

Found this useful? Share it with your team.

Share
Keep reading

Related insights

Talk to CodeSurge AI