RAG

RAG Development Guide: Architecture, Use Cases, Cost and Best Practices

How retrieval-augmented generation actually works, what separates a weekend demo from a system you can trust in production, and what it costs to build properly.

CodeSurge AI Engineering TeamPublished 11 September 202610 min read
Table of contents

Retrieval-augmented generation (RAG) is the pattern of giving an LLM access to your own data at query time — retrieving relevant passages from a knowledge base and feeding them into the prompt — instead of relying only on what the model learned during training. It's the standard approach for building an AI assistant that answers questions from your documents, your product data, or your internal knowledge base without retraining a model.

The gap between a RAG demo and a RAG system your team actually trusts is almost entirely in the parts that don't show up in a weekend proof of concept: chunking strategy, retrieval quality, permission-aware access control, evaluation, and what happens when the model doesn't know the answer. This guide covers the architecture end to end, where teams get it wrong, and what it costs to build at each level of maturity.

Quick answer
Basic RAG prototype
₹8L – ₹15L, 4–8 weeks
Production RAG application
₹15L – ₹35L, 10–16 weeks
Enterprise RAG (permissions, multi-tenant)
₹35L – ₹80L+
Most common failure mode
Poor retrieval, not a weak model

Ranges are indicative and scope-dependent — see the cost section further down for what moves a project between bands.

What RAG actually is

RAG has three moving parts: a retriever that finds relevant content from your data, a context window that content gets placed into, and an LLM that generates an answer grounded in that context rather than purely from its training data. The model itself doesn't need to know anything about your business — it just needs to be handed the right passages and asked to answer using them.

This matters because it solves the two biggest limitations of using an LLM on its own: the model's training data is frozen at some past cutoff, and it has no knowledge of your private, internal, or recently-changed information. RAG closes both gaps without touching the model itself.

RAG vs. a plain LLM application vs. fine-tuning

These three approaches get conflated constantly, and picking the wrong one is one of the more expensive mistakes we see in AI project scoping.

RAG vs. plain LLM app vs. fine-tuning
ApproachWhat it's good forWhat it can't do well
Plain LLM applicationGeneral reasoning, writing, summarizing content you provide directlyDoesn't know your private data unless you paste it in every time
RAGAnswering questions grounded in a large, changing knowledge baseDoesn't teach the model a new *style*, *format*, or *skill* — only gives it facts to reference
Fine-tuningConsistent tone, structured output formats, narrow specialized tasksExpensive to keep current; poor fit for large or frequently-changing knowledge
Most business use cases — internal knowledge assistants, support tools, document Q&A — are RAG problems, not fine-tuning problems. Reserve fine-tuning for cases where the issue is *how* the model responds, not *what* it knows.

Core RAG architecture

At its simplest, a RAG request flow looks like this:

Basic RAG request flow

User question

The query as typed or spoken by the user

Query processing

Cleaning, expansion, or rewriting the query for better retrieval

Retriever

Searches the knowledge base for relevant content

Vector / search layer

Embedding similarity search, keyword search, or both

Relevant context

The top-matching passages, assembled for the prompt

LLM

Generates an answer using only the supplied context

Grounded answer + citations

Returned to the user, ideally with sources shown

That's the request-time flow. Behind it sits a separate, ongoing pipeline that prepares your data to be retrievable in the first place — this is where most of the real engineering effort goes, and where enterprise RAG diverges sharply from a demo.

The ingestion pipeline behind enterprise RAG

Data sources

Documents, wikis, databases, tickets, APIs

Ingestion pipeline

Scheduled or event-driven pulls from each source

Parsing / chunking

Breaking documents into retrievable, appropriately-sized pieces

Embeddings

Converting each chunk into a vector representation

Vector database

Stores embeddings for fast similarity search

Retrieval layer

Combines vector search with filters and keyword matching

Reranker

Re-scores retrieved candidates for relevance before they reach the LLM

LLM

Generates the final answer from the reranked context

Guardrails

Validates the answer before it reaches the application

Application

The chat interface, API, or assistant surfacing the answer

The retrieval quality of the final answer is set almost entirely by everything before the LLM step — a strong model can't compensate for poor chunking or a weak retriever.

Document ingestion and chunking

How you split documents into chunks has more impact on answer quality than almost any other decision in a RAG system.

  • Chunk size — too large, and irrelevant content dilutes the useful part, pushing up token cost and confusing the model; too small, and you lose surrounding context needed to answer correctly.
  • Chunk overlap — a modest overlap between adjacent chunks prevents an answer from being split awkwardly across a chunk boundary.
  • Structure-aware chunking — splitting by heading, section, or logical unit (not just a fixed character count) preserves meaning far better than naive fixed-length splitting.
  • Metadata enrichment — attaching source, date, author, department or document type to each chunk enables filtering before similarity search even runs, which is often more effective than trying to fix bad chunks with a smarter model.

Embeddings and vector databases

Embeddings turn text into a numerical representation that captures semantic meaning, so "cancel my subscription" and "how do I end my plan" land close together in vector space even though they share no words. A vector database (PostgreSQL with pgvector, or dedicated options like Pinecone and Qdrant) stores these embeddings and supports fast similarity search across millions of chunks.

Pure vector (semantic) search is excellent at matching meaning but sometimes misses exact terms — product codes, error messages, specific names. Hybrid search combines vector similarity with traditional keyword search (BM25), typically outperforming either approach alone for real-world enterprise content that mixes prose with structured identifiers.

Reranking and metadata filtering

Initial retrieval usually pulls back more candidates than you'll actually use — a reranker (a smaller, more precise model focused only on relevance scoring) re-orders these candidates before the top few reach the LLM. Metadata filtering narrows the search space before or after retrieval — restricting results to a specific department, document type, or date range — which is both a relevance improvement and, as covered below, a security requirement.

Query transformation

Users don't always phrase questions in a way that matches how information is written in your source documents. Query transformation — rewriting, expanding, or decomposing a question before retrieval — often improves recall meaningfully, especially for vague or multi-part questions.

From retrieval to a grounded answer

Once relevant context is retrieved, prompt construction determines how well the model actually uses it: instructing the model to answer only from the supplied context, to say when it doesn't know rather than guessing, and to cite which passage supported each claim. Citations are not a cosmetic feature — they're what lets a user verify an answer instead of trusting it blindly, and they're one of the fastest ways to build trust in an AI assistant inside an organization.

Guardrails, evaluation and observability

Guardrails catch bad outputs before they reach a user: validating that the answer is actually grounded in the retrieved context (not hallucinated), filtering sensitive content, and having a defined fallback when retrieval comes back empty or low-confidence. Evaluation measures whether the system is actually working — retrieval precision and recall, answer faithfulness to the source material, and citation correctness — ideally against a maintained test set of real questions, not just spot-checked manually. Observability means logging every query, retrieved chunk, and generated answer so failures can be diagnosed rather than guessed at.

Security, access control and multi-tenancy

Enterprise RAG has one requirement that a demo never has to deal with: the retriever must not return information the current user isn't authorized to access. Permission-aware retrieval, document-level access control, and tenant isolation for multi-customer systems are not optional add-ons — they're core to whether the system is safe to deploy at all. This deserves a full treatment on its own; see our enterprise RAG architecture guide for how permission-aware retrieval, RBAC, and audit logging are actually built.

Basic RAG vs. production RAG vs. enterprise RAG

RAG maturity levels
CapabilityBasic RAG (demo/PoC)Production RAGEnterprise RAG
RetrievalVector search onlyHybrid search + rerankingHybrid + reranking + permission filtering
ChunkingFixed-size, naiveStructure-aware, tunedStructure-aware + metadata-rich
Access controlNoneBasic authDocument-level RBAC, tenant isolation
EvaluationManual spot-checksTest-set evaluationContinuous evaluation + user feedback loop
ObservabilityConsole logsStructured loggingFull tracing, audit logs, quality dashboards
GuardrailsNone or minimalGroundedness checksGuardrails + citation validation + fallback logic
Data freshnessOne-time ingestionScheduled re-ingestionEvent-driven, near-real-time updates
Most businesses don't need to jump straight to enterprise RAG — the right level depends on your data sensitivity, user count, and how much a wrong answer actually costs you.

Common implementation mistakes

In rough order of how often we see them cause real problems: poor retrieval quality (the model looks bad when it's actually the retriever's fault), oversized chunks that dilute relevance, stale data from a one-time ingestion nobody re-runs, duplicate context wasting tokens and confusing the model, permission leakage (retrieval returning content the user shouldn't see), and citation quality that doesn't actually match what was retrieved. Most of these are retrieval and data-pipeline problems, not model problems — and switching to a "smarter" LLM rarely fixes any of them.

When NOT to use RAG

RAG is the wrong tool when: your knowledge base is small enough to fit entirely in the model's context window on every request (sometimes a long-context prompt is simpler and cheaper than a retrieval pipeline); the task requires teaching the model a specific output format or style rather than giving it facts, which is a fine-tuning problem; or the answer needs to come from live computation (a database query, a calculation) rather than retrieved text — that's a tool-calling or agent problem, covered in our AI agent development guide.

Cost and timeline

Indicative RAG development cost by maturity level
LevelTypical scopeApproximate cost
Basic RAG prototypeSingle data source, vector search only, no auth₹8L – ₹15L
Production RAG applicationHybrid search, reranking, evaluation, basic access control₹15L – ₹35L
Enterprise RAGPermission-aware retrieval, multi-tenant, full observability₹35L – ₹80L+
Multi-source enterprise searchMultiple connected systems, continuous ingestion, governance₹60L – ₹1.2Cr+
For how these bands compare to other categories of AI project, see our broader [AI development cost guide](/blog/ai-development-cost-india).

Planning a RAG application?

Talk to our engineering team about retrieval strategy, security and implementation options before you commit to a build.

Common use cases

RAG is the underlying pattern behind most of the "AI assistant" projects we scope: internal knowledge assistants over company wikis and policy documents, customer support tools grounded in product documentation, contract and document analysis, compliance Q&A over regulatory text, and research assistants over a curated document set. If you're evaluating where AI automation fits across your business more broadly, our AI automation for business guide organizes use cases by department rather than by underlying technique.

Frequently asked questions

What is RAG (retrieval-augmented generation)?+

RAG is the pattern of retrieving relevant content from your own data at query time and feeding it into an LLM's context, so the model answers using that specific information instead of relying only on its training data.

Is RAG better than fine-tuning?+

They solve different problems. RAG is better for grounding answers in a large or frequently-changing knowledge base. Fine-tuning is better for teaching a model a consistent style, format, or narrow specialized skill. Most business knowledge-assistant use cases are RAG problems.

How much does it cost to build a RAG application?+

A basic prototype typically costs ₹8L–₹15L. A production application with hybrid search and access control runs ₹15L–₹35L. Enterprise RAG with permission-aware retrieval and multi-tenancy typically costs ₹35L–₹80L or more.

Why does my RAG system give wrong or irrelevant answers?+

Almost always a retrieval problem, not a model problem: poor chunking, missing metadata filters, insufficient reranking, or stale data from a one-time ingestion. Improving the retriever and data pipeline fixes far more issues than switching LLMs.

What's the difference between vector search and hybrid search?+

Vector (semantic) search matches meaning even without shared words, but can miss exact terms like product codes or names. Hybrid search combines vector similarity with keyword search (BM25), typically outperforming either alone for real enterprise content.

How do you keep a RAG system secure for enterprise use?+

Through permission-aware retrieval that filters results by the requesting user's actual access rights, document-level RBAC, tenant isolation for multi-customer systems, and audit logging — covered in depth in our enterprise RAG architecture guide.

How long does it take to build a RAG application?+

A basic prototype typically takes 4–8 weeks. A production-grade application with proper retrieval tuning and evaluation typically takes 10–16 weeks. Enterprise RAG with security and multi-tenancy can take 4–7 months.

Do I need a vector database, or can I use my existing database?+

PostgreSQL with the pgvector extension handles vector search well for many workloads without adding a new system. Dedicated vector databases (Pinecone, Qdrant) make sense at larger scale or when you need specialized indexing features they provide.

Written by

CodeSurge AI Engineering Team

The CodeSurge AI team designs and builds AI systems, SaaS products and enterprise integrations for clients in India, the UAE and beyond — this section shares the architecture patterns, cost drivers and implementation tradeoffs we work through on real projects.

AI EngineeringEnterprise ArchitectureSaaSCloudSoftware Development

Found this useful? Share it with your team.

Share
Keep reading

Related insights

Talk to CodeSurge AI