Table of contents
A production AI system can be fully up — every service healthy, every endpoint responding — while quietly giving wrong answers to half its users. Traditional application monitoring (uptime, error rates, response times) doesn't catch this, because the failure isn't a crash, it's a quality problem inside a response that returned successfully. LLM observability is the discipline of monitoring for that category of failure specifically.
This guide covers what to actually monitor in a production AI system beyond standard infrastructure metrics, how tracing works across a multi-step AI request, the difference between observability and evaluation, and how to build an alerting strategy for problems that don't look like outages.
- What standard APM misses
- Answer quality, groundedness, retrieval accuracy
- Most valuable single signal
- Direct user feedback (thumbs up/down, corrections)
- Observability vs. evaluation
- Continuous monitoring vs. periodic structured testing
- Where to start
- Log every query, retrieval and response — you can't fix what you can't see
Why LLM observability is a different discipline
Standard application performance monitoring answers questions like: is the service up, how fast does it respond, what's the error rate. These matter for AI systems too, but they don't answer the question that actually determines whether the system is working: was the answer any good?
An LLM call can return a 200 status code, respond in 400 milliseconds, and still be confidently, fluently wrong — retrieving the wrong context, hallucinating a detail, or answering a question the user didn't actually ask. None of that shows up as an error in traditional monitoring. This is why LLM observability needs its own layer, purpose-built for the specific ways these systems fail.
What to actually monitor
| Signal | What it tells you |
|---|---|
| Token usage and cost | Cost per query, cost trends, and which query patterns are expensive |
| Latency by stage | Where time is actually spent — retrieval, reranking, generation — not just total response time |
| Retrieval logs | Exactly what context was retrieved for each query, essential for debugging bad answers |
| Groundedness / faithfulness | Whether the answer is actually supported by retrieved context or fabricated |
| Citation accuracy | Whether cited sources actually support the specific claim made |
| Error and fallback rates | How often the system fails to answer, times out, or falls back to a default response |
| User feedback | Direct signal — thumbs up/down, corrections, follow-up questions indicating confusion |
| Drift over time | Whether quality is degrading as underlying data, models or usage patterns change |
Tracing a request end to end
A single AI response typically involves multiple steps — query processing, retrieval, reranking, prompt construction, generation, guardrail validation. When an answer is wrong, the question is which step introduced the problem, and answering that requires tracing the full path, not just logging the final input and output.
Query received
The original user question, with a unique trace ID
Query processing
Any rewriting or expansion applied, and why
Retrieval
Exactly which chunks were retrieved, with relevance scores
Reranking
How candidates were re-ordered before reaching the prompt
Prompt construction
The final assembled context sent to the model
Generation
The model's raw response, token count, and latency
Guardrail checks
What validation ran, and whether it passed or modified the output
Response delivered
The final answer, plus the full trace for later debugging
Observability vs. evaluation: different tools for different jobs
These two terms get used interchangeably, but they answer different questions and both are needed.
Evaluation is periodic, structured testing against a maintained set of representative questions — usually run before launch and again whenever the pipeline changes. It answers: "does this system perform acceptably against known cases?" This is covered in depth in our RAG development guide and enterprise RAG architecture guide.
Observability is continuous monitoring of real production traffic. It answers: "how is the system actually performing right now, on real questions we didn't anticipate?" Evaluation catches regressions you can predict and test for. Observability catches the failure modes you didn't think to test — which, in practice, is most of them.
Alerting strategy
Standard alerting (error rate spikes, latency thresholds) still applies, but AI systems need alerts tuned to quality signals too: a sudden drop in average groundedness score, a spike in negative user feedback, an increase in "I don't know" or fallback responses, or a cost spike suggesting the system is retrieving far more context than usual. These thresholds should be based on your own system's baseline behavior, established once you have enough production traffic to know what "normal" looks like — not a generic industry number.
Building this into an agent, not just a RAG pipeline
Everything above applies directly to AI agents as well — in fact, observability matters more for agents, since a multi-step agent has more places a failure can originate, and a wrong tool call can have real-world consequences beyond a bad chat response. See our AI agent development guide for how observability fits into the broader agent architecture, including audit logging for actions taken, not just answers given.
- Every query, retrieval and response logged with a shared trace ID
- Token usage and cost tracked per query, not just in aggregate
- Latency broken down by pipeline stage, not just total response time
- A mechanism for capturing direct user feedback on answer quality
- Groundedness/faithfulness checks logged, not just pass/fail at the guardrail
- Alerting thresholds set from your own system's baseline, reviewed periodically
- A clear, fast path from "user reported a bad answer" to "engineer can see the full trace"
Running an AI system in production?
Talk to our engineering team about observability, evaluation and reliability for your AI deployment.
Frequently asked questions
What is LLM observability?+
Monitoring purpose-built for AI systems — tracking not just uptime and latency, but answer quality signals like groundedness, retrieval accuracy, citation correctness and user feedback, which standard application monitoring doesn't capture.
How is LLM observability different from standard application monitoring?+
Standard monitoring tells you if a service is up and how fast it responds. LLM observability tells you whether the actual content of responses is accurate and grounded — a system can be fully healthy by standard metrics while giving wrong answers.
What's the difference between observability and evaluation for AI systems?+
Evaluation is periodic structured testing against a known set of questions, usually before launch and after pipeline changes. Observability is continuous monitoring of real production traffic, catching failure modes you didn't anticipate or test for.
What should I monitor first in a new AI system?+
Start by logging every query, retrieval and response with a shared trace ID — you can't diagnose quality problems you can't see. Add token cost, latency-by-stage, and a user feedback mechanism next.
How do you detect when an AI system's quality is degrading?+
Through drift tracking — comparing current groundedness scores, user feedback rates and fallback frequency against an established baseline — combined with re-running your evaluation set periodically, especially after data or model changes.
Does observability matter more for RAG systems or AI agents?+
Both need it, but agents arguably need it more — a multi-step agent has more places a failure can originate, and a wrong action (not just a wrong answer) can have real consequences, making audit-quality tracing essential.
Written by
CodeSurge AI Engineering Team
The CodeSurge AI team designs and builds AI systems, SaaS products and enterprise integrations for clients in India, the UAE and beyond — this section shares the architecture patterns, cost drivers and implementation tradeoffs we work through on real projects.