RAG, retrieval-augmented generation, is powerful in demos. In production, it fails in predictable ways: poor retrieval, hallucinated citations, and inconsistent answers. Here is the checklist we use before shipping.
1. Retrieval quality
Chunking strategy, embedding choice, and reranking usually matter more than the LLM itself. Test retrieval in isolation before layering generation on top.
2. Citation and grounding
Every important claim should trace back to a source. Add strict citation rules and fallback behavior when confidence is low.
3. Evaluation pipeline
Build a dataset of real queries and expected behaviors. Track hallucination rate, retrieval recall, and answer consistency over time.
4. Observability
Log retrieval inputs and outputs, token usage, latency, and user feedback. You cannot improve a production system you cannot observe.
5. Guardrails
Validate inputs, filter outputs, and enforce permissions. Production RAG needs boundaries, not just prompt engineering.
We have shipped RAG systems for support copilots, internal search, and document-heavy workflows. The gap between demo and production is this checklist.