RAG systems change as documents, parsers, embeddings, indexes, models, permissions, and user questions evolve. Evaluation sets drawn from representative questions track retrieval relevance, citation support, no-answer behavior, and access boundaries. Retrieved documents are treated as untrusted content and tested for prompt injection. Monitoring and review help identify regressions without promising perfect answers or prevention of every costly error.