Skip to content
BDOT SOFTWAREBDOT Software

Insights / AI

RAG needs provenance, evaluation, and a way to abstain

Retrieval-augmented generation is a data pipeline as much as a prompt pattern. Trust depends on what was retrieved, who could access it, and when the system should not answer.

Kiran Bandarupalli · 2 Oct 2026 · 2 min read

Engineering notes beside a laptop, representing source material and review in a retrieval workflow

A retrieval demo often begins with a folder, a vector index, and a question that happens to match a paragraph. A dependable feature must also handle document updates, access restrictions, ambiguous questions, and cases where the source material does not support an answer.

Retrieval is its own system

Document parsing, chunk boundaries, metadata, embeddings, indexing, and filters all affect what the model sees. Store stable identifiers and source versions with each chunk. When a document is replaced, make the index update atomic enough that users do not see a confusing mixture of old and new material.

Enforce permissions before generation

Filter retrieval by the current user's access scope before sending passages to a model. Hiding a forbidden citation in the interface does not undo disclosure to the model provider. Test authorization with users who have deliberately different access, including revoked and expired permissions.

Preserve provenance

Keep the document ID, version, and passage location that produced each retrieved result. Show sources where a reviewer can inspect them. A citation should refer to material actually provided to the model, not a plausible-looking URL composed after the answer.

Evaluate more than fluency

Maintain a small set of representative questions with expected relevant sources and acceptable answer behavior. Evaluate retrieval recall separately from answer grounding. Include stale documents, conflicting sources, typo-heavy queries, unanswerable questions, and prompt-injection text embedded in documents. Re-run the set when parsing, chunking, embeddings, models, or prompts change.

Know when to stop

Define a minimum evidence threshold and an explicit abstention response. A system can ask for clarification, return the best sources without summarizing, or route the case to a person. Do not solve uncertainty by raising model temperature or writing a stronger instruction.

Production RAG is not a single prompt. It is a versioned data pipeline with permission checks, inspectable evidence, measurable retrieval behavior, and a safe answer for the cases it cannot support.