---
title: "RAG failure modes in production and how to contain them | Dhruv Doshi"
description: "Retrieval augmented generation fails in production in ways that never appear in a demo. The demo uses a clean corpus, a cooperative user, and a forgiving evaluator. Production…"
canonical: https://doshidhruv.com/notes/rag-failure-modes-in-production-and-how-to-contain-them/
author: Dhruv Doshi
---

[Skip to content](#main-content)

[Dhruv Doshi](/)

[Work](/projects)[Experience](/resume)[Notes](/notes)[Guides](/guides)[Research](/research)[Investing](/investing)[About](/about)[Contact](/contact)[Search /](/search)dark mode

[All notes](/notes)

[AI governance](/topics/ai-governance)

# RAG failure modes in production and how to contain them

2026-10-01 · 5 minute read

Retrieval-augmented generation fails in production in ways that never appear in a demo. The demo uses a clean corpus, a cooperative user, and a forgiving evaluator. Production contributes stale documents, ambiguous questions, permission boundaries, adversarial inputs, and users who trust fluent text. Understanding the failure taxonomy is the prerequisite for containing it.

## Retrieval misses

The most common failure is also the least visible: the retriever does not return the chunk that contains the answer, and the model answers anyway from general knowledge. The response reads well and is wrong, or right for the wrong reasons.

Contain it with retrieval evidence requirements. The application should know, for each generated claim, which retrieved chunks support it. When required evidence is missing, the system must say so or abstain rather than synthesize. Measure retrieval recall on a labeled set per document type; a single aggregate number hides that the retriever works for FAQs and fails for policy documents.

Hybrid retrieval — combining keyword and vector search — consistently outperforms either alone on enterprise corpora where exact terms matter. Reranking the candidate set with a cross-encoder before generation is usually the highest-leverage quality improvement available.

## Stale and contradictory sources

Enterprise corpora contain multiple versions of the truth: the old policy and the new policy, the draft and the approved standard, the regional variant and the global one. A RAG system that retrieves all of them will happily cite the outdated version.

Containment starts at ingestion. Every chunk needs effective dates, version identity, supersession links, and authority ranking. Retrieval should prefer current, authoritative sources and the application should surface the version it relied on. Deletion and updates must propagate to derived chunks, embeddings, and caches within a defined objective; otherwise the vector index becomes an uncontrolled copy of information the organization has already retracted.

## Permission leakage

A vector index is a copy of your documents optimized for similarity search. If ingestion strips access metadata, the retriever will happily serve a confidential document to an unauthorized user wrapped in a fluent summary.

Every indexed chunk must carry the access policy of its source, evaluated at query time against the requesting principal. Test this with adversarial queries from low-privilege test accounts, not just with the happy path. Permission-aware retrieval is not a feature to add later; retrofitting it onto an index built without source identity is a rebuild.

## Chunking pathologies

Bad chunking produces failures that look like model failures. A table split across chunks loses its headers. A procedure split mid-step loses its preconditions. A chunk of boilerplate with no substantive content dilutes every query it matches.

Fix chunking before tuning prompts. Parse documents into their natural structure — sections, tables, procedures — and chunk along those boundaries with enough overlap to preserve context. Keep tables intact or convert them to a structured representation the model can actually read. Log which chunks supported each answer so chunking problems show up in production telemetry instead of user complaints.

## Embedding drift and index rot

Embeddings are generated by a model, and models change. When the embedding model is updated, old vectors and new queries live in different spaces, and retrieval quality degrades silently. Similarly, an index that grows without curation accumulates near-duplicates and dead content that crowd out good results.

Version the embedding model with the index. Re-embed on model change or maintain the old model for the old index — but do not mix them. Monitor retrieval quality continuously with a fixed probe set; a drop in recall on probes is the early warning for index rot.

## Evaluation gaps

Teams evaluate the generator and forget the retriever, or evaluate retrieval with the same data they tuned it on. Both produce confidence without evidence.

Evaluate the full path: retrieval recall and precision on held-out queries, faithfulness of claims to cited chunks, abstention rate when evidence is insufficient, and permission correctness under adversarial probing. Keep a regression set of historically observed failures and run it on every change to the model, prompt, index, chunking, or embedding configuration.

## Designing for containment

No RAG system will be failure-free. The design goal is bounded failure: every failure mode has a defined detection mechanism, a containment behavior, and an owner. Retrieval misses trigger abstention. Stale sources trigger version display. Permission questions trigger denial. When the system cannot meet its evidence contract, the correct output is an explicit refusal with a path to a human — not a confident fabrication.

## Turn every failure into a permanent test

The highest-leverage response to a RAG failure is not the fix — it is the regression case. Every confirmed failure mode should become a standing test: the query that retrieved the stale policy, the ambiguous request that produced a confident fabrication, the low-privilege probe that reached restricted content. Run the full set on every change to the model, prompt, index, chunking strategy, or embedding configuration.

Over time this suite becomes the system's immune system and its institutional memory. New team members learn the failure taxonomy by reading the tests. Auditors see evidence rather than assertions. And the organization stops re-learning the same lessons every time the corpus grows or the model changes.

Pair the regression suite with production sampling. A fixed probe set run continuously against production catches silent degradation — index rot, embedding drift, corpus staleness — before users do. The combination is simple: the regression suite guards against known failures, the probe set watches for new ones, and the incident process converts new ones into regression cases.

## Continue reading

- ["AI red teaming: a practical playbook for enterprise teams"](/notes/ai-red-teaming-a-practical-playbook-for-enterprise-teams) · Note
- ["Third-party AI vendor risk management"](/notes/third-party-ai-vendor-risk-management) · Note
- [Design APIs for agent consumers](/notes/design-apis-for-agent-consumers) · Note
- ["Audit logging for AI systems: what to record and why"](/notes/audit-logging-for-ai-systems-what-to-record-and-why) · Note

**Dhruv Doshi** · Toronto, Canada · [work@doshidhruv.com](mailto:work@doshidhruv.com)

[Resume](/resume)[Notes](/notes)[Guides](/guides)[Topics](/topics)[Search](/search)[Research](/research)[Investing](/investing)[LinkedIn](https://www.linkedin.com/in/dhruvdoshi25071999)[GitHub](https://github.com/DhruvDoshi)

[Sitemap](/sitemap.xml)[RSS](/feed.xml)[LLMs](/llms.txt)

© 2026 Dhruv Doshi
