---
title: "Data retention and deletion in AI systems | Dhruv Doshi"
description: "AI systems are data copies all the way down: training sets, vector indexes, caches, logs, evaluation sets, and backups. Each copy has its own lifecycle, and deletion requests — from…"
canonical: https://doshidhruv.com/notes/data-retention-and-deletion-in-ai-systems/
author: Dhruv Doshi
---

[Skip to content](#main-content)

[Dhruv Doshi](/)

[Work](/projects)[Experience](/resume)[Notes](/notes)[Guides](/guides)[Research](/research)[Investing](/investing)[About](/about)[Contact](/contact)[Search /](/search)dark mode

[All notes](/notes)

[AI governance](/topics/ai-governance)

# Data retention and deletion in AI systems

2026-10-01 · 4 minute read

AI systems are data copies all the way down: training sets, vector indexes, caches, logs, evaluation sets, and backups. Each copy has its own lifecycle, and deletion requests — from privacy regulation, from contract terms, or from plain operational hygiene — must reach all of them. Most organizations discover the gaps during an incident rather than during design.

## Map every copy before you need to delete

The first step is an inventory of where data lives in AI form. Source documents and their parsed derivatives. Chunked text and embeddings in the vector store. Prompt and completion logs, including any stored by the model provider. Caches at the application, gateway, and CDN layers. Evaluation datasets derived from production data. Model checkpoints and fine-tuning datasets. Backups of all of the above.

For each copy, record the owner, the retention policy, the deletion mechanism, and the time objective for deletion to take effect. A map that nobody maintains is decoration; tie it to the system inventory so new data stores are registered when they are created, not when they are audited.

## Deletion must propagate through derivatives

Deleting the source document is the easy part. The hard part is the derivatives: the chunks, the embeddings, the cached answers generated from it, the evaluation cases built on it, and the replicas in other regions. A deletion process that stops at the source leaves the information fully retrievable through the index.

Define deletion as an event that flows through the pipeline. When a source is removed or its access changes, derived artifacts must be updated or purged within a defined objective. This needs to be engineered into the ingestion and indexing path — a batch job that runs monthly is not deletion, it is eventual embarrassment.

Pay special attention to backups. A backup that restores deleted data reintroduces the liability. Backup retention windows should be explicit, documented, and justifiable, and restore procedures should respect deletion markers.

## Retention policy is a design input

Not all data should be kept as long as storage allows. Prompt and completion logs are particularly sensitive: they contain user inputs verbatim, potentially including confidential or personal information the user pasted into a chat box. Default retention should be short, with longer retention requiring explicit justification and stronger access controls.

Separate retention by purpose. Operational debugging needs recent logs. Quality evaluation needs sampled, consented data with its own lifecycle. Training and fine-tuning need governed datasets with provenance. Each purpose gets its own policy, and data collected for one purpose is not silently reused for another.

## Handle the model itself honestly

Once data has influenced model weights through fine-tuning, deletion becomes genuinely difficult. The practical controls are preventive: curate fine-tuning data carefully, exclude data subject to deletion requests from training sets, and prefer retrieval-based approaches — where deletion means removing index entries — over baking volatile data into weights.

Be honest in documentation about what deletion can and cannot achieve for each system type. A RAG system can offer strong deletion guarantees because knowledge lives in the index. A fine-tuned model cannot offer the same guarantee, and claiming otherwise creates legal and trust risk.

## Make it routine, not exceptional

Deletion requests will arrive as privacy tickets, customer demands, and internal cleanups. The organizations that handle them well treat deletion as a normal operational capability with runbooks, objectives, and tests — including periodic verification that deleted data is actually gone from every copy. The ones that handle them badly treat each request as a novel crisis.

Design the data lifecycle before the first deletion request arrives. It is one of the few areas where the cost of retrofitting exceeds the cost of doing it right by an order of magnitude.

## Verify deletion instead of assuming it

A deletion process that is never tested is a hope, not a control. Build verification into the lifecycle: periodic probes that attempt to retrieve supposedly deleted content through every access path — the application, the search index, the API, cached responses, and backups within their retention window. A probe that finds deleted data is a high-severity finding, not a curiosity.

Test the full propagation chain, not just the source. Delete a canary document and verify its chunks disappear from the vector index, its cached answers expire, its evaluation derivatives are flagged, and its replicas converge — all within the defined time objective. Automate the canary; run it continuously. Deletion pipelines break silently when an index rebuild skips the tombstone table or a new cache layer is added without the invalidation hook.

Document the results. Deletion verification logs are the evidence that turns a privacy commitment into something defensible — to a regulator, to a customer exercising their rights, or to your own leadership when they ask whether the process actually works. The organizations that can show a year of clean deletion probes sleep well. The ones that cannot are one audit away from discovering what their pipeline really does.

## Continue reading

- ["AI red teaming: a practical playbook for enterprise teams"](/notes/ai-red-teaming-a-practical-playbook-for-enterprise-teams) · Note
- ["Third-party AI vendor risk management"](/notes/third-party-ai-vendor-risk-management) · Note
- [Design APIs for agent consumers](/notes/design-apis-for-agent-consumers) · Note
- ["Audit logging for AI systems: what to record and why"](/notes/audit-logging-for-ai-systems-what-to-record-and-why) · Note

**Dhruv Doshi** · Toronto, Canada · [work@doshidhruv.com](mailto:work@doshidhruv.com)

[Resume](/resume)[Notes](/notes)[Guides](/guides)[Topics](/topics)[Search](/search)[Research](/research)[Investing](/investing)[LinkedIn](https://www.linkedin.com/in/dhruvdoshi25071999)[GitHub](https://github.com/DhruvDoshi)

[Sitemap](/sitemap.xml)[RSS](/feed.xml)[LLMs](/llms.txt)

© 2026 Dhruv Doshi
