Skip to content

Data Retention Runbook

Use this runbook when reviewing customer handoff evidence, audit-log handling, RAG knowledge, restore reports, or coding-agent workspace data.

Policy

The retention policy lives in platform/governance/data-retention.yaml.

Default policy:

  • Gateway and RAG audit logs keep hashes, lengths, IDs, timing, status, usage, and result IDs. They must not store raw prompts, completions, or RAG queries.
  • Generated evidence is retained for release and audit review, with sample files committed and generated run files ignored.
  • RAG knowledge and vector-store collections require review before customer use.
  • Coding-agent workspace PVC data should be purged on tenant offboarding.
  • Model governance evidence has longer retention because it supports model lifecycle audit.

Workflow records and captured content

The gateway uses the existing budget Redis for these records. Configure positive integer gateway environment variables:

Variable Default Applies to
CONTENT_RETENTION_SECONDS 604800 (7 days) Each captured run/step input/output key, from its last write
CONTENT_MAX_BYTES 16384 (16 KB) Each captured input or output field, in UTF-8 bytes after redaction
RUN_RECORD_RETENTION_SECONDS 2592000 (30 days) Terminal run metadata, budget, timeline, capture markers, notifications, and start intent/input

captureContent: none|redacted|full is a team default and an optional per-workflow override in SandboxPolicySet. Capture defaults to none when both are absent; the Compose demo team uses redacted. Full capture still follows gateway admission and output DLP; redacted capture additionally runs the existing redactors on stored text. Receipts remain fingerprints and accounting metadata. Run readers must pass the existing authenticated team/project checks.

Content lives under PREFIX:workflow:TEAM:RUN:content:STEP_HASH, separately from the timeline and receipts. Each value contains text input/output, per-field truncated flags, and the redaction mode. After content expiry the API returns content: null with content_reason: expired; uncaptured steps return capture_off.

The run monitor checks Temporal every 30 seconds and applies a fixed expiry at Temporal's close time plus the run retention period for completed, failed, canceled, terminated, timed-out, and continued-as-new runs. Inspection also applies retention. Reads and duplicate worker initialization do not extend deadlines. Active runs do not get a run-record TTL. Shorter existing TTLs, including content TTLs, are preserved. Expired run IDs are pruned from project indexes during listing and monitoring, without breaking pagination.

An idempotent online state migration runs in the monitor after gateway startup. It scans existing run metadata in batches and asks Temporal for each execution's state, because older Redis metadata did not store terminal status. It backfills terminal TTLs without resetting active runs, audit chains, budgets for other runs, or identity/session keys. Already-old terminal records expire immediately. When Temporal has already purged a run, the migration starts its final Redis retention window at discovery. Transient failures leave the migration pending for retry; check gateway warnings, Redis, and Temporal health.

These TTLs do not erase Temporal history, backups, exported audit logs, or workflow side effects. Configure Temporal namespace retention and backup expiry separately. Changing retention settings affects new deadlines; existing earlier deadlines are never extended.

Validate Retention

Run:

make retention-check

Generate JSON and Markdown retention evidence:

make retention-report

Reports are written under results/retention/.

Customer Handoff

Before handoff, confirm:

  • audit logs do not contain raw prompt, completion, or query text
  • generated evidence is retained according to customer policy
  • RAG knowledge and vector-store collections have been approved for the environment
  • agent workspace PVCs have an offboarding and purge process
  • model governance reports are retained with model approval evidence

Erasing a RAG Source (Right-to-Erasure)

Ingestion is upsert-only, so removing a source from the manifest leaves its vectors in Qdrant. To purge a source's vectors (right-to-erasure or source decommission), delete by source_id:

# Purge across all collection versions:
python scripts/rag-ingest.py --delete --source-id <source-id> \
  --qdrant-url "$QDRANT_URL" --collection "$QDRANT_COLLECTION"

# Or scope the delete to one collection version:
python scripts/rag-ingest.py --delete --source-id <source-id> \
  --qdrant-url "$QDRANT_URL" --collection "$QDRANT_COLLECTION" --collection-version v2

This issues a filtered Qdrant points/delete on the source_id payload field written at ingest time. To re-index a source after a content change, run --delete --source-id <id> followed by --write. Record the deletion in the retention evidence for the environment.

Age-Based Retention Purge

Ingested chunks carry an ingestedAtEpoch timestamp, so the retentionDays policy can be enforced by purging points older than the retention window (run on a schedule, e.g. a CronJob):

python scripts/rag-ingest.py --purge --older-than-days 180 \
  --qdrant-url "$QDRANT_URL" --collection "$QDRANT_COLLECTION"

This deletes every point whose ingestedAtEpoch is older than the cutoff (optionally scoped to a --collection-version). Set --older-than-days to the retentionDays value from platform/governance/data-retention.yaml for the relevant retention class, and retain the purge summary as retention evidence.

Changing Retention

Tune retentionDays and classifications only through reviewed changes to platform/governance/data-retention.yaml. If a customer requires longer retention or stricter classification, update the policy first and regenerate make retention-report.