Skip to content

feature: Add retention-based partition dropping - #33

Draft
codybswaney wants to merge 1 commit into
masterfrom
claude/partition-retention
Draft

codybswaney wants to merge 1 commit into
masterfrom
claude/partition-retention

Conversation

@codybswaney

Copy link
Copy Markdown

Summary

Adds drop_partitions and an opt-in maintain --drop-expired to prune partitions that fall outside a table's retention window.

  • Retention comes from the settings comment as retention:<N>: the current period plus the N periods before it are kept. Tables without it are never pruned. A malformed value (0, -1, 3m, …) is ignored, so a typo can't widen what gets dropped.
  • Two opt-ins: a table is only pruned when its comment has a retention and the job passes --drop-expired (default off).
  • Partitions are selected by bounds, not names: a partition is expired when its range ends at or before the cutoff. Legacy names and irregular layouts (e.g. a year-resetting weekly scheme with a one-day stub) work.
  • Never dropped: DEFAULT, MINVALUE/MAXVALUE, and sub-partitioned children. Nothing is dropped with CASCADE, so a dependent view or other object makes the drop fail loudly.
  • Short locks: dropping a partition takes ACCESS EXCLUSIVE on the parent. Each partition is dropped oldest-first in its own transaction under lock_timeout (shares --lock-timeout, default 5s), and the partition is re-checked inside that transaction before the drop.
    • On a timeout the run stops, so the retained partitions stay contiguous.
    • DropPartitionsError.dropped reports what was already removed; the rest is retried on the next run.
  • maintain extends before it prunes, so a failed drop never costs runway. JSONL records gain partitions.dropped and droppedPartitions (names), and the start record gains dropExpired.
  • --dry-run on drop_partitions lists what would be dropped.

Validation

  • Test suite: 428 tests pass on Postgres 13.20 and 18. That covers the existing suite plus new tests for dropping, the command, maintain integration, settings parsing, and the cutoff math.
  • End-to-end with the built CLI against committed tables on 13.20 and 18, with concurrent sessions:
    • Lock contention: with an open transaction reading the parent, drop_partitions --lock-timeout 1s gave up after ~1s, dropped nothing and exited 1. A writer arriving meanwhile waited ~0.8s, so writer delay is capped by the timeout.
    • Long read on an expired leaf: the drop backs off the same way.
    • Logical replication: pgoutput slots in both publish_via_partition_root modes, one with an actively streaming consumer during maintain --drop-expired.
      • The drop produced an empty transaction in the stream: no DELETEs.
      • Updates and deletes made to rows in the dropped partitions before the drop were still delivered afterwards.
      • Slots stayed healthy and kept advancing, and publication membership shrank automatically.
      • Grants on surviving partitions were intact, and every surviving leaf was still replica-identity safe.
  • Retention cutoffs matched expectations: monthly retention:3 kept 2026-06 onward on 2026-09-22; weekly retention:4 kept the ISO week of 2026-08-24 onward.

Notes / follow-ups

  • The README documents that dropped rows stay in downstream logical-replication copies and can't be recovered by a later re-sync from the source.
  • Concurrent detach: DETACH PARTITION … CONCURRENTLY (PG14+) would avoid the ACCESS EXCLUSIVE lock on the parent. It can't run in a transaction block, so it needs a non-transactional test harness. Left for a follow-up.
  • No per-run cap: turning retention on for a long-lapsed table drops every expired partition in one run, one short transaction each.

🤖 Generated with Claude Code

Adds drop_partitions (and an opt-in maintain --drop-expired) to prune
partitions outside a table's retention window. Retention is read from the
settings comment as `retention:<N>`: the current period plus the N before
it are kept. Tables without a retention are never pruned, and a malformed
value is ignored rather than widening what gets dropped.

Partitions are selected by bounds (range ends at or before the cutoff), so
legacy-named and irregular (e.g. year-resetting weekly) layouts work.
DEFAULT, MINVALUE/MAXVALUE and sub-partitioned children are never dropped,
and nothing is dropped with CASCADE.

Dropping a partition takes ACCESS EXCLUSIVE on the parent, so each
partition is dropped oldest-first in its own transaction under
lock_timeout (default 5s, shared with --lock-timeout). A timeout stops the
run, keeping retained coverage contiguous; DropPartitionsError reports
what was already dropped and the rest is retried on the next run. In
maintain, tables are extended before they are pruned so a failed drop
never costs runway, and the JSONL records gain partitions.dropped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant