PBS: prune, verify and garbage collection that do not fight

In short. Prune removes snapshots by keep rules, per namespace. Garbage collection marks every chunk the remaining indexes reference, then deletes chunks nobody touched in the last 24 hours. Verify re-reads chunks and flags the ones that do not match their digest. Give each its own hour, keep them off the backup window, and keep prune out of the restore drill window, because a prune during a drill is how this estate lost the snapshot it was about to restore.

Backups are slow, verify is red, and a drill lost its snapshot

The symptoms arrive separately and have one cause. Backups that normally finish by 04:00 run until 07:00 on the night garbage collection runs. A verify job logs verification failed on snapshots that were fine last week, and the task list shows GC and verify started within minutes of each other. And during a restore drill the chosen snapshot disappeared between picking it and restoring it, because the daily prune ran at 10:00 and the drill was at 10:00.

What you need

Steps

  1. Understand what each job does before scheduling it. Prune applies keep rules to one group at a time and deletes snapshot directories; it is fast and touches no chunks. Garbage collection runs in two phases: a mark phase that walks every index in the datastore and touches each referenced chunk, then a sweep phase that deletes chunks whose last touch is older than 24 hours and a few minutes. The grace period exists because a backup in progress has uploaded chunks that no finished index references yet. Verify reads every chunk of a snapshot and compares it with its digest; it is pure I/O and takes hours on a full datastore.

  2. Set the keep rules per namespace with prune jobs, not on the datastore. The live job on this estate keeps 14 daily, 8 weekly and 3 monthly for tenants, and 30 daily for platform internals:

    proxmox-backup-manager prune-job create prune-tenants \
      --store backups --ns tenant-x --schedule '06:00' \
      --keep-daily 14 --keep-weekly 8 --keep-monthly 3
    proxmox-backup-manager prune-job create prune-platform \
      --store backups --ns platform --schedule '06:15' \
      --keep-daily 30
    

    The keep rules are not additive. Each rule keeps the newest snapshot in each of its periods, and a snapshot already kept as a daily still counts as that week's weekly. The total kept is at most the sum of the numbers and usually less. keep-last is the one to add for a group with several snapshots a day.

  3. Create the verify job with ignore-verified and outdated-after, so it reads new snapshots every day and re-reads old ones every 30 days instead of everything every night:

    proxmox-backup-manager verify-job create verify-backups \
      --store backups --schedule '09:00' \
      --ignore-verified true --outdated-after 30
    
  4. Schedule garbage collection weekly, in daylight, away from everything else:

    proxmox-backup-manager datastore update backups --gc-schedule 'sat 12:00'
    

    GC on a multi-terabyte datastore takes hours and saturates the disks. Weekly is enough; the chunks it reclaims are only the ones freed by the last week's prunes.

  5. Write the calendar down. This is the one the estate runs on each PBS; the target PBS has the sync job and its own copies of prune, verify and GC shifted by an hour:

    01:00 to 04:00  backup jobs from the clusters, staggered per cluster
    05:00           sync job on the target pulls from the source
    06:00           prune, tenant namespaces (source)
    06:15           prune, platform namespace (source)
    07:00           prune, replica namespaces (target)
    09:00           verify, ignore-verified, outdated-after 30 (both)
    sat 12:00       garbage collection (source)
    sun 12:00       garbage collection (target)
    wed 10:00-14:00 restore drill window: nothing scheduled
    

    The schedule strings are systemd calendar syntax: 06:00 is daily at six, sat 12:00 is weekly, mon..fri 09:00 is weekdays.

  6. Move anything that overlaps. The daily prune used to be at 10:00; the drill is on a Wednesday at 10:00. The prune moved to 06:00 and the drill window is now written into the calendar as a reserved slot. Check with proxmox-backup-manager prune-job list, verify-job list and datastore list that nothing else lands in it.

Verify it worked

The next morning, open the task list and read the start and end time of each job:

proxmox-backup-manager task list --limit 50

Backups should finish inside their window; prune should start after the last backup ends; verify should not start until prune has finished. On Saturday, proxmox-backup-manager garbage-collection status backups shows the last run's duration and how much it removed. A week later, the verify job's log should list only new and outdated snapshots, not the whole store.

Gotchas

FAQ

Why does GC keep chunks for 24 hours instead of checking what is in flight? The mark phase uses the chunk's access time as the marker, which is cheap and needs no coordination with backup clients. Anything a client uploaded in the last day is by definition recently touched and survives, whether or not its index is written yet. A datastore on a filesystem that does not update access times breaks this, which is why PBS checks for it when a datastore is created.

Can prune delete a snapshot that is protected? No. A protected snapshot is skipped by prune and by remove-vanished, and its chunks are kept by GC. It is the right tool for evidence preservation and for the newest snapshot per group before a big verify.

Should the replica use the same keep rules as the source? It should have its own, set deliberately. Here they happen to match. With remove-vanished off, the replica's prune is the only thing that ever removes a group the source has deleted, so its rules also decide how long a deleted guest's history outlives it.

Related