PBS: prune, verify and garbage collection that do not fight
In short. Prune removes snapshots by keep rules, per namespace. Garbage collection marks every chunk the remaining indexes reference, then deletes chunks nobody touched in the last 24 hours. Verify re-reads chunks and flags the ones that do not match their digest. Give each its own hour, keep them off the backup window, and keep prune out of the restore drill window, because a prune during a drill is how this estate lost the snapshot it was about to restore.
Backups are slow, verify is red, and a drill lost its snapshot
The symptoms arrive separately and have one cause. Backups that normally finish by 04:00 run until 07:00 on the night garbage collection runs. A verify job logs verification failed on snapshots that were fine last week, and the task list shows GC and verify started within minutes of each other. And during a restore drill the chosen snapshot disappeared between picking it and restoring it, because the daily prune ran at 10:00 and the drill was at 10:00.
What you need
- PBS 4.x with at least one datastore and one namespace per tenant.
- Root on the PBS host, or a user with
Datastore.Modifyon the datastore. - The backup windows of every cluster writing to this PBS. Here they are staggered between 01:00 and 04:00.
- Names below: datastore
backups, namespacestenant-xandplatform.
Steps
-
Understand what each job does before scheduling it. Prune applies keep rules to one group at a time and deletes snapshot directories; it is fast and touches no chunks. Garbage collection runs in two phases: a mark phase that walks every index in the datastore and touches each referenced chunk, then a sweep phase that deletes chunks whose last touch is older than 24 hours and a few minutes. The grace period exists because a backup in progress has uploaded chunks that no finished index references yet. Verify reads every chunk of a snapshot and compares it with its digest; it is pure I/O and takes hours on a full datastore.
-
Set the keep rules per namespace with prune jobs, not on the datastore. The live job on this estate keeps 14 daily, 8 weekly and 3 monthly for tenants, and 30 daily for platform internals:
proxmox-backup-manager prune-job create prune-tenants \ --store backups --ns tenant-x --schedule '06:00' \ --keep-daily 14 --keep-weekly 8 --keep-monthly 3 proxmox-backup-manager prune-job create prune-platform \ --store backups --ns platform --schedule '06:15' \ --keep-daily 30The keep rules are not additive. Each rule keeps the newest snapshot in each of its periods, and a snapshot already kept as a daily still counts as that week's weekly. The total kept is at most the sum of the numbers and usually less.
keep-lastis the one to add for a group with several snapshots a day. -
Create the verify job with
ignore-verifiedandoutdated-after, so it reads new snapshots every day and re-reads old ones every 30 days instead of everything every night:proxmox-backup-manager verify-job create verify-backups \ --store backups --schedule '09:00' \ --ignore-verified true --outdated-after 30 -
Schedule garbage collection weekly, in daylight, away from everything else:
proxmox-backup-manager datastore update backups --gc-schedule 'sat 12:00'GC on a multi-terabyte datastore takes hours and saturates the disks. Weekly is enough; the chunks it reclaims are only the ones freed by the last week's prunes.
-
Write the calendar down. This is the one the estate runs on each PBS; the target PBS has the sync job and its own copies of prune, verify and GC shifted by an hour:
01:00 to 04:00 backup jobs from the clusters, staggered per cluster 05:00 sync job on the target pulls from the source 06:00 prune, tenant namespaces (source) 06:15 prune, platform namespace (source) 07:00 prune, replica namespaces (target) 09:00 verify, ignore-verified, outdated-after 30 (both) sat 12:00 garbage collection (source) sun 12:00 garbage collection (target) wed 10:00-14:00 restore drill window: nothing scheduledThe schedule strings are systemd calendar syntax:
06:00is daily at six,sat 12:00is weekly,mon..fri 09:00is weekdays. -
Move anything that overlaps. The daily prune used to be at 10:00; the drill is on a Wednesday at 10:00. The prune moved to 06:00 and the drill window is now written into the calendar as a reserved slot. Check with
proxmox-backup-manager prune-job list,verify-job listanddatastore listthat nothing else lands in it.
Verify it worked
The next morning, open the task list and read the start and end time of each job:
proxmox-backup-manager task list --limit 50
Backups should finish inside their window; prune should start after the last backup ends; verify should not start until prune has finished. On Saturday, proxmox-backup-manager garbage-collection status backups shows the last run's duration and how much it removed. A week later, the verify job's log should list only new and outdated snapshots, not the whole store.
Gotchas
- Prune during a restore: the snapshot being restored is locked and survives, but the one you chose and have not opened yet does not. Reserve the drill window.
- GC during verify doubles the disk load and the verify starts reporting chunk load failures under contention. Those are not corruption, but you cannot tell from the log until it runs again cleanly.
- GC during a backup window does not break the backup; the 24-hour grace covers it. It makes the backup three times slower, which is how it gets noticed.
- Verify without
ignore-verifiedre-reads everything every day. It will not finish before the next backup window on a datastore of any size. - Prune jobs are per namespace, and a namespace with no prune job keeps everything forever. New tenant, new prune job; put it in the onboarding checklist.
FAQ
Why does GC keep chunks for 24 hours instead of checking what is in flight? The mark phase uses the chunk's access time as the marker, which is cheap and needs no coordination with backup clients. Anything a client uploaded in the last day is by definition recently touched and survives, whether or not its index is written yet. A datastore on a filesystem that does not update access times breaks this, which is why PBS checks for it when a datastore is created.
Can prune delete a snapshot that is protected?
No. A protected snapshot is skipped by prune and by remove-vanished, and its chunks are kept by GC. It is the right tool for evidence preservation and for the newest snapshot per group before a big verify.
Should the replica use the same keep rules as the source?
It should have its own, set deliberately. Here they happen to match. With remove-vanished off, the replica's prune is the only thing that ever removes a group the source has deleted, so its rules also decide how long a deleted guest's history outlives it.