PBS: a restore drill for one file, one guest and one site
In short. Restore one file through the file browser or
proxmox-backup-client restore, one guest withqmrestoreto a fresh VMID with its network disconnected, and one guest on the other site's cluster from the replica namespace through the read-only storage entry. Time each one, write the times next to the targets, and reserve the window so prune and garbage collection cannot run inside it. The first drill here found a prune race and a script that hid failures behind a pipe, which is exactly what a drill is for.
Has anyone restored from this?
The backup jobs are green, the sync is green, verify is green, and nobody has restored anything in a year. The written targets are a recovery point of 30 minutes for files, from the CephFS snapshots, and a recovery time of 4 hours for a guest. Neither number has been measured. This post is the drill that measures them.
flowchart LR
subgraph A["Site A"]
PBSA["PBS A<br/>ns tenant-x"]
PVEA["PVE cluster A"]
PBSA -- "1. one file<br/>file browser or client restore" --> F["scratch directory"]
PBSA -- "2. one guest<br/>qmrestore to new VMID" --> G["VM 9104, NIC down"]
end
subgraph B["Site B"]
PBSB["PBS B<br/>ns site-a/tenant-x"]
PVEB["PVE cluster B"]
PBSB -- "3. one guest, other site<br/>read-only storage entry" --> H["VM 9104 on cluster B"]
end
PBSB -. "pull sync" .-> PBSAWhat you need
- PBS 4.x at both sites, PVE 9.2 at both, the cross-site pull in place and verified.
- A reserved window: Wednesday 10:00 to 14:00 here, with no prune, verify or GC scheduled inside it.
- A VMID range nobody uses,
9000and up, and a storage with room for one more guest at each site. - A stopwatch, or the task timestamps, and somewhere to write results. Names below: datastore
backups, namespacetenant-x, guest104, archive grouphost/files-tenant-x.
Steps
-
Pick the targets in advance and write them down: which file, which guest, which snapshot. Then protect those snapshots so the drill cannot lose them:
export PBS_REPOSITORY=backups proxmox-backup-client snapshot protected update vm/104/2026-08-13T01:40:12Z true --ns tenant-x proxmox-backup-client snapshot protected update host/files-tenant-x/2026-08-13T03:35:12Z true --ns tenant-x -
Restore one file. Start the clock. In the PBS GUI, open the datastore, the namespace, the snapshot, and use the file browser on
data.pxarto download one named file. From the command line, restore to a scratch directory and compare:proxmox-backup-client restore host/files-tenant-x/2026-08-13T03:35:12Z data.pxar /tmp/drill --ns tenant-x diff /tmp/drill/<uuid>/reports/2026-08.csv /mnt/cephfs/volumes/tenant-x/files/<uuid>/reports/2026-08.csvIf the CephFS snapshot is still there,
cpfrom.snapis faster and should be timed too; it is the 30-minute recovery point in practice. -
Restore one guest to a new VMID on the local cluster. Find the volume id, restore with a fresh VMID, regenerate MAC addresses, and keep the network down until you have looked at it:
pvesm list pbs-a | grep 'vm/104/' qmrestore pbs-a:backup/vm/104/2026-08-13T01:40:12Z 9104 --storage ceph-rbd --unique qm set 9104 --net0 virtio,bridge=vmbr0,tag=110,link_down=1 qm start 9104Open the console, log in, check the application starts and the data is the date you expect. Stop the clock. Destroy it:
qm stop 9104 && qm destroy 9104 --purge. -
Restore the same guest on the other cluster from the replica. On a node of cluster B, the read-only storage entry points at
site-a/tenant-xon PBS B. The volume ids look local because the namespace is in the storage definition:pvesm list pbs-replica-a | grep 'vm/104/' qmrestore pbs-replica-a:backup/vm/104/2026-08-13T01:40:12Z 9104 --storage ceph-rbd --unique qm set 9104 --net0 virtio,bridge=vmbr0,tag=110,link_down=1 qm start 9104This is the number that matters: it is the recovery time if site A is gone. Include the minutes spent finding the storage entry and the VLAN tag on the other site; a drill that skips the fumbling measures the wrong thing.
-
For a container with a bind mount, the guest restore gives you the runtime and not the data. Restore the pxar into a fresh subvolume, point a scratch container at it, then swap the bind mount. That is a fourth timing and the longest one here.
-
Record it. One block per restore, appended to the drill log next to the runbooks:
drill 2026-08-13, window wed 10:00-14:00, operator: on-call one file, pxar, GUI browser: target n/a measured __:__ one file, cp from .snap: target RPO 0:30 measured snapshot age _:__ one guest, local, qmrestore: target RTO 4:00 measured __:__ one guest, other site, replica: target RTO 4:00 measured __:__ found: prune removed chosen snapshot; verify script hid failureThe targets come from the written RPO and RTO; the measured column is filled in during the drill, never ahead of it.
-
Unprotect the snapshots and schedule the next drill: quarterly, rotating the tenant namespace, with the same window reserved and checked against
proxmox-backup-manager prune-job list,verify-job listanddatastore list.
Verify it worked
The drill verified itself if each restore produced a thing you could open: the file matched its original, the guest booted with the right data, the guest on the other cluster did the same from the replica. Beyond that, the measured times sit under the written targets, the drill log has a dated entry, and the task list on both PBS hosts shows no prune or GC inside the window. A restore that failed is a successful drill with a finding, and the finding becomes an issue with an owner.
Gotchas
- The first drill here chose a snapshot at 09:50 and the prune job ran at 10:00. By the time the restore started the snapshot was gone. Prune now runs at 06:00 and the window is reserved; step 1 protects the targets as well.
- The verification script that checked the archive did
proxmox-backup-client ... | grep -c pxarand reported green when the client failed, because the pipe's exit status was grep's.set -o pipefailat the top of every script, and a test that makes the first command fail on purpose. qmrestorewithout--uniquebrings the original MAC addresses up on the same VLAN as the live guest. Always--unique, alwayslink_down=1until you have looked.- The VLAN tag and bridge name differ between sites. Write them in the runbook so the other-site restore does not start with a search.
- A container restore (
pct restore) from the replica needs the bind-mount directory to exist on the other site's CephFS, or the restore has to drop the mount point. Decide which before the drill.
FAQ
How often?
Quarterly for the full three, monthly for the one-file restore because it is ten minutes and it exercises the file archive that vzdump does not cover. After any change to the PBS hosts, the sync, or the storage entries, run the other-site restore once.
Why a new VMID and not restore over the original? Restoring over the live guest is the production procedure for a real incident and the drill should never be the thing that causes one. A new VMID in a reserved range, with its NIC down, cannot collide with anything.
What about the evidence case, where a guest must be preserved rather than restored?
Same primitives. A CephFS snapshot of the exported artefacts plus a PBS backup marked protected, which prune, remove-vanished and garbage collection all leave alone. Read one back during the drill so the preservation procedure is rehearsed too.