PBS: a restore drill for one file, one guest and one site

In short. Restore one file through the file browser or proxmox-backup-client restore, one guest with qmrestore to a fresh VMID with its network disconnected, and one guest on the other site's cluster from the replica namespace through the read-only storage entry. Time each one, write the times next to the targets, and reserve the window so prune and garbage collection cannot run inside it. The first drill here found a prune race and a script that hid failures behind a pipe, which is exactly what a drill is for.

Has anyone restored from this?

The backup jobs are green, the sync is green, verify is green, and nobody has restored anything in a year. The written targets are a recovery point of 30 minutes for files, from the CephFS snapshots, and a recovery time of 4 hours for a guest. Neither number has been measured. This post is the drill that measures them.

flowchart LR
  subgraph A["Site A"]
    PBSA["PBS A<br/>ns tenant-x"]
    PVEA["PVE cluster A"]
    PBSA -- "1. one file<br/>file browser or client restore" --> F["scratch directory"]
    PBSA -- "2. one guest<br/>qmrestore to new VMID" --> G["VM 9104, NIC down"]
  end
  subgraph B["Site B"]
    PBSB["PBS B<br/>ns site-a/tenant-x"]
    PVEB["PVE cluster B"]
    PBSB -- "3. one guest, other site<br/>read-only storage entry" --> H["VM 9104 on cluster B"]
  end
  PBSB -. "pull sync" .-> PBSA

What you need

Steps

  1. Pick the targets in advance and write them down: which file, which guest, which snapshot. Then protect those snapshots so the drill cannot lose them:

    export PBS_REPOSITORY=backups
    proxmox-backup-client snapshot protected update vm/104/2026-08-13T01:40:12Z true --ns tenant-x
    proxmox-backup-client snapshot protected update host/files-tenant-x/2026-08-13T03:35:12Z true --ns tenant-x
    
  2. Restore one file. Start the clock. In the PBS GUI, open the datastore, the namespace, the snapshot, and use the file browser on data.pxar to download one named file. From the command line, restore to a scratch directory and compare:

    proxmox-backup-client restore host/files-tenant-x/2026-08-13T03:35:12Z data.pxar /tmp/drill --ns tenant-x
    diff /tmp/drill/<uuid>/reports/2026-08.csv /mnt/cephfs/volumes/tenant-x/files/<uuid>/reports/2026-08.csv
    

    If the CephFS snapshot is still there, cp from .snap is faster and should be timed too; it is the 30-minute recovery point in practice.

  3. Restore one guest to a new VMID on the local cluster. Find the volume id, restore with a fresh VMID, regenerate MAC addresses, and keep the network down until you have looked at it:

    pvesm list pbs-a | grep 'vm/104/'
    qmrestore pbs-a:backup/vm/104/2026-08-13T01:40:12Z 9104 --storage ceph-rbd --unique
    qm set 9104 --net0 virtio,bridge=vmbr0,tag=110,link_down=1
    qm start 9104
    

    Open the console, log in, check the application starts and the data is the date you expect. Stop the clock. Destroy it: qm stop 9104 && qm destroy 9104 --purge.

  4. Restore the same guest on the other cluster from the replica. On a node of cluster B, the read-only storage entry points at site-a/tenant-x on PBS B. The volume ids look local because the namespace is in the storage definition:

    pvesm list pbs-replica-a | grep 'vm/104/'
    qmrestore pbs-replica-a:backup/vm/104/2026-08-13T01:40:12Z 9104 --storage ceph-rbd --unique
    qm set 9104 --net0 virtio,bridge=vmbr0,tag=110,link_down=1
    qm start 9104
    

    This is the number that matters: it is the recovery time if site A is gone. Include the minutes spent finding the storage entry and the VLAN tag on the other site; a drill that skips the fumbling measures the wrong thing.

  5. For a container with a bind mount, the guest restore gives you the runtime and not the data. Restore the pxar into a fresh subvolume, point a scratch container at it, then swap the bind mount. That is a fourth timing and the longest one here.

  6. Record it. One block per restore, appended to the drill log next to the runbooks:

    drill 2026-08-13, window wed 10:00-14:00, operator: on-call
    one file, pxar, GUI browser:        target n/a      measured __:__
    one file, cp from .snap:            target RPO 0:30 measured snapshot age _:__
    one guest, local, qmrestore:        target RTO 4:00 measured __:__
    one guest, other site, replica:     target RTO 4:00 measured __:__
    found: prune removed chosen snapshot; verify script hid failure
    

    The targets come from the written RPO and RTO; the measured column is filled in during the drill, never ahead of it.

  7. Unprotect the snapshots and schedule the next drill: quarterly, rotating the tenant namespace, with the same window reserved and checked against proxmox-backup-manager prune-job list, verify-job list and datastore list.

Verify it worked

The drill verified itself if each restore produced a thing you could open: the file matched its original, the guest booted with the right data, the guest on the other cluster did the same from the replica. Beyond that, the measured times sit under the written targets, the drill log has a dated entry, and the task list on both PBS hosts shows no prune or GC inside the window. A restore that failed is a successful drill with a finding, and the finding becomes an issue with an owner.

Gotchas

FAQ

How often? Quarterly for the full three, monthly for the one-file restore because it is ten minutes and it exercises the file archive that vzdump does not cover. After any change to the PBS hosts, the sync, or the storage entries, run the other-site restore once.

Why a new VMID and not restore over the original? Restoring over the live guest is the production procedure for a real incident and the drill should never be the thing that causes one. A new VMID in a reserved range, with its NIC down, cannot collide with anything.

What about the evidence case, where a guest must be preserved rather than restored? Same primitives. A CephFS snapshot of the exported artefacts plus a PBS backup marked protected, which prune, remove-vanished and garbage collection all leave alone. Read one back during the drill so the preservation procedure is rehearsed too.

Related