Proxmox Backup Server that leaves the building

In short. One standalone PBS per site on the storage VLAN. One namespace per tenant, retention by resource pool. Then the remote site pulls the datastore over the mesh VPN with a read-only token, pins the source fingerprint, lands the copy in a per-source namespace, and verifies on its own schedule. The replica has its own prune rules and never deletes what the source deletes.
The uncomfortable discovery
An audit question: "which backups exist at a site other than the one the guest runs at?" The answer was none. Every cluster had a PBS, every PBS had years of retention, and a fire in one rack would have taken the backups with the guests.
This post is what fixed that, and the placement decisions that made the fix straightforward.
Placement
- Standalone. The backup server is never a cluster member. It does not run guests and it does not share the cluster's failure modes.
- On the storage VLAN. Backup traffic runs over the same 10G jumbo-frame VLAN as Ceph, never over management. Management is for out-of-band controllers and SSH; it is not sized for terabytes.
- Access port, untagged. The PBS is the one host class that is untagged on the storage VLAN. Get that wrong and LACP still comes up while nothing passes.
- Old hardware is fine. The first site's PBS is a chassis two generations older than the compute nodes, with a RAID controller that cannot be flashed to IT mode. For PBS that does not matter.
Namespaces and retention
One namespace per tenant. Retention is set by the resource pool the guest sits in, not per guest:
| Pool | Keep |
|---|---|
| Platform internals | 30 days |
| Tenants | 14 days, by agreement |
The live job keeps 14 daily, 8 weekly and 3 monthly. Jobs are staggered between one and four in the morning so that no two clusters hit the same PBS at once.
Intra-day RBD snapshots run three times a day and keep three. They give fast rollback for a bad change; they are not backups and are never counted as such. The cron job that runs them once "succeeded" for months while snapshotting nothing because cron's PATH did not contain the guest tools. Errors went to /dev/null. The job now checks that a snapshot exists afterwards and fails loudly if it does not.
Replication: pull, not push
The remote PBS pulls. The source never has credentials for the target.
- Token scoped to one datastore, read-only. If the token leaks, the worst case is a read of one site's backups.
- Fingerprint pinned. The sync job carries the source PBS's certificate fingerprint. Renaming the source host regenerates its certificate and breaks every fingerprint that referenced it, which is a feature: a changed identity should fail closed.
- Per-source namespace. Guest IDs collide across clusters. Guest 100 at site A and guest 100 at site B are different machines, so the replica lands in a namespace named for the source.
remove-vanishedoff. If the source loses a backup group, the replica keeps its copy. A compromised source cannot reach through the sync and erase history.- Own prune rules, own verify, own garbage collection. The replica verifies daily, re-verifying anything older than 30 days, and collects garbage weekly.
- Transport is the mesh VPN. The two sites are peers on the overlay; the sync job talks to the source by its overlay name.
flowchart LR
subgraph A["Site A"]
PA["PVE cluster A"] -- "vzdump nightly, VLAN 5" --> PBSA["PBS A<br/>namespace per tenant<br/>retention by pool"]
end
subgraph B["Site B"]
PB["PVE cluster B"] -- "vzdump nightly, VLAN 5" --> PBSB["PBS B"]
PBSB --> R["Replica namespace per source<br/>remove-vanished off, own prune"]
R --> V["verify daily<br/>GC weekly"]
PB -. "read-only storage entry<br/>restore from replica" .-> R
end
PBSB -- "pull sync over the mesh VPN<br/>read-only token, pinned fingerprint" --> PBSARestoring from the other side
A replica you cannot restore from is a tape in a drawer. The surviving cluster has a read-only storage entry pointing at the replica namespace on its local PBS. Restore is the normal restore dialog; the guest comes up with a new ID in the local range.
The restore drill found two things the sync job did not:
- A prune race. A prune ran while a restore was reading the group it was pruning. Prune windows now avoid the drill window.
- A masked pipeline failure. A verification script piped a failing command into one that succeeded and reported green.
set -o pipefailfixed the script; the lesson is that any verification that does not itself fail loudly is decoration.
What vzdump does not capture
- No snapshots: only the current disk state.
- No RAM.
- No bind mounts in containers. If your user data is on a CephFS bind mount, the container backup is a few hundred megabytes of runtime and none of the data. That has its own post.
Targets, written down
| Tier | RPO | RTO |
|---|---|---|
| Platform services | 30 minutes (snapshots) | 4 hours |
| Tenant guests | 24 hours | 4 hours |
Having the numbers written down is what turns "we should replicate" into a job with a deadline.
FAQ
Why not push from the source? A push needs the source to hold a credential for the target. If the source is compromised, the attacker holds a credential for your off-site backups.
Why one PBS per site instead of one central PBS? Because the first copy must be fast and local. The central copy is the sync, and it can be slow.
Could the replica be a different product, like object storage? It could. PBS to PBS keeps one restore procedure and one verify mechanism, and both sites already had the hardware.