Proxmox Backup Server that leaves the building

Isometric illustration of two distant small buildings with servers inside, a dotted arc showing one pulling a copy from the other.

In short. One standalone PBS per site on the storage VLAN. One namespace per tenant, retention by resource pool. Then the remote site pulls the datastore over the mesh VPN with a read-only token, pins the source fingerprint, lands the copy in a per-source namespace, and verifies on its own schedule. The replica has its own prune rules and never deletes what the source deletes.

The uncomfortable discovery

An audit question: "which backups exist at a site other than the one the guest runs at?" The answer was none. Every cluster had a PBS, every PBS had years of retention, and a fire in one rack would have taken the backups with the guests.

This post is what fixed that, and the placement decisions that made the fix straightforward.

Placement

Namespaces and retention

One namespace per tenant. Retention is set by the resource pool the guest sits in, not per guest:

PoolKeep
Platform internals30 days
Tenants14 days, by agreement

The live job keeps 14 daily, 8 weekly and 3 monthly. Jobs are staggered between one and four in the morning so that no two clusters hit the same PBS at once.

Intra-day RBD snapshots run three times a day and keep three. They give fast rollback for a bad change; they are not backups and are never counted as such. The cron job that runs them once "succeeded" for months while snapshotting nothing because cron's PATH did not contain the guest tools. Errors went to /dev/null. The job now checks that a snapshot exists afterwards and fails loudly if it does not.

Replication: pull, not push

The remote PBS pulls. The source never has credentials for the target.

flowchart LR
  subgraph A["Site A"]
    PA["PVE cluster A"] -- "vzdump nightly, VLAN 5" --> PBSA["PBS A<br/>namespace per tenant<br/>retention by pool"]
  end
  subgraph B["Site B"]
    PB["PVE cluster B"] -- "vzdump nightly, VLAN 5" --> PBSB["PBS B"]
    PBSB --> R["Replica namespace per source<br/>remove-vanished off, own prune"]
    R --> V["verify daily<br/>GC weekly"]
    PB -. "read-only storage entry<br/>restore from replica" .-> R
  end
  PBSB -- "pull sync over the mesh VPN<br/>read-only token, pinned fingerprint" --> PBSA

Restoring from the other side

A replica you cannot restore from is a tape in a drawer. The surviving cluster has a read-only storage entry pointing at the replica namespace on its local PBS. Restore is the normal restore dialog; the guest comes up with a new ID in the local range.

The restore drill found two things the sync job did not:

What vzdump does not capture

Targets, written down

TierRPORTO
Platform services30 minutes (snapshots)4 hours
Tenant guests24 hours4 hours

Having the numbers written down is what turns "we should replicate" into a job with a deadline.

FAQ

Why not push from the source? A push needs the source to hold a credential for the target. If the source is compromised, the attacker holds a credential for your off-site backups.

Why one PBS per site instead of one central PBS? Because the first copy must be fast and local. The central copy is the sync, and it can be slow.

Could the replica be a different product, like object storage? It could. PBS to PBS keeps one restore procedure and one verify mechanism, and both sites already had the hardware.

Related