PBS: pull replication between two sites with a read-only token
In short. Create a read-only token on the source, read its certificate fingerprint, define the source as a remote on the target with that fingerprint, and create a sync job that pulls into a namespace named for the source with
remove-vanishedoff. Run it once, count the groups, then give the replica its own prune, verify and garbage collection. Finish by adding a read-only PBS storage entry on the surviving cluster so a restore from the replica is a normal restore.
Getting a copy of every backup to the other site
The question that started this was "which backups exist somewhere other than where the guest runs?" and the answer was none. The fix below is a pull: the target reaches out to the source with a credential that can read and nothing else. The source never holds a credential for the target.
flowchart LR
subgraph A["Site A"]
PVEA["PVE cluster A<br/>10.1.5.51+"] -- "vzdump 01:00 to 04:00" --> PBSA["PBS A, datastore backups<br/>10.1.5.49, ns per tenant"]
TOK["token sync@pbs!siteb<br/>DatastoreReader on backups"] -.-> PBSA
end
subgraph B["Site B"]
PBSB["PBS B, datastore replica"] --> NS["ns site-a/tenant-x<br/>remove-vanished off<br/>own prune, verify, GC"]
PVEB["PVE cluster B"] -. "read-only storage entry<br/>restore only" .-> NS
end
PBSB -- "pull over the mesh VPN<br/>pinned fingerprint" --> PBSAWhat you need
- PBS 3.4 at both sites, each a standalone host on its site's storage VLAN, neither a PVE cluster member.
- A route between the two PBS hosts. Here it is a self-hosted WireGuard mesh and each PBS is a peer; the target addresses the source by its overlay name,
pbs-a.mesh.internal.example. - Root on both PBS hosts and on one node of the surviving PVE 8.4 cluster.
- Names used below: source datastore
backups, target datastorereplica, remotesite-a, target namespacesite-a.
Steps
-
On the source, create a user and a token, and grant the token read access to exactly one datastore. The secret prints once; put it in your vault, not a file.
proxmox-backup-manager user create sync@pbs --comment "pulled by site B" proxmox-backup-manager user generate-token sync@pbs siteb proxmox-backup-manager acl update /datastore/backups DatastoreReader --auth-id 'sync@pbs!siteb'DatastoreReadercan list and read; it cannot write, prune or delete. The ACL path is the whole datastore so every tenant namespace is visible to the pull. -
On the source, read the certificate fingerprint:
proxmox-backup-manager cert infoCopy the
Fingerprint (sha256)line. The target will refuse to talk to anything presenting a different certificate. -
On the target, define the remote. Read the secret into a variable so it does not land in shell history:
read -rs SECRET proxmox-backup-manager remote create site-a \ --host pbs-a.mesh.internal.example \ --auth-id 'sync@pbs!siteb' \ --password "$SECRET" \ --fingerprint 'AA:BB:...' proxmox-backup-manager remote list -
On the target, create the per-source namespace and the sync job. Leave
remote-nsunset so the job starts at the source root and recurses into every tenant namespace; each lands assite-a/<tenant>.proxmox-backup-client namespace create site-a --repository replica proxmox-backup-manager sync-job create pull-site-a \ --store replica --ns site-a \ --remote site-a --remote-store backups \ --remove-vanished false \ --schedule '05:00' \ --comment "pull site A, read-only token"The schedule is after the last backup job on the source finishes at 04:00 and before the source prunes.
remove-vanishedstays off: if the source loses a group, by accident or by someone with a stolen credential, the replica keeps its copy. -
Run it once by hand and watch the task:
proxmox-backup-manager sync-job run pull-site-aThe first run copies everything and takes hours over a site link. Later runs copy only new snapshots.
-
Give the replica its own retention and checking. Prune per namespace with its own keep rules, verify daily with re-verification after 30 days, and collect garbage weekly:
proxmox-backup-manager prune-job create prune-site-a \ --store replica --ns site-a --max-depth 2 \ --schedule '07:00' --keep-daily 14 --keep-weekly 8 --keep-monthly 3 proxmox-backup-manager verify-job create verify-replica \ --store replica --schedule '09:00' \ --ignore-verified true --outdated-after 30 proxmox-backup-manager datastore update replica --gc-schedule 'sun 12:00' -
Let the surviving cluster see the replica. On the target PBS, create a second read-only token scoped to the replica namespace; on a node of the site B cluster, add a PBS storage entry pointed at the tenant namespace you expect to restore from:
proxmox-backup-manager user create restore@pbs proxmox-backup-manager user generate-token restore@pbs site-a proxmox-backup-manager acl update /datastore/replica/site-a DatastoreReader --auth-id 'restore@pbs!site-a'pvesm add pbs pbs-replica-a \ --server 10.2.5.49 --datastore replica --namespace site-a/tenant-x \ --username 'restore@pbs!site-a' --password "$SECRET" \ --fingerprint 'CC:DD:...' --content backupThe fingerprint here is PBS B's, from
cert infoon that host. A backup job pointed at this storage fails, which is the point: the entry exists for restores.
Verify it worked
On the target, count groups and look at the newest snapshot per tenant:
proxmox-backup-client list --repository replica --ns site-a/tenant-x
proxmox-backup-client snapshot list --repository replica --ns site-a/tenant-x
Compare with the same commands on the source without the site-a/ prefix. Then on the site B cluster, pvesm list pbs-replica-a should print the site A guests' backups as volumes. The next morning, check the task log for the scheduled run and confirm it transferred at least one snapshot.
Gotchas
- Guest IDs collide across clusters. VM 100 at site A and VM 100 at site B are different machines; the per-source namespace is what keeps them apart. Never pull into the target's root.
- Renaming either PBS regenerates its certificate and breaks every pinned fingerprint. Treat a rename as a change with a runbook.
- A sync that reports TASK OK with an empty target is almost always a namespace pair problem; see the related post before touching the token.
- The sync token's secret is in the remote definition on the target, in
/etc/proxmox-backup/remote.cfg. That file is root-only; keep it that way and keep it out of backups of the PBS host itself. - The surviving cluster's storage entry is per namespace. Add entries for other tenants at restore time; a dozen idle entries slow every
pvesm status.
FAQ
Why pull and not push? A push needs the source to hold a credential that can write to the target. A compromised source could then write, and with the right role delete, at the other site. A pull with a read-only token means a compromised source exposes one datastore to being read and nothing else.
Does the replica need its own prune rules?
Yes, and they should be the replica's, not a copy of the source's. The replica keeps 14 daily, 8 weekly and 3 monthly here, the same as the live job, but with remove-vanished off it also keeps groups the source has deleted until its own prune ages them out.
What about bandwidth over the site link?
The sync job accepts a rate limit (--rate-in on the pulling side). The first full copy ran over a weekend; daily increments for a few tenants are small.