PBS: pull replication between two sites with a read-only token

In short. Create a read-only token on the source, read its certificate fingerprint, define the source as a remote on the target with that fingerprint, and create a sync job that pulls into a namespace named for the source with remove-vanished off. Run it once, count the groups, then give the replica its own prune, verify and garbage collection. Finish by adding a read-only PBS storage entry on the surviving cluster so a restore from the replica is a normal restore.

Getting a copy of every backup to the other site

The question that started this was "which backups exist somewhere other than where the guest runs?" and the answer was none. The fix below is a pull: the target reaches out to the source with a credential that can read and nothing else. The source never holds a credential for the target.

flowchart LR
  subgraph A["Site A"]
    PVEA["PVE cluster A<br/>10.1.5.51+"] -- "vzdump 01:00 to 04:00" --> PBSA["PBS A, datastore backups<br/>10.1.5.49, ns per tenant"]
    TOK["token sync@pbs!siteb<br/>DatastoreReader on backups"] -.-> PBSA
  end
  subgraph B["Site B"]
    PBSB["PBS B, datastore replica"] --> NS["ns site-a/tenant-x<br/>remove-vanished off<br/>own prune, verify, GC"]
    PVEB["PVE cluster B"] -. "read-only storage entry<br/>restore only" .-> NS
  end
  PBSB -- "pull over the mesh VPN<br/>pinned fingerprint" --> PBSA

What you need

Steps

  1. On the source, create a user and a token, and grant the token read access to exactly one datastore. The secret prints once; put it in your vault, not a file.

    proxmox-backup-manager user create sync@pbs --comment "pulled by site B"
    proxmox-backup-manager user generate-token sync@pbs siteb
    proxmox-backup-manager acl update /datastore/backups DatastoreReader --auth-id 'sync@pbs!siteb'
    

    DatastoreReader can list and read; it cannot write, prune or delete. The ACL path is the whole datastore so every tenant namespace is visible to the pull.

  2. On the source, read the certificate fingerprint:

    proxmox-backup-manager cert info
    

    Copy the Fingerprint (sha256) line. The target will refuse to talk to anything presenting a different certificate.

  3. On the target, define the remote. Read the secret into a variable so it does not land in shell history:

    read -rs SECRET
    proxmox-backup-manager remote create site-a \
      --host pbs-a.mesh.internal.example \
      --auth-id 'sync@pbs!siteb' \
      --password "$SECRET" \
      --fingerprint 'AA:BB:...'
    proxmox-backup-manager remote list
    
  4. On the target, create the per-source namespace and the sync job. Leave remote-ns unset so the job starts at the source root and recurses into every tenant namespace; each lands as site-a/<tenant>.

    proxmox-backup-client namespace create site-a --repository replica
    proxmox-backup-manager sync-job create pull-site-a \
      --store replica --ns site-a \
      --remote site-a --remote-store backups \
      --remove-vanished false \
      --schedule '05:00' \
      --comment "pull site A, read-only token"
    

    The schedule is after the last backup job on the source finishes at 04:00 and before the source prunes. remove-vanished stays off: if the source loses a group, by accident or by someone with a stolen credential, the replica keeps its copy.

  5. Run it once by hand and watch the task:

    proxmox-backup-manager sync-job run pull-site-a
    

    The first run copies everything and takes hours over a site link. Later runs copy only new snapshots.

  6. Give the replica its own retention and checking. Prune per namespace with its own keep rules, verify daily with re-verification after 30 days, and collect garbage weekly:

    proxmox-backup-manager prune-job create prune-site-a \
      --store replica --ns site-a --max-depth 2 \
      --schedule '07:00' --keep-daily 14 --keep-weekly 8 --keep-monthly 3
    proxmox-backup-manager verify-job create verify-replica \
      --store replica --schedule '09:00' \
      --ignore-verified true --outdated-after 30
    proxmox-backup-manager datastore update replica --gc-schedule 'sun 12:00'
    
  7. Let the surviving cluster see the replica. On the target PBS, create a second read-only token scoped to the replica namespace; on a node of the site B cluster, add a PBS storage entry pointed at the tenant namespace you expect to restore from:

    proxmox-backup-manager user create restore@pbs
    proxmox-backup-manager user generate-token restore@pbs site-a
    proxmox-backup-manager acl update /datastore/replica/site-a DatastoreReader --auth-id 'restore@pbs!site-a'
    
    pvesm add pbs pbs-replica-a \
      --server 10.2.5.49 --datastore replica --namespace site-a/tenant-x \
      --username 'restore@pbs!site-a' --password "$SECRET" \
      --fingerprint 'CC:DD:...' --content backup
    

    The fingerprint here is PBS B's, from cert info on that host. A backup job pointed at this storage fails, which is the point: the entry exists for restores.

Verify it worked

On the target, count groups and look at the newest snapshot per tenant:

proxmox-backup-client list --repository replica --ns site-a/tenant-x
proxmox-backup-client snapshot list --repository replica --ns site-a/tenant-x

Compare with the same commands on the source without the site-a/ prefix. Then on the site B cluster, pvesm list pbs-replica-a should print the site A guests' backups as volumes. The next morning, check the task log for the scheduled run and confirm it transferred at least one snapshot.

Gotchas

FAQ

Why pull and not push? A push needs the source to hold a credential that can write to the target. A compromised source could then write, and with the right role delete, at the other site. A pull with a read-only token means a compromised source exposes one datastore to being read and nothing else.

Does the replica need its own prune rules? Yes, and they should be the replica's, not a copy of the source's. The replica keeps 14 daily, 8 weekly and 3 monthly here, the same as the live job, but with remove-vanished off it also keeps groups the source has deleted until its own prune ages them out.

What about bandwidth over the site link? The sync job accepts a rate limit (--rate-in on the pulling side). The first full copy ran over a weekend; daily increments for a few tenants are small.

Related