PBS: sync job says TASK OK but the datastore is empty
In short. TASK OK means the sync finished without an error, not that it copied anything. Read the task log for the "found N groups to sync" line first. Then check the namespace pair in the job, look for a previous run that never ended, and list the source with the job's own token to see what it can actually reach. Most of the time it is the namespace.
"TASK OK" and nothing arrived
The task log on the pulling PBS ends like this:
sync datastore 'replica' from 'site-a/backups'
found 0 groups to sync (out of 0 total)
TASK OK
The GUI shows a green tick, the schedule says it ran last night, and the target datastore has no groups in it. This was the state of the estate's first replica for two days before anyone opened the log. The sync had nothing to do, did nothing, and said so politely.
What you need
- Two PBS 3.4 hosts, a remote defined on the pulling side, a sync job that runs green.
- Root on the pulling PBS and the sync token's secret (or the ability to read the remote definition).
- The repository names used below:
backupson the source,replicaon the target, tenant namespaces on the source, asite-anamespace on the target.
Steps
-
Read the log, not the status. Find the task and print it:
proxmox-backup-manager task list --limit 30 proxmox-backup-manager task log 'UPID:pbs-b:...:syncjob:...'Three lines matter.
found 0 groups to sync (out of 0 total)means the job saw an empty source.found 0 groups to sync (out of 12 total)means it saw groups and filtered all of them out. A line per group saying the lock could not be acquired means another task has them. -
Compare the namespace pair. Show the job and read
remote-nsandnstogether:proxmox-backup-manager sync-job show pull-site-aremote-nsis where the job looks on the source,nsis where it lands on the target, andmax-depthis how far belowremote-nsit recurses. On this estate every backup group lives in a tenant namespace, so a job withremote-nspointed at a namespace that exists but holds only child namespaces, andmax-depthset to 0, sees nothing and is content. Aremote-nsthat is spelled differently from the real one (tenant_xfortenant-x) either errors or finds nothing, depending on the release. -
List the source with the job's own token. This is the dry run. It uses the same credential, the same fingerprint and the same path the job uses, and transfers nothing:
export PBS_REPOSITORY='sync@pbs!siteb@pbs-a.mesh.internal.example:backups' export PBS_FINGERPRINT='<fingerprint from the remote definition>' read -rs PBS_PASSWORD && export PBS_PASSWORD proxmox-backup-client namespace list proxmox-backup-client list --ns tenant-xIf
namespace listis empty but the source GUI shows namespaces, the token's ACL is the problem: a token grantedDatastoreReaderon/datastore/backups/tenant-xsees only that namespace, and one granted on the wrong datastore sees nothing. Iflist --ns tenant-xis empty while the GUI shows groups, the snapshots were pruned on the source before the sync ran, or the job'sgroup-filterexcludes them. -
Check for a run that never ended. A sync that lost its VPN mid-transfer can sit as a running task for hours, holding the group locks. Later runs skip what is locked.
cat /var/log/proxmox-backup/tasks/active ls /run/proxmox-backup/locks/ 2>/dev/nullStop the stuck task with
proxmox-backup-manager task stop <UPID>. On releases before the lock directory moved under/run, the lock was a.lockfile inside the group directory in the datastore; a reboot clears either. -
Fix the one cause you found and run the job by hand:
proxmox-backup-manager sync-job update pull-site-a --max-depth 2 proxmox-backup-manager sync-job run pull-site-aWatch the task this time. The first real run of a replica is slow; the log prints a line per snapshot as it lands.
Verify it worked
List groups on the target, in the namespace the job writes to, and count them against the source:
proxmox-backup-client list --repository replica --ns site-a/tenant-x
proxmox-backup-client snapshot list --repository replica --ns site-a/tenant-x
The group count should match what step 3 showed on the source. The task log should say found N groups to sync (out of N total) with N greater than zero, followed by sync snapshot lines. If the counts match and the next scheduled run also lands something, the job is fixed.
Gotchas
- The GUI's green tick is the exit status of the task, not a count of bytes. Put the group count in your monitoring, not the task state.
- A
remote-nsleft blank means the source root with full recursion, which is usually what you want for a replica. Setting it to a namespace and forgettingmax-depthis the common mistake. transfer-lastlimits how many snapshots per group come across. Set to 1 for a first run, it looks like a full sync and is not.- The sync token is read-only by design. If step 3 fails with a permission error, do not widen the token; fix the ACL path.
- If the source prunes at 06:00 and the target pulls at 06:30, a snapshot can vanish between listing and transfer. The log says so; move the pull earlier.
FAQ
Does TASK OK ever mean the sync copied nothing when it should have? Only in the sense above. The job is a listing followed by a copy of the difference; if the listing is empty, the copy is empty and the job is correct about itself. The error is in what the listing was pointed at.
Can I make an empty sync fail loudly? Not from the job itself. A small timer on the target that counts groups in the replica namespace and alerts if the count drops, or if the newest snapshot is older than two days, is what this estate runs. It caught the two-day gap above the second time, not the first.
The job pulls some tenants and not others. Same problem?
Same check. Either the missing tenants are below max-depth, or the token's ACL covers only some namespaces, or those tenants' groups are filtered by group-filter. Step 3 with --ns set to a missing tenant tells you which.