PBS: sync job says TASK OK but the datastore is empty

In short. TASK OK means the sync finished without an error, not that it copied anything. Read the task log for the "found N groups to sync" line first. Then check the namespace pair in the job, look for a previous run that never ended, and list the source with the job's own token to see what it can actually reach. Most of the time it is the namespace.

"TASK OK" and nothing arrived

The task log on the pulling PBS ends like this:

sync datastore 'replica' from 'site-a/backups'
found 0 groups to sync (out of 0 total)
TASK OK

The GUI shows a green tick, the schedule says it ran last night, and the target datastore has no groups in it. This was the state of the estate's first replica for two days before anyone opened the log. The sync had nothing to do, did nothing, and said so politely.

What you need

Steps

  1. Read the log, not the status. Find the task and print it:

    proxmox-backup-manager task list --limit 30
    proxmox-backup-manager task log 'UPID:pbs-b:...:syncjob:...'
    

    Three lines matter. found 0 groups to sync (out of 0 total) means the job saw an empty source. found 0 groups to sync (out of 12 total) means it saw groups and filtered all of them out. A line per group saying the lock could not be acquired means another task has them.

  2. Compare the namespace pair. Show the job and read remote-ns and ns together:

    proxmox-backup-manager sync-job show pull-site-a
    

    remote-ns is where the job looks on the source, ns is where it lands on the target, and max-depth is how far below remote-ns it recurses. On this estate every backup group lives in a tenant namespace, so a job with remote-ns pointed at a namespace that exists but holds only child namespaces, and max-depth set to 0, sees nothing and is content. A remote-ns that is spelled differently from the real one (tenant_x for tenant-x) either errors or finds nothing, depending on the release.

  3. List the source with the job's own token. This is the dry run. It uses the same credential, the same fingerprint and the same path the job uses, and transfers nothing:

    export PBS_REPOSITORY='sync@pbs!siteb@pbs-a.mesh.internal.example:backups'
    export PBS_FINGERPRINT='<fingerprint from the remote definition>'
    read -rs PBS_PASSWORD && export PBS_PASSWORD
    proxmox-backup-client namespace list
    proxmox-backup-client list --ns tenant-x
    

    If namespace list is empty but the source GUI shows namespaces, the token's ACL is the problem: a token granted DatastoreReader on /datastore/backups/tenant-x sees only that namespace, and one granted on the wrong datastore sees nothing. If list --ns tenant-x is empty while the GUI shows groups, the snapshots were pruned on the source before the sync ran, or the job's group-filter excludes them.

  4. Check for a run that never ended. A sync that lost its VPN mid-transfer can sit as a running task for hours, holding the group locks. Later runs skip what is locked.

    cat /var/log/proxmox-backup/tasks/active
    ls /run/proxmox-backup/locks/ 2>/dev/null
    

    Stop the stuck task with proxmox-backup-manager task stop <UPID>. On releases before the lock directory moved under /run, the lock was a .lock file inside the group directory in the datastore; a reboot clears either.

  5. Fix the one cause you found and run the job by hand:

    proxmox-backup-manager sync-job update pull-site-a --max-depth 2
    proxmox-backup-manager sync-job run pull-site-a
    

    Watch the task this time. The first real run of a replica is slow; the log prints a line per snapshot as it lands.

Verify it worked

List groups on the target, in the namespace the job writes to, and count them against the source:

proxmox-backup-client list --repository replica --ns site-a/tenant-x
proxmox-backup-client snapshot list --repository replica --ns site-a/tenant-x

The group count should match what step 3 showed on the source. The task log should say found N groups to sync (out of N total) with N greater than zero, followed by sync snapshot lines. If the counts match and the next scheduled run also lands something, the job is fixed.

Gotchas

FAQ

Does TASK OK ever mean the sync copied nothing when it should have? Only in the sense above. The job is a listing followed by a copy of the difference; if the listing is empty, the copy is empty and the job is correct about itself. The error is in what the listing was pointed at.

Can I make an empty sync fail loudly? Not from the job itself. A small timer on the target that counts groups in the replica namespace and alerts if the count drops, or if the newest snapshot is older than two days, is what this estate runs. It caught the two-day gap above the second time, not the first.

The job pulls some tenants and not others. Same problem? Same check. Either the missing tenants are below max-depth, or the token's ACL covers only some namespaces, or those tenants' groups are filtered by group-filter. Step 3 with --ns set to a missing tenant tells you which.

Related