PBS: a lot of "verify failed", and the protect-latest trick

In short. Read the verify log for which chunk failed and how many snapshots share it; deduplication means one bad chunk fails every snapshot that references it. Before re-running anything, protect the newest snapshot per group so prune cannot remove your last good copy while you are busy. Then re-run the backup to re-upload the chunk, or pull it from the replica at the other site, and move verify and garbage collection apart so they stop producing load failures that look like corruption.

"verification failed" on forty snapshots overnight

The morning task list shows the verify job in red and the log ends like this:

verify backups:tenant-x/vm/104/2026-09-03T01:40:12Z
  can't verify chunk, load failed - ...
  verified 2140 chunks, 1 failed
Failed to verify the following snapshots/groups:
  tenant-x/vm/104/2026-09-03T01:40:12Z
  tenant-x/vm/104/2026-09-02T01:40:09Z
  ...
TASK ERROR: verification failed - please check the log for details

Forty snapshots in a dozen groups failed, all between 02:00 and 03:30, which is when garbage collection was also running. Some of these are real and some are not, and the log alone does not say which.

What you need

Steps

  1. Read the log for the chunk, not the snapshot list. Find the first can't verify chunk line and copy the digest. Then count how many failed snapshots are in the same group and whether they share it. One digest appearing under many snapshots is one corrupt or unreadable chunk; many different digests across unrelated groups during the same hour is contention or a failing disk.

  2. Check what else was running. proxmox-backup-manager task list --limit 100 for the night; if garbage collection, a large sync, or a second verify overlapped, the load failures may be I/O timeouts. Check the disk underneath:

    dmesg -T | grep -iE 'i/o error|ata|nvme|reset' | tail
    smartctl -a /dev/sda | grep -iE 'reallocated|pending|uncorrect'
    zpool status
    

    A reallocated-sector count that moved, or a ZFS checksum error, means step 6 is the real fix and everything else is damage control.

  3. Protect the newest snapshot in every group before touching anything. A verify failure marks the snapshot, and the next prune still treats it as a normal snapshot. If the only snapshot in a group that still verifies clean is the oldest one, you do not want it pruned tomorrow. The GUI toggle is under the datastore's Content tab, per snapshot. For a whole datastore, loop:

    set -o pipefail
    export PBS_REPOSITORY=backups
    for ns in $(proxmox-backup-client namespace list --output-format json | jq -r '.[].ns'); do
      for g in $(proxmox-backup-client list --ns "$ns" --output-format json \
                 | jq -r '.[] | "\(.["backup-type"])/\(.["backup-id"])"'); do
        latest=$(proxmox-backup-client snapshot list "$g" --ns "$ns" --output-format json \
                 | jq -r 'sort_by(.["backup-time"]) | last | "\(.["backup-type"])/\(.["backup-id"])/\(.["backup-time"] | todate)"')
        proxmox-backup-client snapshot protected update "$latest" true --ns "$ns"
      done
    done
    

    Adapt the namespace depth to your layout. Run it with echo in front of the last command first. The set -o pipefail is there because the estate's earlier verification script piped a failing command into a succeeding one and reported green for weeks.

  4. Re-run verify on one failed group with nothing else running. Create a one-off verify job scoped to the namespace, or use the Verify button on the group in the GUI, and run it in daylight:

    proxmox-backup-manager verify-job create verify-once-tenant-x \
      --store backups --ns tenant-x --ignore-verified false
    proxmox-backup-manager verify-job run verify-once-tenant-x
    

    Snapshots that pass now were load failures under contention. Snapshots that fail again have a bad chunk.

  5. Re-upload the chunk from the source. PBS renames a chunk that fails verification with a .bad suffix, so it is no longer present, and it will not use a snapshot that failed verification as the base for the next incremental. Run the guest's backup job again from PVE:

    vzdump 104 --storage pbs-a --mode snapshot
    

    The client uploads the missing chunk, the new snapshot references it, and the older failed snapshots stay failed because their index points at a chunk that no longer exists under that name. They age out through prune.

  6. If the guest is gone, or the data is a file archive whose source snapshot has rotated, pull from the other site. The replica has its own copy of the chunk if it was synced before the corruption. Define the replica as a remote on this PBS and run a one-off sync job scoped to the one group, or restore the guest on the other cluster through its read-only storage entry. Then replace the disk.

  7. Move the windows apart. Verify at 09:00, GC on Saturday noon, and never both in the backup window. The related post has the calendar.

Verify it worked

Run the normal verify job by hand after the re-upload and read its log:

proxmox-backup-manager verify-job run verify-backups

The new snapshot of each affected group verifies clean. Older failed snapshots remain marked failed and that is expected. proxmox-backup-client snapshot list --ns tenant-x vm/104 shows the newest snapshot with a passing verification state. After the next weekly GC, find /mnt/datastore/backups/.chunks -name '*.bad' | wc -l should stop growing.

Gotchas

FAQ

Why protect before re-verifying rather than after? Because the re-verify takes hours and the prune job is scheduled. Protecting first takes a minute and removes the race. Unprotect afterwards; it is a flag, not a copy.

Should I delete the failed snapshots? Not yet. They may still restore most of their files, and in the case of a pxar archive the file browser works around a single bad chunk. Let prune age them out once a clean snapshot exists.

Can verify itself corrupt anything? No. It reads and compares. The only write it does is renaming a chunk that fails its digest to .bad, which is why a missing chunk and a corrupt chunk look the same to the next backup client.

Related