PBS: a lot of "verify failed", and the protect-latest trick
In short. Read the verify log for which chunk failed and how many snapshots share it; deduplication means one bad chunk fails every snapshot that references it. Before re-running anything, protect the newest snapshot per group so prune cannot remove your last good copy while you are busy. Then re-run the backup to re-upload the chunk, or pull it from the replica at the other site, and move verify and garbage collection apart so they stop producing load failures that look like corruption.
"verification failed" on forty snapshots overnight
The morning task list shows the verify job in red and the log ends like this:
verify backups:tenant-x/vm/104/2026-09-03T01:40:12Z
can't verify chunk, load failed - ...
verified 2140 chunks, 1 failed
Failed to verify the following snapshots/groups:
tenant-x/vm/104/2026-09-03T01:40:12Z
tenant-x/vm/104/2026-09-02T01:40:09Z
...
TASK ERROR: verification failed - please check the log for details
Forty snapshots in a dozen groups failed, all between 02:00 and 03:30, which is when garbage collection was also running. Some of these are real and some are not, and the log alone does not say which.
What you need
- PBS 4.x with a verify job, root on the host, and
jqinstalled. - The task list for the night in question, including whatever else ran.
- If the datastore is replicated, access to the other site's PBS as well.
- Names below: datastore
backups, namespacetenant-x.
Steps
-
Read the log for the chunk, not the snapshot list. Find the first
can't verify chunkline and copy the digest. Then count how many failed snapshots are in the same group and whether they share it. One digest appearing under many snapshots is one corrupt or unreadable chunk; many different digests across unrelated groups during the same hour is contention or a failing disk. -
Check what else was running.
proxmox-backup-manager task list --limit 100for the night; if garbage collection, a large sync, or a second verify overlapped, the load failures may be I/O timeouts. Check the disk underneath:dmesg -T | grep -iE 'i/o error|ata|nvme|reset' | tail smartctl -a /dev/sda | grep -iE 'reallocated|pending|uncorrect' zpool statusA reallocated-sector count that moved, or a ZFS checksum error, means step 6 is the real fix and everything else is damage control.
-
Protect the newest snapshot in every group before touching anything. A verify failure marks the snapshot, and the next prune still treats it as a normal snapshot. If the only snapshot in a group that still verifies clean is the oldest one, you do not want it pruned tomorrow. The GUI toggle is under the datastore's Content tab, per snapshot. For a whole datastore, loop:
set -o pipefail export PBS_REPOSITORY=backups for ns in $(proxmox-backup-client namespace list --output-format json | jq -r '.[].ns'); do for g in $(proxmox-backup-client list --ns "$ns" --output-format json \ | jq -r '.[] | "\(.["backup-type"])/\(.["backup-id"])"'); do latest=$(proxmox-backup-client snapshot list "$g" --ns "$ns" --output-format json \ | jq -r 'sort_by(.["backup-time"]) | last | "\(.["backup-type"])/\(.["backup-id"])/\(.["backup-time"] | todate)"') proxmox-backup-client snapshot protected update "$latest" true --ns "$ns" done doneAdapt the namespace depth to your layout. Run it with
echoin front of the last command first. Theset -o pipefailis there because the estate's earlier verification script piped a failing command into a succeeding one and reported green for weeks. -
Re-run verify on one failed group with nothing else running. Create a one-off verify job scoped to the namespace, or use the Verify button on the group in the GUI, and run it in daylight:
proxmox-backup-manager verify-job create verify-once-tenant-x \ --store backups --ns tenant-x --ignore-verified false proxmox-backup-manager verify-job run verify-once-tenant-xSnapshots that pass now were load failures under contention. Snapshots that fail again have a bad chunk.
-
Re-upload the chunk from the source. PBS renames a chunk that fails verification with a
.badsuffix, so it is no longer present, and it will not use a snapshot that failed verification as the base for the next incremental. Run the guest's backup job again from PVE:vzdump 104 --storage pbs-a --mode snapshotThe client uploads the missing chunk, the new snapshot references it, and the older failed snapshots stay failed because their index points at a chunk that no longer exists under that name. They age out through prune.
-
If the guest is gone, or the data is a file archive whose source snapshot has rotated, pull from the other site. The replica has its own copy of the chunk if it was synced before the corruption. Define the replica as a remote on this PBS and run a one-off sync job scoped to the one group, or restore the guest on the other cluster through its read-only storage entry. Then replace the disk.
-
Move the windows apart. Verify at 09:00, GC on Saturday noon, and never both in the backup window. The related post has the calendar.
Verify it worked
Run the normal verify job by hand after the re-upload and read its log:
proxmox-backup-manager verify-job run verify-backups
The new snapshot of each affected group verifies clean. Older failed snapshots remain marked failed and that is expected. proxmox-backup-client snapshot list --ns tenant-x vm/104 shows the newest snapshot with a passing verification state. After the next weekly GC, find /mnt/datastore/backups/.chunks -name '*.bad' | wc -l should stop growing.
Gotchas
- A protected snapshot survives prune,
remove-vanishedand GC. Remember to unprotect once the group has a clean newer snapshot, or the datastore slowly fills with pinned history. - Verify marks the snapshot, not the chunk. A chunk shared by twenty snapshots fails all twenty, and fixing it once fixes none of them retroactively; only a new snapshot gets a good reference.
- GC and verify on the same night produce
load failedlines that are not corruption. Do not replace a disk on the strength of one overlapping night. ignore-verifiedmeans a snapshot marked failed is not re-checked untiloutdated-afterexpires. Use a one-off job without it when you need an answer today.- The replica only helps if its copy predates the corruption. A sync after the chunk went bad copies the bad chunk's absence, not its content.
FAQ
Why protect before re-verifying rather than after? Because the re-verify takes hours and the prune job is scheduled. Protecting first takes a minute and removes the race. Unprotect afterwards; it is a flag, not a copy.
Should I delete the failed snapshots? Not yet. They may still restore most of their files, and in the case of a pxar archive the file browser works around a single bad chunk. Let prune age them out once a clean snapshot exists.
Can verify itself corrupt anything?
No. It reads and compares. The only write it does is renaming a chunk that fails its digest to .bad, which is why a missing chunk and a corrupt chunk look the same to the next backup client.