Proxmox: intra-day snapshots from cron that actually run

In short. qm and pct live in /usr/sbin, which is not on cron's default PATH. A script that calls them from cron with its output sent to /dev/null fails every run with exit 127 and nobody knows. Set PATH in the script, use set -euo pipefail, confirm the snapshot exists after taking it, rotate to keep three, and send the output somewhere a human or a monitor reads it.

"qm: command not found" from cron, or no error at all

You wrote a script that snapshots every guest on the node, put it in /etc/cron.d, and the log said nothing. Months later a restore was needed and qm listsnapshot was empty. Proxmox VE 8.4 with Ceph Reef RBD underneath, where snapshots are cheap and instant.

What you need

Steps

  1. Write the script with PATH set at the top. This one line is the fix for the original fault; the rest is what makes the next fault visible.
#!/bin/bash
# /usr/local/sbin/rbd-snap
set -euo pipefail
export PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin

KEEP=3
PREFIX=auto
NAME="${PREFIX}-$(date +%Y%m%d-%H%M)"
FAILED=0

log() { logger -t rbd-snap -- "$*"; printf '%s %s\n' "$(date -Is)" "$*"; }

snap_one() {
  local tool=$1 vmid=$2
  if "$tool" config "$vmid" | grep -q '^template: 1'; then
    return 0
  fi
  if ! "$tool" snapshot "$vmid" "$NAME" --description "intra-day, keep $KEEP"; then
    log "FAIL $tool $vmid: snapshot command failed"
    return 1
  fi
  if ! "$tool" listsnapshot "$vmid" | grep -q " $NAME "; then
    log "FAIL $tool $vmid: snapshot $NAME not present after create"
    return 1
  fi
  "$tool" listsnapshot "$vmid" \
    | awk -v p="$PREFIX-" '$1 == "`->" && index($2, p) == 1 { print $2 }' \
    | sort | head -n -"$KEEP" \
    | while read -r old; do "$tool" delsnapshot "$vmid" "$old"; done
  log "ok $tool $vmid $NAME"
}

for id in $(qm list | awk 'NR > 1 { print $1 }'); do
  snap_one qm "$id" || FAILED=1
done
for id in $(pct list | awk 'NR > 1 { print $1 }'); do
  snap_one pct "$id" || FAILED=1
done

if [ "$FAILED" -ne 0 ]; then
  log "FAIL: at least one guest was not snapshotted"
  exit 1
fi
log "all guests snapshotted as $NAME"
  1. Install it and run it by hand once, as root, from a shell with the minimal cron environment so you see what cron will see.
chmod 755 /usr/local/sbin/rbd-snap
env -i HOME=/root SHELL=/bin/sh PATH=/usr/bin:/bin /usr/local/sbin/rbd-snap
  1. Schedule it. Output goes to a log file and, because of MAILTO, any output also goes to root's mail. The original line sent everything to /dev/null; the whole lesson is in that difference.
# /etc/cron.d/rbd-snap
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
MAILTO=root
0 6,12,18 * * * root /usr/local/sbin/rbd-snap >> /var/log/rbd-snap.log 2>&1 || echo "rbd-snap failed on $(hostname), see /var/log/rbd-snap.log"
  1. Rotate the log so it does not grow forever.
# /etc/logrotate.d/rbd-snap
/var/log/rbd-snap.log {
    weekly
    rotate 8
    compress
    missingok
    notifempty
}
  1. Add an independent check. A script that reports its own success is only as honest as its last edit, so the monitoring side asks Proxmox directly: for each guest, is the newest auto- snapshot younger than eight hours? One qm listsnapshot per guest, run from the monitoring host through the API VIP, is enough.

Verify it worked

After the next scheduled run:

tail -n 20 /var/log/rbd-snap.log
qm listsnapshot 111
rbd -p rbd-guests snap ls vm-111-disk-0

The log should show one ok line per guest and the final all guests snapshotted line. qm listsnapshot should show three auto- entries at most, and the RBD image should carry the same three snapshots, which proves the snapshot is on the storage and not only in the guest config.

Then make it fail on purpose: rename /usr/sbin/qm for one run, or point PATH back at cron's default, and confirm root receives mail and the log says FAIL. A monitor you have never seen fire is a decoration.

Gotchas

FAQ

Are three RBD snapshots a day a backup? No. They are on the same pool, in the same cluster, behind the same credentials as the data. They cover operator error inside the guest. The nightly backup to the backup server at the other site covers everything else, and the decision post linked below is about that.

Why not the Proxmox backup schedule with snapshot mode? Because a backup job is a full copy to another storage, and three a day is a lot of traffic and retention for a protection that RBD gives for free. Snapshots and backups answer different questions.

Do snapshots slow the guest down? On RBD, the first write to each object after a snapshot is a copy, so there is a cost that fades through the day. Keeping only three and deleting the oldest before taking the next keeps the chain short.

Related