Ceph
8 posts on Ceph: design decisions and how-to guides from small infrastructure estates run on a very low budget.
- PBS: back up a directory with proxmox-backup-client from a snapshot
vzdump skips bind mounts, so the files a container serves from CephFS never reach the backup server. A nightly proxmox-backup-client run takes a pxar archive of the latest CephFS snapshot into the tenant's namespace. The token, the command, the timer, and the one-file restore that proves it.
- Proxmox: bind-mount CephFS into an unprivileged container
Create a CephFS subvolume with a quota, scope a cephx key to its path, mount it on every node with a systemd unit, and bind-mount it into an unprivileged container that holds no Ceph key. df inside the container shows the quota, not the cluster.
- Proxmox: fix "cannot migrate local bind mount point"
The error means Proxmox cannot promise the directory behind mp0 exists on the target node. Mount the same path on every node with a systemd unit, check it, and only then add shared=1.
CephFS subvolumes under LXC: snapshots your backups can seeGuests are disposable; the data is not in them. Put user files on a CephFS subvolume, bind-mount it into an unprivileged container, snapshot on a schedule, and back the snapshot up, because vzdump skips bind mounts.
- Proxmox: replace a failed OSD, or reinstall and rejoin a node
Two runbooks that share a shape. For a dead disk: identify, out, stop, destroy, swap, create, watch recovery. For a dead node: drain it, remove its Ceph roles, pvecm delnode, reinstall by PXE, pvecm add, recreate the roles, and keep HA fencing in mind throughout.
- Proxmox: a three-node Ceph cluster on used Dell servers
From three second-hand R630-class servers with IT-mode HBAs to a quorate cluster with one replicated RBD pool, HA on every guest and no local storage for guests. The commands, in order, and what to check after each.
Ex-lease Dell Servers, one Ceph pool, no local-zfsSecond-hand servers are cheap. What costs money later is a storage decision that blocks migration. Here is the cluster layout that keeps HA, migration and backup working.
- Proxmox: intra-day snapshots from cron that actually run
A snapshot script that ran from cron three times a day, reported success for months, and never created a snapshot, because cron's PATH does not include /usr/sbin and the output went to /dev/null. The fixed script, and the checks that make it fail loudly.