Proxmox: replace a failed OSD, or reinstall and rejoin a node

In short. Ceph tells you which OSD is down; ceph-volume lvm list tells you which disk. Mark it out, stop it, pveceph osd destroy, swap the drive, pveceph osd create, and watch ceph -s until it is clean. A whole node is the same sequence with the monitor, manager and cluster membership added, and the reinstall comes from the answer file so the node returns identical to how it left.

"1 osds down" or "HEALTH_WARN 1 host down": what to do on a three-node cluster

A three-node cluster with size 3 pools has no spare host. While one node or one disk is out, every write is still safe (min size 2) but you have no margin. The job is to restore the margin without making it worse. Proxmox VE 9.2, Ceph Squid 19.2.

What you need

Steps

A failed OSD

  1. Identify the OSD and the physical disk behind it.
ceph health detail
ceph osd tree | grep down
ceph-volume lvm list | grep -B2 -A8 'osd id.*7'   # on the node that owns osd.7
smartctl -a /dev/sdd | grep -iE 'reallocated|pending|result'
  1. Mark it out and stop it. Out triggers rebalancing of its placement groups to the other OSDs on the same host, which keeps three copies across three hosts.
ceph osd out 7
systemctl stop ceph-osd@7
  1. Destroy it through Proxmox so the cluster map, the keyring and the LVM volumes go together. --cleanup wipes the disk if it is still readable; if it is not, the command will say so and you carry on.
pveceph osd destroy 7 --cleanup
  1. Swap the drive. On an R730xd backplane the slot light helps; ls -l /dev/disk/by-path/ before and after tells you which /dev/sdX the new drive took.

  2. Create the OSD on the raw device. Proxmox handles the LVM, the keyring and the CRUSH placement under the host.

pveceph osd create /dev/sdd
  1. Watch recovery until the cluster is clean.
watch -n 5 ceph -s

A failed or compromised node

  1. Drain it. Enable maintenance so the HA manager migrates everything off and will not place anything back.
ha-manager crm-command node-maintenance enable pve03
  1. Remove its Ceph roles, from the node if it is up, otherwise from a survivor. Out and destroy each of its OSDs as above; if the node is gone for good, ceph osd purge <id> --yes-i-really-mean-it removes a dead OSD from the map without a host to run pveceph on.
pveceph mds destroy pve03    # only if CephFS is in use
pveceph mon destroy pve03
pveceph mgr destroy pve03
  1. Remove it from the cluster. Power it off first; a node that is removed and still running corosync will confuse the survivors.
pvecm delnode pve03
rm -r /etc/pve/nodes/pve03
  1. Reinstall it by PXE with the answer file. Same hostname, same management and storage addresses, same bonds from the baseline. Run the configuration baseline before going further.

  2. Rejoin, from the new node, pointing at a survivor.

pvecm add 10.1.1.51 --link0 10.1.1.53 --link1 10.1.5.53
  1. Recreate the Ceph roles and the OSDs.
pveceph install --repository no-subscription --version squid
pveceph mon create
pveceph mgr create
pveceph mds create
pveceph osd create /dev/sdc
pveceph osd create /dev/sdd
  1. Clear maintenance once Ceph is clean and the CephFS mount units are active.
ha-manager crm-command node-maintenance disable pve03
flowchart LR
  D["OSD down"] --> O["out, stop"] --> X["pveceph osd destroy"] --> S["swap disk"] --> C["pveceph osd create"] --> R["watch recovery"]
  N["Node lost"] --> M["node-maintenance enable"] --> RR["mon, mgr, mds, OSDs destroyed"] --> DN["pvecm delnode"] --> PX["PXE reinstall"] --> AD["pvecm add"] --> RC["roles recreated"] --> R

Verify it worked

ceph -s
ceph osd tree
pvecm status | grep -E 'Quorate|Nodes'
ha-manager status | grep pve03
systemctl is-active 'mnt-cephfs-*.mount'

HEALTH_OK, every OSD up and in under the right host, three quorate nodes, the node listed as online by HA, and the mount units active. Then migrate one guest onto the node and back, because a node that is in the cluster but cannot run a guest is not finished.

Gotchas

FAQ

Can I skip out and go straight to destroy? pveceph osd destroy refuses an OSD that is not out and stopped. The order exists so the cluster rebalances before the data holder disappears, and on a dead disk it is still the order that keeps the map consistent.

Why reinstall rather than repair the node? Because the install takes twenty minutes from the answer file and leaves no mystery. A repaired node carries whatever the failure did to it; a reinstalled one matches the other two.

Does a reinstalled node need the same name? It needs the same addresses because the monitor map and the mount units use them, and the same name because the guest configs reference nodes by name. Remove /etc/pve/nodes/<name> on a survivor before the rejoin or the stale directory collides.

Related