Proxmox: a three-node Ceph cluster on used Dell servers

In short. Check the HBA is in IT mode before anything else. Install each node with a ZFS mirror boot, wire three bonds, form the cluster on the management VLAN, install Ceph with the storage VLAN as its network, add one monitor and one manager per node, one OSD per data disk, one replicated pool of size 3, and remove local-zfs so nobody can put a guest disk on it. Then enrol the first VM in HA and migrate it, because a cluster you have not migrated on is not a cluster yet.

Goal: three used Dell servers into one Proxmox cluster with Ceph

You have three R630, R720xd or R730xd class machines, bought used, and you want shared storage without buying a SAN. The commands below are for Proxmox VE 9.2 and Ceph Squid 19.2. Each step ends with the thing to look at before moving on.

What you need

Steps

  1. Confirm the controller presents raw disks. The kernel should see the HBA as a plain SAS adapter and lsblk should list every physical drive with its real model and serial. If you see one large virtual disk, the controller is still in RAID mode.
lspci | grep -i 'SAS\|RAID'
lsblk -o NAME,SIZE,MODEL,SERIAL,ROTA,TYPE
for d in /dev/sd?; do smartctl -H "$d" | grep -E 'overall|result'; done
  1. Write the network configuration on each node. Three bonds: bond2 for management, bond0 under the VLAN-aware vmbr0 for guests, bond1 for Ceph with jumbo frames. Only the addresses differ between nodes.
# /etc/network/interfaces (pve01, abridged)
auto bond2
iface bond2 inet static
    address 10.1.1.51/24
    gateway 10.1.1.254
    bond-slaves eno3 eno4
    bond-mode active-backup

auto bond0
iface bond0 inet manual
    bond-slaves eno1 eno2
    bond-mode 802.3ad
    bond-xmit-hash-policy layer3+4

auto vmbr0
iface vmbr0 inet manual
    bridge-ports bond0
    bridge-stp off
    bridge-fd 0
    bridge-vlan-aware yes
    bridge-vids 2-4094

auto bond1
iface bond1 inet manual
    bond-slaves enp5s0f0 enp5s0f1
    bond-mode 802.3ad
    mtu 9000

auto bond1.5
iface bond1.5 inet static
    address 10.1.5.51/24
    mtu 9000

Apply with ifreload -a and check jumbo frames end to end with ping -M do -s 8972 10.1.5.52.

  1. Create the cluster on the first node, then join the others. Corosync rides the management VLAN; the storage VLAN is a second link so a management outage alone cannot break quorum.
# on pve01
pvecm create sitea --link0 10.1.1.51 --link1 10.1.5.51
# on pve02 and pve03
pvecm add 10.1.1.51 --link0 10.1.1.52 --link1 10.1.5.52
  1. Install Ceph on every node, then initialise it once with the storage VLAN as the cluster and public network.
pveceph install --repository no-subscription --version squid   # every node
pveceph init --network 10.1.5.0/24                             # pve01 only
  1. One monitor and one manager per node.
pveceph mon create   # on each node
pveceph mgr create   # on each node
  1. One OSD per data disk, on raw devices. Repeat per disk, per node.
pveceph osd create /dev/sdc
pveceph osd create /dev/sdd
  1. One replicated pool, size 3, min size 2, and let --add_storages register it as the RBD storage.
pveceph pool create rbd-guests --size 3 --min_size 2 --pg_autoscale_mode on --add_storages
pvesm status
  1. Ban local storage for guests. The installer registers local-zfs for the boot pool; remove it and restrict local to things that are not guest disks.
pvesm remove local-zfs
pvesm set local --content iso,vztmpl,snippets
  1. First VM, HA-enrolled. Clone a template (the cloud-init post covers building one) and add it as an HA resource. With three identical nodes no HA group or node-affinity rule is needed; any node may run any guest.
qm clone 9000 100 --full --name canary
qm set 100 --net0 virtio=BC:24:11:65:01:0B,bridge=vmbr0,tag=101
ha-manager add vm:100 --state started
flowchart TB
  subgraph Node["Each node"]
    B2["bond2<br/>management VLAN 1<br/>corosync link0"]
    B0["bond0 LACP<br/>vmbr0 VLAN-aware<br/>guest traffic"]
    B1["bond1 LACP<br/>bond1.5, MTU 9000<br/>Ceph, corosync link1"]
  end
  B2 --> MGMT["Management VLAN<br/>API VIP .50, nodes .51+"]
  B0 --> STACK["Switch stack"]
  B1 --> STOR["Storage VLAN 5<br/>mon, mgr, OSD"]
  STOR --> POOL["rbd-guests<br/>size 3, min 2"]

Verify it worked

pvecm status | grep -E 'Quorate|Nodes'
ceph -s
ceph osd tree
ha-manager status
qm migrate 100 pve02 --online

You want Quorate: Yes with three nodes, HEALTH_OK with three monitors, three managers and all OSDs up, and a live migration that finishes with the VM still answering ping. If the migration fails with a storage error, a disk is on local and step 8 was skipped.

Gotchas

FAQ

Why not local-zfs with replication instead of Ceph? ZFS replication is asynchronous and per guest. HA, migration and backup all assume every node can see every disk right now. With three nodes, Ceph gives that for the price of one bond.

Can I run this with two nodes and a QDevice? You can form quorum that way, but Ceph with size 3 needs three hosts to place replicas. Two nodes means size 2, and size 2 loses data on a double fault. Three is the minimum that makes the storage decision honest.

Does the cluster need the second corosync link? Not strictly. It costs one line per node and it is the difference between a management switch reboot and a fencing event.

Related