Proxmox: a three-node Ceph cluster on used Dell servers
In short. Check the HBA is in IT mode before anything else. Install each node with a ZFS mirror boot, wire three bonds, form the cluster on the management VLAN, install Ceph with the storage VLAN as its network, add one monitor and one manager per node, one OSD per data disk, one replicated pool of size 3, and remove
local-zfsso nobody can put a guest disk on it. Then enrol the first VM in HA and migrate it, because a cluster you have not migrated on is not a cluster yet.
Goal: three used Dell servers into one Proxmox cluster with Ceph
You have three R630, R720xd or R730xd class machines, bought used, and you want shared storage without buying a SAN. The commands below are for Proxmox VE 9.2 and Ceph Squid 19.2. Each step ends with the thing to look at before moving on.
What you need
- Three nodes, each with an IT-mode HBA (SAS2308, SAS3008 or similar), two small SSDs for boot and at least two new drives for OSDs. A node whose RAID controller cannot be flashed to IT mode is not a Ceph node; it becomes the backup server.
- A switch stack that can do LACP across members, and either a separate 10G switch or a tagged storage VLAN for Ceph.
- Proxmox VE 9.2 installed on each node (the automated install post below does that), Ceph Squid 19.2 from the Proxmox repository.
- The addressing plan: management VLAN 1, nodes at
10.1.1.51to.53, storage VLAN 5 at10.1.5.51to.53, gateway.254.
Steps
- Confirm the controller presents raw disks. The kernel should see the HBA as a plain SAS adapter and
lsblkshould list every physical drive with its real model and serial. If you see one large virtual disk, the controller is still in RAID mode.
lspci | grep -i 'SAS\|RAID'
lsblk -o NAME,SIZE,MODEL,SERIAL,ROTA,TYPE
for d in /dev/sd?; do smartctl -H "$d" | grep -E 'overall|result'; done
- Write the network configuration on each node. Three bonds:
bond2for management,bond0under the VLAN-awarevmbr0for guests,bond1for Ceph with jumbo frames. Only the addresses differ between nodes.
# /etc/network/interfaces (pve01, abridged)
auto bond2
iface bond2 inet static
address 10.1.1.51/24
gateway 10.1.1.254
bond-slaves eno3 eno4
bond-mode active-backup
auto bond0
iface bond0 inet manual
bond-slaves eno1 eno2
bond-mode 802.3ad
bond-xmit-hash-policy layer3+4
auto vmbr0
iface vmbr0 inet manual
bridge-ports bond0
bridge-stp off
bridge-fd 0
bridge-vlan-aware yes
bridge-vids 2-4094
auto bond1
iface bond1 inet manual
bond-slaves enp5s0f0 enp5s0f1
bond-mode 802.3ad
mtu 9000
auto bond1.5
iface bond1.5 inet static
address 10.1.5.51/24
mtu 9000
Apply with ifreload -a and check jumbo frames end to end with ping -M do -s 8972 10.1.5.52.
- Create the cluster on the first node, then join the others. Corosync rides the management VLAN; the storage VLAN is a second link so a management outage alone cannot break quorum.
# on pve01
pvecm create sitea --link0 10.1.1.51 --link1 10.1.5.51
# on pve02 and pve03
pvecm add 10.1.1.51 --link0 10.1.1.52 --link1 10.1.5.52
- Install Ceph on every node, then initialise it once with the storage VLAN as the cluster and public network.
pveceph install --repository no-subscription --version squid # every node
pveceph init --network 10.1.5.0/24 # pve01 only
- One monitor and one manager per node.
pveceph mon create # on each node
pveceph mgr create # on each node
- One OSD per data disk, on raw devices. Repeat per disk, per node.
pveceph osd create /dev/sdc
pveceph osd create /dev/sdd
- One replicated pool, size 3, min size 2, and let
--add_storagesregister it as the RBD storage.
pveceph pool create rbd-guests --size 3 --min_size 2 --pg_autoscale_mode on --add_storages
pvesm status
- Ban local storage for guests. The installer registers
local-zfsfor the boot pool; remove it and restrictlocalto things that are not guest disks.
pvesm remove local-zfs
pvesm set local --content iso,vztmpl,snippets
- First VM, HA-enrolled. Clone a template (the cloud-init post covers building one) and add it as an HA resource. With three identical nodes no HA group or node-affinity rule is needed; any node may run any guest.
qm clone 9000 100 --full --name canary
qm set 100 --net0 virtio=BC:24:11:65:01:0B,bridge=vmbr0,tag=101
ha-manager add vm:100 --state started
flowchart TB
subgraph Node["Each node"]
B2["bond2<br/>management VLAN 1<br/>corosync link0"]
B0["bond0 LACP<br/>vmbr0 VLAN-aware<br/>guest traffic"]
B1["bond1 LACP<br/>bond1.5, MTU 9000<br/>Ceph, corosync link1"]
end
B2 --> MGMT["Management VLAN<br/>API VIP .50, nodes .51+"]
B0 --> STACK["Switch stack"]
B1 --> STOR["Storage VLAN 5<br/>mon, mgr, OSD"]
STOR --> POOL["rbd-guests<br/>size 3, min 2"]Verify it worked
pvecm status | grep -E 'Quorate|Nodes'
ceph -s
ceph osd tree
ha-manager status
qm migrate 100 pve02 --online
You want Quorate: Yes with three nodes, HEALTH_OK with three monitors, three managers and all OSDs up, and a live migration that finishes with the VM still answering ping. If the migration fails with a storage error, a disk is on local and step 8 was skipped.
Gotchas
- An HBA that cannot be flashed to IT mode gives Ceph a cached, lying block device. Do not try to work around it; move that chassis to backup duty.
- Monitor addresses are fixed at creation. If you change the storage VLAN later, you are destroying and recreating monitors, so get the addressing right first.
- Size 3 with min size 2 means one node down is fine and two nodes down stops writes. That is the correct behaviour for three nodes; do not lower
min_sizeto make a bad day look better. pveceph pool createwithout--pg_autoscale_mode onleaves you choosing PG counts by hand. Let the autoscaler do it on a cluster this small.- Used drives for OSDs save money once and cost it back in rebalances. Boot SSDs can be used; OSD drives are bought new.
FAQ
Why not local-zfs with replication instead of Ceph? ZFS replication is asynchronous and per guest. HA, migration and backup all assume every node can see every disk right now. With three nodes, Ceph gives that for the price of one bond.
Can I run this with two nodes and a QDevice? You can form quorum that way, but Ceph with size 3 needs three hosts to place replicas. Two nodes means size 2, and size 2 loses data on a double fault. Three is the minimum that makes the storage decision honest.
Does the cluster need the second corosync link? Not strictly. It costs one line per node and it is the difference between a management switch reboot and a fencing event.