Ex-lease Dell Servers, one Ceph pool, no local-zfs

In short. Minimum of three R630 or R720xd class nodes with IT-mode HBAs. A VLAN-aware bridge on an LACP bond for guests, a second bond for Ceph. One replicated RBD pool for every guest disk. Every guest HA-enrolled. Node-local storage banned, in writing, with reasons.
Hardware rules
- IT-mode HBAs or nothing. Ceph wants raw disks. A RAID controller that cannot be flashed to pass-through mode goes in the backup server, where it does not matter.
- OSD drives are bought new. Everything else can be second-hand. A used chassis fails loudly and gets swapped; a used disk fails slowly and takes data with it.
- Reuse the chassis, not the design. The first cluster's R720xd nodes, with 24 bays and SAS2308 controllers, became the third site's nodes. Their original storage layout did not come with them.
- Out-of-band licences cost almost nothing on the second-hand market. An enterprise iDRAC licence is a ten to twenty dollar item and it turns a server you have to visit into one you can reinstall from your desk.
- Power is a design input. A home staging rack tripped its circuit with the firewalls and two servers powered on together. Count amps before you count cores.
Node networking
flowchart LR
subgraph Node["PVE node: ex-lease R630 or R720xd, IT-mode HBA"]
b0["bond0 LACP"] --> v0["vmbr0 VLAN-aware<br/>guest and management traffic"]
b1["bond1"] --> v1["vmbr1 or bond1.5<br/>Ceph public + cluster"]
v0 --> G["Guests: VMs and LXC<br/>every disk on one RBD pool<br/>every guest HA-enrolled"]
v1 --> OSD["OSDs, new drives only"]
K["keepalived VRID 50<br/>API VIP with pveproxy check"] -.-> v0
end
v0 --> SW["Switch stack"]
v1 --> SW
v1 --> PBS["Backup server on VLAN 5"]
X["local-zfs: banned<br/>breaks HA, migration, backup"]vmbr0is VLAN-aware onbond0, an LACP bundle across both stack members. Guests get a VLAN tag on their virtual NIC; the bridge carries all of them.- Ceph gets its own bond. At the first site that is an active-backup pair on a separate unmanaged 10G switch. At later sites it is LACP to the stack, tagged for the storage VLAN as
bond1.5. - Public and cluster networks are the same VLAN. Splitting them on a three-node cluster adds a failure mode and buys nothing measurable.
- The boot disk is a ZFS mirror on two small SSDs, which is the one place ZFS is allowed.
Storage layout
One replicated RBD pool, minimum size 3, holds every guest disk. CephFS is also standard on every cluster for bulk file data, with one filesystem per cluster and isolation by subvolume; that gets its own post.
Rule 13 in the cluster standard reads: no local-zfs, no local-lvm for guests. The reasons, so nobody re-litigates it:
- HA cannot restart a guest on another node if its disk is on the dead one.
- Live migration has to copy the disk, which turns a ten-second move into an hour.
- Backups from local storage compete with the guest for the same spindles.
A guest on local storage is a pet with a chronic condition. The tooling that creates clusters refuses to create one.
Every guest is HA-enrolled
The exception list is empty. If a guest is important enough to exist, it is important enough to restart somewhere else when its node dies. The anti-affinity rules keep replicas of the same service on different nodes.
The cluster API address
Operators and tools talk to a cluster VIP, not to a node. It is keepalived with VRRP, with two details that matter:
- The VRID is kept numerically clear of the firewall's CARP IDs. VRRP and CARP both use IP protocol 112 on the same layer 2. A collision is not an error; it is a silent election between two things that do not know about each other.
- It has a track script on the API service. CARP-style failover only notices the host going away. A node whose API daemon is wedged still answers ARP. The script moves the VIP when the API stops answering, not when the node stops pinging.
Provisioning from the jumphost
Each site's jumphost serves PXE and hosts the install media for out-of-band virtual media. Two things cost an afternoon each:
- Older out-of-band controllers reject HTTP/1.0. Serve virtual media from something that speaks HTTP/1.1.
- The install answer file is generated, not edited. A hand-patched network interface on one node was overwritten by the next install. Fix the template.
What was evaluated and rejected
- A container-native cloud stack was evaluated on paper against Proxmox. It lost on operational familiarity and on backup tooling. The decision record notes the one gap Proxmox left open: per-tenant inbound NAT still lives on the firewall by hand.
- Nested virtualisation for the lab. No. The lab is a smaller copy of production on real hardware, or it is not a lab.
FAQ
Three nodes is enough? Three is the minimum for Ceph quorum and for the cluster vote. It is also the number where a single failure leaves you with a working, if nervous, cluster.
Why not SSDs for OSDs across the board? Budget. Spinning OSDs with SSD boot and a fast network are adequate for a small estate. The money went to a second site instead.
Which Proxmox and Ceph versions? Proxmox VE 9.2 with Ceph Squid at the time of writing. Versions are pinned in the inventory and the lab is always upgraded first.