Ex-lease Dell Servers, one Ceph pool, no local-zfs

Isometric illustration of three open rack servers with disks visible and three connected spheres floating above them.

In short. Minimum of three R630 or R720xd class nodes with IT-mode HBAs. A VLAN-aware bridge on an LACP bond for guests, a second bond for Ceph. One replicated RBD pool for every guest disk. Every guest HA-enrolled. Node-local storage banned, in writing, with reasons.

Hardware rules

Node networking

flowchart LR
  subgraph Node["PVE node: ex-lease R630 or R720xd, IT-mode HBA"]
    b0["bond0 LACP"] --> v0["vmbr0 VLAN-aware<br/>guest and management traffic"]
    b1["bond1"] --> v1["vmbr1 or bond1.5<br/>Ceph public + cluster"]
    v0 --> G["Guests: VMs and LXC<br/>every disk on one RBD pool<br/>every guest HA-enrolled"]
    v1 --> OSD["OSDs, new drives only"]
    K["keepalived VRID 50<br/>API VIP with pveproxy check"] -.-> v0
  end
  v0 --> SW["Switch stack"]
  v1 --> SW
  v1 --> PBS["Backup server on VLAN 5"]
  X["local-zfs: banned<br/>breaks HA, migration, backup"]

Storage layout

One replicated RBD pool, minimum size 3, holds every guest disk. CephFS is also standard on every cluster for bulk file data, with one filesystem per cluster and isolation by subvolume; that gets its own post.

Rule 13 in the cluster standard reads: no local-zfs, no local-lvm for guests. The reasons, so nobody re-litigates it:

  1. HA cannot restart a guest on another node if its disk is on the dead one.
  2. Live migration has to copy the disk, which turns a ten-second move into an hour.
  3. Backups from local storage compete with the guest for the same spindles.

A guest on local storage is a pet with a chronic condition. The tooling that creates clusters refuses to create one.

Every guest is HA-enrolled

The exception list is empty. If a guest is important enough to exist, it is important enough to restart somewhere else when its node dies. The anti-affinity rules keep replicas of the same service on different nodes.

The cluster API address

Operators and tools talk to a cluster VIP, not to a node. It is keepalived with VRRP, with two details that matter:

Provisioning from the jumphost

Each site's jumphost serves PXE and hosts the install media for out-of-band virtual media. Two things cost an afternoon each:

What was evaluated and rejected

FAQ

Three nodes is enough? Three is the minimum for Ceph quorum and for the cluster vote. It is also the number where a single failure leaves you with a working, if nervous, cluster.

Why not SSDs for OSDs across the board? Budget. Spinning OSDs with SSD boot and a fast network are adequate for a small estate. The money went to a second site instead.

Which Proxmox and Ceph versions? Proxmox VE 9.2 with Ceph Squid at the time of writing. Versions are pinned in the inventory and the lab is always upgraded first.

Related