Proxmox: a cluster API address that follows a healthy node
In short. Install keepalived on every node, give them one VRRP instance with VRID 50 and unicast peers, and a
vrrp_scriptthat fetcheshttps://127.0.0.1:8006/and expects a 200. When pveproxy dies the node drops out and the address moves. Keep the VRID clear of the firewall's CARP VHIDs, because VRRP and CARP share IP protocol 112 on the same VLAN.
Goal: one stable address for the Proxmox API and web UI
Automation, the backup server and the humans all need an address for the cluster that does not depend on which node is up. A DNS round robin fails badly when one node's API is hung but its network stack is fine. This is Proxmox VE 9.2; keepalived is the Debian package.
What you need
- A three-node PVE 9.2 cluster with the management bond
bond2on VLAN 1. Nodes at10.1.1.51,.52,.53; the VIP will be10.1.1.50. - The list of VRRP VRIDs and CARP VHIDs already in use on the management VLAN, from the firewall's configuration.
curlon each node (already present on PVE).
Steps
- Install keepalived on every node.
apt install keepalived
- Write the health check.
-fmakes curl fail on any HTTP status of 400 or above,-kaccepts the node's self-signed certificate, and the timeout keeps a hung pveproxy from stalling keepalived.
# /usr/local/sbin/check-pveproxy
#!/bin/sh
exec curl -ksf --max-time 2 -o /dev/null https://127.0.0.1:8006/
chmod 755 /usr/local/sbin/check-pveproxy
chown root:root /usr/local/sbin/check-pveproxy
- Write the keepalived configuration. Each node has a different
router_id,priorityandunicast_src_ip; everything else is identical. With noweighton the script, a failed check puts the instance into FAULT and the node leaves the election entirely, which is what you want: a node with a broken API should not hold the address at any priority.
# /etc/keepalived/keepalived.conf on pve01
global_defs {
router_id pve01
enable_script_security
script_user root
}
vrrp_script chk_pveproxy {
script "/usr/local/sbin/check-pveproxy"
interval 2
timeout 3
fall 2
rise 2
}
vrrp_instance PVE_API {
state BACKUP
interface bond2
virtual_router_id 50
priority 150
advert_int 1
unicast_src_ip 10.1.1.51
unicast_peer {
10.1.1.52
10.1.1.53
}
virtual_ipaddress {
10.1.1.50/24 dev bond2
}
track_script {
chk_pveproxy
}
}
Use priority 150 on pve01, 140 on pve02, 130 on pve03, and swap the unicast_src_ip and peers accordingly. Every node starts as BACKUP and the election sorts out the master.
- Enable and start on every node, then look for the master.
systemctl enable --now keepalived
journalctl -u keepalived --since -1min | grep -E 'MASTER|BACKUP'
ip -br addr show bond2
flowchart LR
C["Clients, Ansible, PBS"] --> VIP["10.1.1.50<br/>VRID 50, unicast"]
VIP --> N1["pve01 prio 150<br/>check: curl 127.0.0.1:8006"]
VIP -.-> N2["pve02 prio 140"]
VIP -.-> N3["pve03 prio 130"]
FW["Firewall pair<br/>CARP VHIDs, protocol 112"] -. "same VLAN" .- VIPVerify it worked
Prove the check is what moves the address, not a node outage. On the current master, stop pveproxy and watch from a different node:
# on pve01 (master)
systemctl stop pveproxy
# on pve02
journalctl -u keepalived -f
curl -ks -o /dev/null -w '%{http_code}\n' https://10.1.1.50:8006/
Within about six seconds (two failed checks at two-second intervals, plus an advertisement interval) pve02 logs Entering MASTER STATE and the curl returns 200. Start pveproxy again on pve01 and watch the address return, since preemption is on by default. A web browser pointed at https://10.1.1.50:8006 should show the login page throughout, apart from one refresh.
Gotchas
- VRRP and CARP both use IP protocol 112 and both send to the same multicast group. If the firewall's CARP pair sits on the management VLAN with VHID 50, it will see your advertisements and may log or, worse, act on them. The VRID is kept numerically clear of every VHID, and unicast peers mean the firewall never sees the packets at all.
- A node whose API is hung still answers ARP and ping. A check that pings the node, or that only tests the kernel's view of the interface, keeps the address on a dead API. The check has to speak HTTPS to 8006.
- The node certificate does not include the VIP, so browsers warn on
10.1.1.50. Either accept that for a management address or issue a certificate that names it withpvenode cert set. enable_script_securityrefuses to run a script that is writable by anyone but root. If keepalived logs that the script is insecure, fix the mode; do not remove the option.- The VIP is for the API, not for corosync. Corosync has its own links and its own membership; do not point
pvecmat the VIP.
FAQ
Why keepalived rather than a DNS round robin or a load balancer? Round robin cannot tell a hung API from a healthy one. A load balancer is another machine to keep alive in front of the thing that keeps machines alive. Keepalived runs on the nodes themselves and needs nothing else.
Should the backup server use the VIP? Yes. PBS talks to the cluster API to enumerate guests, and the VIP means a backup job does not depend on one node. The backup data path is separate and goes straight to the node that owns the guest.
What if two nodes both think they are master? That is a split brain on the management VLAN, and it means the unicast peers cannot reach each other. Check the bond and the switch before touching keepalived. The cluster would be losing corosync at the same time, so you will have other symptoms.