Proxmox: a cluster API address that follows a healthy node

In short. Install keepalived on every node, give them one VRRP instance with VRID 50 and unicast peers, and a vrrp_script that fetches https://127.0.0.1:8006/ and expects a 200. When pveproxy dies the node drops out and the address moves. Keep the VRID clear of the firewall's CARP VHIDs, because VRRP and CARP share IP protocol 112 on the same VLAN.

Goal: one stable address for the Proxmox API and web UI

Automation, the backup server and the humans all need an address for the cluster that does not depend on which node is up. A DNS round robin fails badly when one node's API is hung but its network stack is fine. This is Proxmox VE 9.2; keepalived is the Debian package.

What you need

Steps

  1. Install keepalived on every node.
apt install keepalived
  1. Write the health check. -f makes curl fail on any HTTP status of 400 or above, -k accepts the node's self-signed certificate, and the timeout keeps a hung pveproxy from stalling keepalived.
# /usr/local/sbin/check-pveproxy
#!/bin/sh
exec curl -ksf --max-time 2 -o /dev/null https://127.0.0.1:8006/
chmod 755 /usr/local/sbin/check-pveproxy
chown root:root /usr/local/sbin/check-pveproxy
  1. Write the keepalived configuration. Each node has a different router_id, priority and unicast_src_ip; everything else is identical. With no weight on the script, a failed check puts the instance into FAULT and the node leaves the election entirely, which is what you want: a node with a broken API should not hold the address at any priority.
# /etc/keepalived/keepalived.conf on pve01
global_defs {
    router_id pve01
    enable_script_security
    script_user root
}

vrrp_script chk_pveproxy {
    script "/usr/local/sbin/check-pveproxy"
    interval 2
    timeout 3
    fall 2
    rise 2
}

vrrp_instance PVE_API {
    state BACKUP
    interface bond2
    virtual_router_id 50
    priority 150
    advert_int 1
    unicast_src_ip 10.1.1.51
    unicast_peer {
        10.1.1.52
        10.1.1.53
    }
    virtual_ipaddress {
        10.1.1.50/24 dev bond2
    }
    track_script {
        chk_pveproxy
    }
}

Use priority 150 on pve01, 140 on pve02, 130 on pve03, and swap the unicast_src_ip and peers accordingly. Every node starts as BACKUP and the election sorts out the master.

  1. Enable and start on every node, then look for the master.
systemctl enable --now keepalived
journalctl -u keepalived --since -1min | grep -E 'MASTER|BACKUP'
ip -br addr show bond2
flowchart LR
  C["Clients, Ansible, PBS"] --> VIP["10.1.1.50<br/>VRID 50, unicast"]
  VIP --> N1["pve01 prio 150<br/>check: curl 127.0.0.1:8006"]
  VIP -.-> N2["pve02 prio 140"]
  VIP -.-> N3["pve03 prio 130"]
  FW["Firewall pair<br/>CARP VHIDs, protocol 112"] -. "same VLAN" .- VIP

Verify it worked

Prove the check is what moves the address, not a node outage. On the current master, stop pveproxy and watch from a different node:

# on pve01 (master)
systemctl stop pveproxy
# on pve02
journalctl -u keepalived -f
curl -ks -o /dev/null -w '%{http_code}\n' https://10.1.1.50:8006/

Within about six seconds (two failed checks at two-second intervals, plus an advertisement interval) pve02 logs Entering MASTER STATE and the curl returns 200. Start pveproxy again on pve01 and watch the address return, since preemption is on by default. A web browser pointed at https://10.1.1.50:8006 should show the login page throughout, apart from one refresh.

Gotchas

FAQ

Why keepalived rather than a DNS round robin or a load balancer? Round robin cannot tell a hung API from a healthy one. A load balancer is another machine to keep alive in front of the thing that keeps machines alive. Keepalived runs on the nodes themselves and needs nothing else.

Should the backup server use the VIP? Yes. PBS talks to the cluster API to enumerate guests, and the VIP means a backup job does not depend on one node. The backup data path is separate and goes straight to the node that owns the guest.

What if two nodes both think they are master? That is a split brain on the management VLAN, and it means the unicast peers cannot reach each other. Check the bond and the switch before touching keepalived. The cluster would be losing corosync at the same time, so you will have other symptoms.

Related