OPNsense HA that actually fails over: VHID 1 everywhere, and the sync that is not automatic

Isometric illustration of two identical firewall units, one wearing a crown, with a half-printed sheet of paper passing between them under a wall clock.

In short. Every CARP VIP is VHID 1. The WAN has one CARP VIP and every other public address is an alias on it. Config sync from master to backup is not triggered by API writes; only an explicit filter sync runs it. Apply changes to the backup first, verify, then the master plus sync, and audit for drift on a timer.

The pair

Two identical firewalls. On every VLAN, fw01 is .252, fw02 is .253, and the shared gateway is .254. Guests only ever know .254. If the master dies, the backup answers for .254 within a second or two and nothing downstream notices.

That is the brochure. The rest of this post is what it took to make the brochure true.

VHID numbering: why "VHID equals VLAN" is a trap

The obvious scheme is to set each VLAN's CARP virtual host ID to the VLAN number. It works until it does not:

The replacement, recorded as a decision:

Sync transport

The first site used a crossover cable between the pair for state sync. Later sites have a dedicated sync VLAN through the switch stack, so state and config sync do not share the management network and the pair can be in different racks.

Config sync is not automatic

This is the finding that justified the post. The high-availability sync from master to backup, done over XMLRPC, is triggered by the GUI when you save a rule. It is not triggered by a write through the API. A tool that updates rules through the API changes fw01 and leaves fw02 exactly as it was.

The symptom: weeks of rules on the master only. The backup looked healthy, had the same VLANs and VIPs, and would have failed over to a policy from a month earlier.

The fix has three parts.

  1. The change tool applies to the backup first, reads back and verifies, then applies to the master and calls configctl filter sync explicitly.
  2. A drift audit runs on a timer, fetches the rule set from both nodes, diffs them, and alerts to a push channel.
  3. CARP VIP deletes are done on both nodes by hand, because deletes do not sync even when creates do.
sequenceDiagram
  participant Op as Change tool
  participant B as fw02 backup
  participant A as fw01 master
  participant D as Drift audit timer
  Op->>B: apply rule change via API
  Op->>B: read back, verify
  Op->>A: apply the same change
  Op->>A: configctl filter sync
  A-->>B: XMLRPC config push
  Note over A,B: No API write triggers this sync on its own
  D->>A: fetch rules
  D->>B: fetch rules
  D-->>Op: diff, alert on drift

Other things a pair needs

The hypervisors' VIP is on the same wire

The cluster API address on the hypervisors uses VRRP. VRRP and CARP both speak IP protocol 112 on the same layer 2, and they do not know about each other. The VRRP ID is kept numerically clear of every CARP VHID, and the VRRP instance has a track script on the API service because, as the decision record puts it, CARP hides userland failure: a node whose API is wedged still wins the election.

When a cleanup takes the site down

An "unused" provider-assigned alias was deleted during a tidy-up. It was the next-hop the provider used for a routed /28, and no local check could see that dependency. The routed block was unreachable for about 26 hours.

Rules that came out of it:

FAQ

Does the GUI's "synchronise config to backup" button cover this? Yes, for people. The problem is tools. If any automation writes to the firewall, it must call the sync itself.

Is pfsync state sync affected? No. Connection state syncs continuously. It is the configuration that does not.

How often does the drift audit run? Every fifteen minutes, with a quiet period after a known change so it does not alert on its own work.

Related