OPNsense HA that actually fails over: VHID 1 everywhere, and the sync that is not automatic

In short. Every CARP VIP is VHID 1. The WAN has one CARP VIP and every other public address is an alias on it. Config sync from master to backup is not triggered by API writes; only an explicit filter sync runs it. Apply changes to the backup first, verify, then the master plus sync, and audit for drift on a timer.
The pair
Two identical firewalls. On every VLAN, fw01 is .252, fw02 is .253, and the shared gateway is .254. Guests only ever know .254. If the master dies, the backup answers for .254 within a second or two and nothing downstream notices.
That is the brochure. The rest of this post is what it took to make the brochure true.
VHID numbering: why "VHID equals VLAN" is a trap
The obvious scheme is to set each VLAN's CARP virtual host ID to the VLAN number. It works until it does not:
- CARP builds a virtual MAC from the VHID. The VHID field is one byte. At VLAN 256 the generated MAC is malformed and the VIP silently stops working.
- The WAN needs a VIP per public address. With VHID per address you run out at 255, and every added address is a new election to debug.
The replacement, recorded as a decision:
- Every CARP VIP is VHID 1. Switches learn MAC addresses per VLAN, so the same virtual MAC on every VLAN is fine. The change was gated on proving that on the actual switch stack before touching production.
- The WAN has exactly one CARP VIP. Every other public address is an IP alias on that VIP. All public addresses fail over as one group, in one election.
Sync transport
The first site used a crossover cable between the pair for state sync. Later sites have a dedicated sync VLAN through the switch stack, so state and config sync do not share the management network and the pair can be in different racks.
Config sync is not automatic
This is the finding that justified the post. The high-availability sync from master to backup, done over XMLRPC, is triggered by the GUI when you save a rule. It is not triggered by a write through the API. A tool that updates rules through the API changes fw01 and leaves fw02 exactly as it was.
The symptom: weeks of rules on the master only. The backup looked healthy, had the same VLANs and VIPs, and would have failed over to a policy from a month earlier.
The fix has three parts.
- The change tool applies to the backup first, reads back and verifies, then applies to the master and calls
configctl filter syncexplicitly. - A drift audit runs on a timer, fetches the rule set from both nodes, diffs them, and alerts to a push channel.
- CARP VIP deletes are done on both nodes by hand, because deletes do not sync even when creates do.
sequenceDiagram
participant Op as Change tool
participant B as fw02 backup
participant A as fw01 master
participant D as Drift audit timer
Op->>B: apply rule change via API
Op->>B: read back, verify
Op->>A: apply the same change
Op->>A: configctl filter sync
A-->>B: XMLRPC config push
Note over A,B: No API write triggers this sync on its own
D->>A: fetch rules
D->>B: fetch rules
D-->>Op: diff, alert on driftOther things a pair needs
- Certificates renew on the master only. The ACME client runs where it was configured. A daily job copies the renewed certificate to the backup, and the audit checks the two fingerprints match.
- The API cannot assign interfaces. Creating a VLAN interface and giving it an address is a GUI step. It has a checklist, and the checklist ends with "do it on both".
- After any WAN change, probe every published port on every public address from two vantage points. Inside the estate is not a vantage point; split-horizon DNS will test the internal path.
The hypervisors' VIP is on the same wire
The cluster API address on the hypervisors uses VRRP. VRRP and CARP both speak IP protocol 112 on the same layer 2, and they do not know about each other. The VRRP ID is kept numerically clear of every CARP VHID, and the VRRP instance has a track script on the API service because, as the decision record puts it, CARP hides userland failure: a node whose API is wedged still wins the election.
When a cleanup takes the site down
An "unused" provider-assigned alias was deleted during a tidy-up. It was the next-hop the provider used for a routed /28, and no local check could see that dependency. The routed block was unreachable for about 26 hours.
Rules that came out of it:
- Public addresses are retired by a tool that re-reads the config and warns about every rule that still references the address, including hand-made ones it did not create.
- A reference guard checks outbound NAT, inbound NAT, filter rules and BGP before a VIP is removed.
- Nothing is "unused" until the provider says so in writing.
FAQ
Does the GUI's "synchronise config to backup" button cover this? Yes, for people. The problem is tools. If any automation writes to the firewall, it must call the sync itself.
Is pfsync state sync affected? No. Connection state syncs continuously. It is the configuration that does not.
How often does the drift audit run? Every fifteen minutes, with a quiet period after a known change so it does not alert on its own work.