OPNsense: fix HA config sync that stopped syncing

In short. OPNsense pushes configuration from master to backup only when something asks it to: a save in the GUI, or configctl filter sync from a shell. Writes through the API do not ask. Compare the rule sets on both nodes, check the sync settings and credentials, run the sync by hand, and then put a drift check on a timer so it cannot creep back.

My backup firewall has fewer rules than the master

The symptom is quiet. Both nodes are up, CARP shows one master and one backup, pfsync is happy and a failover test passes. Then you open the backup's rule list and it is missing everything added since some date you can almost put your finger on. Nothing alerted, because from pfsync's point of view nothing was wrong: connection state was syncing fine. It was the configuration that had stopped.

In my case the cause was my own tooling. A Python change tool wrote rules through the API on the master. The GUI triggers the XMLRPC sync after a save; the API does not. Weeks of rules lived on fw01 only, and fw02 would have failed over to a policy from a month earlier.

What you need

Steps

  1. Confirm it is really stale before touching anything. Count the filter rules on each node through the API.
for fw in 10.1.2.252 10.1.2.253; do
  curl -s -k -u "$KEY:$SECRET" \
    "https://$fw/api/firewall/filter/searchRule" | jq '.rowCount'
done

Different numbers mean drift. Equal numbers do not mean no drift, which is why step 7 does a real diff.

  1. Check the sync settings on the master under System > High Availability > Settings. The backup host address must be the backup's sync-VLAN address (10.1.2.253), not a VLAN that might be busy or filtered. The remote username and password must belong to an account on the backup with administrator rights. Tick every service you want pushed: Firewall Rules, NAT, Aliases, Virtual IPs, Kea DHCP, Unbound, Users and Groups, and whichever plugins list themselves there. Leave the backup's own sync target empty. Sync is one-way; only the master pushes.

  2. Check that the credentials still work. The most common silent breakage is a rotated admin password on the backup. Look in System > Log Files > Backend on the master for lines mentioning xmlrpc and authentication. A sync that fails with a bad password logs one line and otherwise says nothing.

  3. Run the sync by hand.

configctl filter sync

The same thing is available in the GUI: System > High Availability > Status has a button that synchronises the configuration and reconfigures the backup's services. Use whichever you can script.

  1. If you have tooling that writes to the firewall, make it call the sync itself. The order that has held up for me is: apply to the backup first, read back and verify, apply to the master, then sync. Applying to the backup first means a bad change breaks the node that is not carrying traffic.

  2. Accept the list of things that never sync, and handle them by hand. CARP VIP deletes do not sync; the row stays on the backup until you delete it there too. ACME certificate renewals land on the node that ran the renewal. Interface assignments and VLAN devices are never pushed, so a rule that references an interface the backup does not have will be skipped. Plugins that do not list themselves in the sync settings are on their own.

  3. Put a drift check on a timer. The minimal version fetches the rules from both nodes, strips the per-node fields, and diffs.

import json, requests, sys

NODES = {"fw01": "10.1.2.252", "fw02": "10.1.2.253"}
AUTH = ("KEY", "SECRET")
IGNORE = {"uuid", "sequence"}

def rules(host):
    r = requests.get(f"https://{host}/api/firewall/filter/searchRule",
                     auth=AUTH, verify=False, params={"rowCount": -1})
    rows = r.json()["rows"]
    return sorted(json.dumps({k: v for k, v in row.items() if k not in IGNORE},
                             sort_keys=True) for row in rows)

a, b = rules(NODES["fw01"]), rules(NODES["fw02"])
only_a, only_b = set(a) - set(b), set(b) - set(a)
for line in only_a: print("master only:", line)
for line in only_b: print("backup only:", line)
sys.exit(1 if only_a or only_b else 0)

Run it every fifteen minutes from a cron host and send a non-zero exit to whatever push notification service you already use. Give it a quiet period after a known change so it does not alert on your own work.

sequenceDiagram
  participant T as Change tool
  participant B as fw02 backup
  participant A as fw01 master
  participant D as Drift check
  T->>B: write rule via API
  T->>B: read back
  T->>A: write the same rule
  T->>A: configctl filter sync
  A-->>B: XMLRPC config push
  D->>A: fetch rules
  D->>B: fetch rules
  D-->>T: diff, alert if different

Verify it worked

Run the rule count from step 1 again; the numbers should match. Run the diff from step 7; it should exit zero. On the backup, pfctl -sr | wc -l should be within a handful of lines of the master (the difference is per-node automatic rules). System > High Availability > Status on the master should list the backup's services without errors. Then do the test that matters: change a rule description in the GUI on the master, save, and watch it appear on the backup within a few seconds.

Gotchas

FAQ

Does the sync run on a schedule? No. It runs when a GUI save triggers it or when something calls configctl filter sync. There is no timer. If you want one, add a cron entry that calls the sync, and keep the drift check as the independent witness.

Will pfsync hide this from a failover test? Yes, for a while. State sync means established connections survive the failover, so a short test looks fine. New connections hit the backup's stale rule set. Test with a new connection that depends on a recent rule.

Can I diff the config.xml files instead of using the API? You can, but /conf/config.xml contains node-specific values (hostname, addresses, CARP skew, the sync settings themselves) that make a raw diff noisy. The API gives you the rules as the firewall sees them, which is what you are trying to compare.

Related