OPNsense: keep ACME certificates valid on both HA nodes
In short. Run the ACME client plugin on the master with a DNS-01 challenge so renewal does not depend on which node holds the WAN. The renewed certificate lands on the master only. A daily job copies the certificate and key to the backup and imports them, and an audit compares SHA-256 fingerprints on both nodes and alerts when they differ.
openssl x509 -fingerprintis the whole check.
Certificate expired on the backup firewall after failover
Everything was fine for months. The master renewed on schedule, the GUI and the reverse proxy served a fresh certificate, and nobody looked at the backup. Then a failover, and every browser in the estate complained about an expired certificate for .254. The backup had the certificate from the day it was built. The HA sync had not carried the renewal across, and nothing had been checking.
What you need
- Two OPNsense 26.1 nodes in a CARP pair, with the
os-acme-clientplugin installed on the master from System > Firmware > Plugins. - A DNS provider with an API the plugin supports, and an API token scoped to the zone, stored in the estate's secret manager. The token goes into the plugin's challenge type and nowhere else.
- SSH from the master to the backup over the sync VLAN (
10.1.2.252to10.1.2.253), with a key restricted to the copy job. opensslon both nodes, which FreeBSD ships.
Steps
-
Set up the account and challenge on the master. Services > ACME Client > Settings: enable the plugin. Accounts: add one with the contact address and the production directory. Challenge Types: add a DNS-01 type for your provider with the scoped token. DNS-01 rather than HTTP-01, because an HTTP challenge needs port 80 on the WAN VIP to reach the node running the client, which only holds while the master is master.
-
Issue the certificate. Services > ACME Client > Certificates: add the name (
fw.site1.example.net, with any alternative names), pick the account and the challenge, and issue it. The plugin stores the result under System > Trust > Certificates, where the GUI, the reverse proxy and anything else on the node can select it. The ACME working files live under/var/etc/acme-client/, which matters for step 4. -
Confirm that this does not reach the backup. Open System > Trust > Certificates on the backup. If your release lists certificates among the HA sync services and you have ticked it, a copy may be there, but in my experience the renewed one did not reliably replace it. Treat the backup as stale until the fingerprint check in step 5 says otherwise.
-
Add the daily copy. On the master, a small script exports the current certificate and key and pushes them to the backup, where an import script installs them. The export can be done through the trust API (
/api/trust/cert/) or by reading the plugin's files; the plugin's files are simpler and change only on renewal.
#!/bin/sh
# /usr/local/etc/cert-sync.sh on the master, run daily from cron
NAME="fw.site1.example.net"
SRC="/var/etc/acme-client/home/${NAME}"
DST="root@10.1.2.253:/var/db/cert-sync/"
scp -q "${SRC}/${NAME}.cer" "${SRC}/${NAME}.key" "${SRC}/fullchain.cer" "${DST}"
ssh -q root@10.1.2.253 /usr/local/etc/cert-import.sh "${NAME}"
The import on the backup is the part that needs care. OPNsense stores certificates in config.xml as base64, so the clean way is the trust API: look up the existing certificate's uuid by its description, then replace its certificate and key through the cert endpoints and reload the services that use it (the GUI and whichever proxy). The quick way is the GUI import under System > Trust > Certificates once, then a scripted update after that. Whichever you choose, keep the description identical on both nodes so the services keep pointing at the same object after an update rather than at a new one.
Install the master's script under System > Settings > Cron. The plugin's own cron entry runs renewal early in the morning; run the copy an hour later so it picks up the fresh file the same day.
- Add the fingerprint check. Run it from the same cron host that runs the drift audit, daily, and alert on any difference.
#!/bin/sh
# compare the certificate each node is actually serving
for fw in 10.1.2.252 10.1.2.253; do
printf '%s ' "$fw"
openssl s_client -connect "${fw}:443" -servername fw.site1.example.net </dev/null 2>/dev/null \
| openssl x509 -noout -fingerprint -sha256 -enddate
done
Two matching fingerprints and a sensible notAfter is a pass. Anything else pages someone. Checking what is served, not what is on disk, catches the case where the file was copied but the service was never reloaded.
sequenceDiagram
participant L as Let's Encrypt
participant A as fw01 master
participant B as fw02 backup
participant C as Fingerprint check
A->>L: renew with DNS-01
L-->>A: new certificate
A->>B: scp cert and key, run import
B->>B: replace cert object, reload services
C->>A: fingerprint of served cert
C->>B: fingerprint of served cert
C-->>C: equal and not near expiry, or alertVerify it worked
Run the fingerprint script by hand. Then run it from inside the estate against .254 while in CARP maintenance mode on the master, so the backup is answering: the fingerprint must still match. On each node, openssl x509 -in /var/etc/acme-client/home/fw.site1.example.net/fw.site1.example.net.cer -noout -enddate on the master and the equivalent path on the backup should agree. Last, open the GUI on the backup directly by its own address and look at the certificate the browser shows: same serial number as the master.
Gotchas
- DNS-01, not HTTP-01. An HTTP challenge depends on the WAN VIP being on the node that runs the client. It will work for a year and then fail during a maintenance window.
- The copy is not enough; the service has to reload. A certificate replaced in
config.xmlwhile the web server still holds the old one in memory passes a file check and fails the served check. - Keep the certificate's description identical on both nodes. Services reference certificates by uuid; an import that creates a new object leaves the proxy pointing at the old one.
- If the master is down for more than the renewal window, nothing renews. Put the expiry date in the audit output so a certificate with twenty days left is a warning before it is an outage.
- The DNS API token is a secret with the power to issue certificates for your zone. Scope it to the one zone, store it in the secret manager, and never paste it into a chat or a ticket.
FAQ
Why not run the ACME client on both nodes? Rate limits and two sources of truth. Two clients issue two certificates with different serials, which makes the fingerprint check meaningless and doubles the issuance against a shared limit. One client, one copy job, one audit.
Can the HA sync carry the certificate? Newer releases list certificates among the services the sync can push. Tick it if you have it, and keep the fingerprint check regardless, because the question is never "did the sync run" but "is the backup serving the right certificate".
What about the certificate the reverse proxy on the tenant ingress uses? That is a different certificate on a different machine, issued by its own ACME client, and it is not the firewall's problem. This post is about the firewall's own certificate for its GUI and anything the firewall itself terminates.