OPNsense: retire a public IP address without a blackhole

In short. Nothing is unused until you have listed every reference on both nodes, checked with the provider in writing, and had the replacement address running alongside for long enough to be sure. Only then delete the alias, on both nodes, and prove with a drift audit and external probes that everything still answers. The 26 hours this cost me are the reason the procedure has steps you will be tempted to skip.

I deleted a virtual IP on OPNsense and now a whole block is unreachable

Here is how it happened. A cleanup of old provider-assigned addresses found an IP alias on the WAN VIP that no port forward, no outbound NAT rule and no filter rule mentioned. It looked like leftover. It was deleted on the master, the backup followed the sync (or did not; it did not matter by then), and about twenty minutes later a tenant reported that nothing in their routed /28 answered.

The deleted address was the next hop the provider used to route the /28 to the estate. It did not appear in any rule because the firewall did not need a rule for it: the provider's router just needed it to respond to ARP. No check on the firewall could have seen that dependency. It took about 26 hours, most of them waiting for the provider, to understand and restore it.

What you need

Steps

  1. Inventory every reference on both nodes, including the ones you did not make. Search aliases, port forwards, outbound NAT, filter rules and the FRR configuration for the literal address. The search has to read the current config, not a tool's idea of it.
for fw in 10.1.2.252 10.1.2.253; do
  echo "== $fw"
  ssh "root@$fw" "grep -n '203\.0\.113\.17' /conf/config.xml; vtysh -c 'show running-config' | grep -n '203.0.113.17'"
done

An empty result is the start of the investigation, not the end of it. Also search DNS: the public zone, the internal overrides, and any monitoring configuration that targets the address.

  1. Ask the provider, in writing, what they route to or via that address. The question is specific: "Do you use 203.0.113.17 as a next hop, a monitoring target or anything else on your side?" Wait for the answer. If the provider assigned the address, assume it has a role on their side until they say otherwise.

  2. Put the replacement into service first and leave the old address in place. Every forward, outbound rule and DNS record that used 203.0.113.17 is changed to 203.0.113.33 while the old address still exists. Lower the DNS TTL a day beforehand. Nothing is removed yet. This is the dual-run, and it runs for at least a week so that weekly jobs, monthly reports and the one client with a hard-coded address have a chance to show up in the logs.

  3. Watch the old address during the dual-run. Firewall > Log Files > Live View filtered on the address shows whoever is still using it. pfctl -ss | grep 203.0.113.17 shows live states. Zero hits for a week, plus the provider's written answer, is the bar.

  4. Run the reference guard one more time, then delete on both nodes. Firewall > Virtual IPs > Settings, delete the alias on the master and apply. CARP VIP deletes do not sync, so repeat on the backup. If a tool created the address, use the tool's retire command, which re-reads both configs and refuses if any rule still references the address, including hand-made ones.

  5. Finish with the drift audit and external probes. The audit confirms the two nodes agree. The probes confirm the public is still being served.

flowchart TB
  A["Inventory references<br/>both nodes, DNS, monitoring"] --> B["Ask the provider<br/>in writing"]
  B --> C["Move everything to<br/>the replacement address"]
  C --> D["Dual-run for a week<br/>watch logs and states"]
  D --> E["Reference guard<br/>NAT, rules, BGP"]
  E --> F["Delete alias on master<br/>then on backup"]
  F --> G["Drift audit and<br/>probes from two vantage points"]

Verify it worked

From two vantage points outside the estate, probe every published port on every remaining public address, not just the replacement. The script is dull and that is the point:

for addr in 203.0.113.10 203.0.113.33; do
  for port in 443 25 993; do
    nc -vz -w 5 "$addr" "$port" 2>&1 | tail -1
  done
done

Then confirm the address is gone from both nodes (ifconfig | grep 203.0.113.17 returns nothing on either), that the drift audit exits zero, and that the provider's routed blocks still answer. Check DNS against the authoritative server, not your laptop. Leave the probe running on a timer for a few days; the failure you are guarding against might not show until the next failover.

Gotchas

FAQ

Why not just add the address back if something breaks? Because you may not know what broke until someone tells you, and on a Friday the someone may be a tenant on Monday. Adding it back also does not restore the provider's state if their router has already aged out the route. Prevention is cheaper.

Does this apply to addresses in my own prefix, not provider-assigned ones? Yes, with one difference: your own prefix is announced as a whole, so the provider does not need an address inside it as a next hop. The inventory and dual-run steps are the same; the provider question becomes "do you monitor or route anything to this address" rather than "do you depend on it".

How does the tool decide an address is safe to retire? It fetches aliases, port forwards, outbound NAT and filter rules from both nodes through the API, reads the FRR configuration, and lists every object whose fields contain the address. If the list is not empty it stops. The human checks the provider and the dual-run; the tool checks the firewall.

Related