OPNsense: retire a public IP address without a blackhole
In short. Nothing is unused until you have listed every reference on both nodes, checked with the provider in writing, and had the replacement address running alongside for long enough to be sure. Only then delete the alias, on both nodes, and prove with a drift audit and external probes that everything still answers. The 26 hours this cost me are the reason the procedure has steps you will be tempted to skip.
I deleted a virtual IP on OPNsense and now a whole block is unreachable
Here is how it happened. A cleanup of old provider-assigned addresses found an IP alias on the WAN VIP that no port forward, no outbound NAT rule and no filter rule mentioned. It looked like leftover. It was deleted on the master, the backup followed the sync (or did not; it did not matter by then), and about twenty minutes later a tenant reported that nothing in their routed /28 answered.
The deleted address was the next hop the provider used to route the /28 to the estate. It did not appear in any rule because the firewall did not need a rule for it: the provider's router just needed it to respond to ARP. No check on the firewall could have seen that dependency. It took about 26 hours, most of them waiting for the provider, to understand and restore it.
What you need
- Two OPNsense 26.1 nodes in a CARP pair, with the WAN's public addresses as IP aliases on the single WAN CARP VIP.
- The address to retire. In the example it is
203.0.113.17, and the replacement already in service is203.0.113.33. - API access to both nodes for the reference search, and shell access for the final check.
- A contact at the provider who can answer in writing, and two external vantage points for probing.
- Patience measured in days, not minutes. The dual-run is deliberately slow.
Steps
- Inventory every reference on both nodes, including the ones you did not make. Search aliases, port forwards, outbound NAT, filter rules and the FRR configuration for the literal address. The search has to read the current config, not a tool's idea of it.
for fw in 10.1.2.252 10.1.2.253; do
echo "== $fw"
ssh "root@$fw" "grep -n '203\.0\.113\.17' /conf/config.xml; vtysh -c 'show running-config' | grep -n '203.0.113.17'"
done
An empty result is the start of the investigation, not the end of it. Also search DNS: the public zone, the internal overrides, and any monitoring configuration that targets the address.
-
Ask the provider, in writing, what they route to or via that address. The question is specific: "Do you use
203.0.113.17as a next hop, a monitoring target or anything else on your side?" Wait for the answer. If the provider assigned the address, assume it has a role on their side until they say otherwise. -
Put the replacement into service first and leave the old address in place. Every forward, outbound rule and DNS record that used
203.0.113.17is changed to203.0.113.33while the old address still exists. Lower the DNS TTL a day beforehand. Nothing is removed yet. This is the dual-run, and it runs for at least a week so that weekly jobs, monthly reports and the one client with a hard-coded address have a chance to show up in the logs. -
Watch the old address during the dual-run. Firewall > Log Files > Live View filtered on the address shows whoever is still using it.
pfctl -ss | grep 203.0.113.17shows live states. Zero hits for a week, plus the provider's written answer, is the bar. -
Run the reference guard one more time, then delete on both nodes. Firewall > Virtual IPs > Settings, delete the alias on the master and apply. CARP VIP deletes do not sync, so repeat on the backup. If a tool created the address, use the tool's retire command, which re-reads both configs and refuses if any rule still references the address, including hand-made ones.
-
Finish with the drift audit and external probes. The audit confirms the two nodes agree. The probes confirm the public is still being served.
flowchart TB
A["Inventory references<br/>both nodes, DNS, monitoring"] --> B["Ask the provider<br/>in writing"]
B --> C["Move everything to<br/>the replacement address"]
C --> D["Dual-run for a week<br/>watch logs and states"]
D --> E["Reference guard<br/>NAT, rules, BGP"]
E --> F["Delete alias on master<br/>then on backup"]
F --> G["Drift audit and<br/>probes from two vantage points"]Verify it worked
From two vantage points outside the estate, probe every published port on every remaining public address, not just the replacement. The script is dull and that is the point:
for addr in 203.0.113.10 203.0.113.33; do
for port in 443 25 993; do
nc -vz -w 5 "$addr" "$port" 2>&1 | tail -1
done
done
Then confirm the address is gone from both nodes (ifconfig | grep 203.0.113.17 returns nothing on either), that the drift audit exits zero, and that the provider's routed blocks still answer. Check DNS against the authoritative server, not your laptop. Leave the probe running on a timer for a few days; the failure you are guarding against might not show until the next failover.
Gotchas
- The absent rule is the trap. An address used as a next hop needs no firewall rule, so a rule search finds nothing and the address looks unused. Only the provider can tell you.
- Delete on both nodes. The sync carries VIP creates and edits but not deletes, so the backup keeps answering for the address until you remove it there too. Which sounds harmless until it is the backup's ARP reply that keeps a stale route alive on the provider's side.
- A tool that only knows about its own rules is not a reference guard. The retire command has to read every rule from the firewall, including the ones someone made in the GUI at two in the morning.
- Monitoring that pings the old address will go red when you remove it. Change monitoring during the dual-run, not after, or you will not know whether red means "retired" or "broken".
- Shortening the dual-run to save a week is how you spend 26 hours.
FAQ
Why not just add the address back if something breaks? Because you may not know what broke until someone tells you, and on a Friday the someone may be a tenant on Monday. Adding it back also does not restore the provider's state if their router has already aged out the route. Prevention is cheaper.
Does this apply to addresses in my own prefix, not provider-assigned ones? Yes, with one difference: your own prefix is announced as a whole, so the provider does not need an address inside it as a next hop. The inventory and dual-run steps are the same; the provider question becomes "do you monitor or route anything to this address" rather than "do you depend on it".
How does the tool decide an address is safe to retire? It fetches aliases, port forwards, outbound NAT and filter rules from both nodes through the API, reads the FRR configuration, and lists every object whose fields contain the address. If the list is not empty it stops. The human checks the provider and the dual-run; the tool checks the firewall.