OPNsense: split-horizon DNS with Unbound forward zones

In short. Unbound on the firewall forwards each internal zone to the site's PowerDNS and resolves everything else over TLS to an upstream. Public names that internal clients must reach without hairpinning get a host override with the internal address. Test with dig against the firewall for the internal view and against the authoritative server for the public one, because a laptop inside the estate can only ever see the internal answer.

Internal clients get the public IP for my own service

Or the reverse: they get the internal address when you wanted to test the public path. Both are the same problem seen from different sides. Unbound on the firewall is the resolver for every VLAN, so what it answers is what the estate believes. This post sets it up so that internal zones come from the source of record, public names resolve internally where that is wanted, and you know which answer you are looking at.

What you need

Steps

  1. Forward each internal zone to the site's authoritative server. Services > Unbound DNS > Query Forwarding, add: domain site1.internal.example, server IP 10.1.10.53, port 53. Leave "Forward TLS upstream" unticked for an internal server on a VLAN you control. Add the reverse zones the same way, for example 1.10.in-addr.arpa, so that PTR lookups also come from the source of record. Repeat for each site's zone if sites must resolve each other.

  2. Forward everything else over TLS. Services > Unbound DNS > DNS over TLS, add the upstream's address, port 853 and its verify name. Add a second upstream from a different operator. Then, under General, make sure Unbound is not also using the WAN interface's DHCP-assigned resolvers in parallel (the system's own DNS servers under System > Settings > General are for the firewall itself; untick the option that lets Unbound use them if you want TLS to be the only path).

  3. Add host overrides for public names that must resolve internally. Services > Unbound DNS > Overrides, Host Overrides, add: host portal, domain example.net (the public zone), type A, address 10.1.101.10. This is the tenant ingress at .10. An internal client asking for portal.example.net now gets 10.1.101.10 and goes straight to the VM across the VLANs. The port forward on the WAN VIP still exists for the public, and NAT reflection stays off.

  4. Make sure every VLAN's clients use the firewall. The Kea DHCP subnet for a tenant hands out 10.<site>.<vlan>.254 or the tenant's own .2 resolver as the DNS server; the tenant resolver in turn forwards to the firewall. Either way the firewall's view is the estate's view. Hosts with a hard-coded public resolver bypass all of this, which the tenant firewall rules should block (only the resolver and the firewall may talk to port 53 and 853 outside).

  5. Apply, and do it on both nodes. Unbound settings are in the HA sync list under System > High Availability > Settings, so a GUI save on the master pushes them. Confirm on the backup anyway; it has to answer the same way after failover, and a client whose resolver is .254 will be talking to it.

flowchart LR
  C["Tenant client<br/>10.1.101.50"] -->|"portal.example.net?"| U["Unbound on fw<br/>10.1.101.254"]
  U -->|"host override"| C
  U -->|"site1.internal.example<br/>forward zone"| P["PowerDNS<br/>10.1.10.53"]
  U -->|"everything else<br/>DNS over TLS"| X["Upstream resolver<br/>port 853"]
  I["Internet client"] -->|"portal.example.net?"| A["Public authoritative<br/>answers 203.0.113.10"]

Verify it worked

Three queries, three different answers, and you need all three to be right.

# 1. The internal view: ask the firewall from inside a tenant VLAN
dig +short portal.example.net @10.1.101.254
# expect 10.1.101.10

# 2. The internal zone: ask the firewall, then ask the source of record directly
dig +short host01.site1.internal.example @10.1.101.254
dig +short host01.site1.internal.example @10.1.10.53
# expect the same address from both

# 3. The public view: ask the public authoritative server, over TCP, not your laptop's resolver
dig +tcp +short portal.example.net @ns1.example.net
# expect 203.0.113.10

Query 3 is the one people skip. Running dig portal.example.net on a laptop inside the estate asks the firewall and gets 10.1.101.10, which proves the override and nothing else. Running it on a laptop at home asks the home resolver, which may be caching an answer from last week. Asking the authoritative server over TCP removes both the cache and the resolver from the picture. If you cannot reach the authoritative server, a DNS over HTTPS query to a public resolver with a fresh cache is the next best thing. For a reverse lookup, dig -x 10.1.101.10 @10.1.101.254 should return the name from PowerDNS.

Last, check the TLS path is actually in use: Services > Unbound DNS > Log File, or unbound-control stats_noreset | grep -i tls from the shell, should show forwarded queries going to port 853 and none to port 53 upstream.

Gotchas

FAQ

Why forward to PowerDNS instead of putting the records in Unbound? Because the records come from NetBox. The source of record generates the zone; Unbound just has to know where to ask. Putting host entries in the firewall GUI is a second copy that drifts.

Could the tenant's own .2 resolver do the split horizon? It could, and for tenant-specific names it does. The firewall's override is for estate-wide names that every VLAN should resolve the same way, and it is the fallback when a tenant resolver is down.

Why TCP for the authoritative query? Because some middleboxes and resolvers treat UDP DNS specially, and a TCP answer from the authoritative server over port 53 is hard to misattribute. It also sidesteps truncation on large answers.

Related