Self-hosted NetBird as the operator plane: firewalls as routing peers, policies as transport

In short. Management, signal and relay on one Docker VM behind a reverse proxy, SQLite shipped by Litestream to a home DR box, failover by DNS change. The firewall pair at each site is the routing peer, advertising one resource per VLAN with explicit SNAT into the overlay. Groups per human role, no policy sourced from "All", exit nodes per region and opt-in.
What it replaced
A jumphost per site, reached over SSH, with everything else reached through it. That works until two people need two sites at once, or a backup server needs to pull from a peer site, or a monitoring relay needs to phone home. A WireGuard mesh with a self-hosted controller gives every operator and every service a direct, encrypted path, and the jumphost becomes the break-glass instead of the front door.
Where it lives
- One VM on the external ops VLAN, running management, signal and relay containers behind a reverse proxy.
- Published on 443 and 80 for the control plane and 3478 TCP and UDP for STUN and TURN. The relay carries traffic for peers that cannot reach each other directly.
- The overlay is
10.254.0.0/16, reserved in the addressing scheme and never assigned to a VLAN. - The management API key lives in the secrets store. Scripts read it at run time.
Durability without a cluster
The management service keeps its state in SQLite. Litestream ships the write-ahead log to a small box at a home DR site roughly every 30 seconds. If the data centre VM is lost, the standby is started from the replica and the public DNS name is pointed at it.
That is not high availability, and it does not need to be. Peers that are already connected keep their tunnels when the control plane is away. What breaks is logins: operators are logged out after about 48 hours of control-plane outage, because their sessions expire and cannot be renewed. The warm standby exists to make sure the outage is shorter than that.
Routing peers on the firewall
flowchart TB
subgraph CP["Control plane on the external ops VLAN"]
MG["Management + Signal + Relay<br/>behind Traefik, SQLite"]
LS["Litestream ships the WAL every 30 s"]
MG --> LS
end
LS --> DR["Warm standby at a home DR site<br/>failover = public DNS change"]
subgraph SA["Site A"]
FWA["OPNsense pair = routing peer<br/>one resource per VLAN<br/>explicit SNAT into overlay"]
PBSA["PBS A"]
end
subgraph SB["Site B"]
FWB["OPNsense pair = routing peer"]
PBSB["PBS B"]
end
OPS["Operator laptop<br/>role group, hardware key login"]
OPS -- "WireGuard mesh 10.254/16" --> FWA
OPS -- "WireGuard mesh" --> FWB
PBSB -- "pull sync rides the mesh" --> PBSA
MG -. "registration, groups, policies" .- OPS
MG -.- FWA
MG -.- FWB
EX["Exit node per region<br/>opt-in, only where own space is announced"] -.- FWARather than install an agent on every host, each site's firewall pair runs the agent and advertises the site's VLANs into the mesh.
- One resource per VLAN, not a site-wide /16. Access to the storage VLAN and access to a tenant VLAN are different grants.
- Explicit SNAT into the overlay. The agent's masquerade option covers traffic leaving the overlay into the site. The return direction, from the site into the overlay, needs its own NAT rule on the firewall, or replies go nowhere.
- No peer may use a tunnel endpoint inside a prefix it advertises. A firewall that advertises a VLAN and also has its own tunnel endpoint in that VLAN will encapsulate its own tunnel traffic. Large packets are dropped and small ones work, which is a miserable thing to debug.
Policies are transport, not authorisation
Mesh policies decide who can open a connection to what. They are not the access-control layer for applications, and they are not where "who can log in" lives. That belongs to the identity provider and to the packet filter on the firewall.
- Groups are per human role. There are three.
- No policy has "All" as its source. A new peer can reach nothing until it is placed in a group.
- Service peers (backup servers, monitoring relays) have their own groups and can reach exactly the peer they need.
Exit nodes
An exit node lets a travelling laptop route its internet through a site. There is one per region, offered to a per-device travellers group, and it is opt-in on each device. Exit nodes exist only at sites that announce their own address space, so that traffic leaving through one comes from an address the organisation owns and can answer for.
Retiring the old public address
When the controller moved from provider-assigned to portable address space, retiring the old address was its own runbook:
- Delete the port forwards the tooling created.
- Find the hand-made legacy forwards the tooling cannot see, matching them loosely.
- Run a reference guard across outbound NAT, inbound NAT, filter rules and BGP before touching the VIP.
- Delete the VIP on both firewalls explicitly, because VIP deletes do not sync.
- Finish with a drift audit of the pair.
Gotchas
- The API silently drops rule fields it does not understand, and a partial PUT wipes the rules it does. Read back after every write.
- Refer to peers by name. Overlay addresses change when the overlay is renumbered; names do not.
- DNS inside the mesh points at the site resolvers, so a "public" URL tested from a mesh-connected laptop tests the internal path.
- Pin image versions on the standby and pull nightly, or the failover brings up a version that does not match the replica's schema.
FAQ
Why not the hosted version? The controller holds the map of every site and every peer. That belongs on infrastructure the organisation controls, and the self-hosted version costs a small VM.
Why not an agent on every host? Agents on hosts mean a tunnel endpoint on every host, a credential on every host, and a dependency on the control plane for every host. Routing peers on the firewall keep the blast radius at two devices per site.
What runs over the mesh besides people? Backup replication between sites, monitoring proxies at secondary sites, and the secrets store's disaster-recovery voters.