Self-hosted NetBird as the operator plane: firewalls as routing peers, policies as transport

Isometric illustration of a laptop, two brick firewall walls and a small house joined by thin gold threads into a mesh, with a small control box and a dotted thread to the house.

In short. Management, signal and relay on one Docker VM behind a reverse proxy, SQLite shipped by Litestream to a home DR box, failover by DNS change. The firewall pair at each site is the routing peer, advertising one resource per VLAN with explicit SNAT into the overlay. Groups per human role, no policy sourced from "All", exit nodes per region and opt-in.

What it replaced

A jumphost per site, reached over SSH, with everything else reached through it. That works until two people need two sites at once, or a backup server needs to pull from a peer site, or a monitoring relay needs to phone home. A WireGuard mesh with a self-hosted controller gives every operator and every service a direct, encrypted path, and the jumphost becomes the break-glass instead of the front door.

Where it lives

Durability without a cluster

The management service keeps its state in SQLite. Litestream ships the write-ahead log to a small box at a home DR site roughly every 30 seconds. If the data centre VM is lost, the standby is started from the replica and the public DNS name is pointed at it.

That is not high availability, and it does not need to be. Peers that are already connected keep their tunnels when the control plane is away. What breaks is logins: operators are logged out after about 48 hours of control-plane outage, because their sessions expire and cannot be renewed. The warm standby exists to make sure the outage is shorter than that.

Routing peers on the firewall

flowchart TB
  subgraph CP["Control plane on the external ops VLAN"]
    MG["Management + Signal + Relay<br/>behind Traefik, SQLite"]
    LS["Litestream ships the WAL every 30 s"]
    MG --> LS
  end
  LS --> DR["Warm standby at a home DR site<br/>failover = public DNS change"]
  subgraph SA["Site A"]
    FWA["OPNsense pair = routing peer<br/>one resource per VLAN<br/>explicit SNAT into overlay"]
    PBSA["PBS A"]
  end
  subgraph SB["Site B"]
    FWB["OPNsense pair = routing peer"]
    PBSB["PBS B"]
  end
  OPS["Operator laptop<br/>role group, hardware key login"]
  OPS -- "WireGuard mesh 10.254/16" --> FWA
  OPS -- "WireGuard mesh" --> FWB
  PBSB -- "pull sync rides the mesh" --> PBSA
  MG -. "registration, groups, policies" .- OPS
  MG -.- FWA
  MG -.- FWB
  EX["Exit node per region<br/>opt-in, only where own space is announced"] -.- FWA

Rather than install an agent on every host, each site's firewall pair runs the agent and advertises the site's VLANs into the mesh.

Policies are transport, not authorisation

Mesh policies decide who can open a connection to what. They are not the access-control layer for applications, and they are not where "who can log in" lives. That belongs to the identity provider and to the packet filter on the firewall.

Exit nodes

An exit node lets a travelling laptop route its internet through a site. There is one per region, offered to a per-device travellers group, and it is opt-in on each device. Exit nodes exist only at sites that announce their own address space, so that traffic leaving through one comes from an address the organisation owns and can answer for.

Retiring the old public address

When the controller moved from provider-assigned to portable address space, retiring the old address was its own runbook:

  1. Delete the port forwards the tooling created.
  2. Find the hand-made legacy forwards the tooling cannot see, matching them loosely.
  3. Run a reference guard across outbound NAT, inbound NAT, filter rules and BGP before touching the VIP.
  4. Delete the VIP on both firewalls explicitly, because VIP deletes do not sync.
  5. Finish with a drift audit of the pair.

Gotchas

FAQ

Why not the hosted version? The controller holds the map of every site and every peer. That belongs on infrastructure the organisation controls, and the self-hosted version costs a small VM.

Why not an agent on every host? Agents on hosts mean a tunnel endpoint on every host, a credential on every host, and a dependency on the control plane for every host. Routing peers on the firewall keep the blast radius at two devices per site.

What runs over the mesh besides people? Backup replication between sites, monitoring proxies at secondary sites, and the secrets store's disaster-recovery voters.

Related