Self-hosted NetBird end to end: zero-trust admin access and an encrypted mesh between sites

In short. Run the combined NetBird server, the dashboard and Traefik on one VM, and ship its SQLite databases to a warm standby with Litestream. Enrol one admin laptop. Make each site's OPNsense pair a routing peer that advertises one resource per VLAN, with an outbound NAT rule for traffic entering the overlay. Join the sites with one self-policy over the ops networks, then give people access by role, with no policy sourced from "All". Every policy only opens a tunnel; the firewall's packet filter and the service itself decide what is allowed.

One walkthrough instead of five

NetBird's own documentation and OPNsense's covers each part well: installing the server, the OPNsense plugin, site-to-site routing, site-to-VPN access. I did not find a single walkthrough that joins them in the order you actually build them, for an estate where the same mesh carries both people and machines. This is that order. The design reasons are in Self-hosted NetBird as the operator plane, and the firewall detail is in OPNsense as the routing peer; this post links to both rather than repeating them.

flowchart TB
  CP["Control plane VM<br/>Traefik, dashboard,<br/>netbird-server, SQLite"] -- "Litestream,<br/>about 30 s" --> DR["Warm standby<br/>second relay"]
  CP -. "login, peers,<br/>policies" .-> OPS["Operator laptop<br/>role group, passkey"]
  OPS -- "WireGuard<br/>10.254.0.0/16" --> FWA["Site A OPNsense pair<br/>routing peers,<br/>one resource per VLAN"]
  FWA -- "pf decides" --> VA["Site A VLANs"]
  FWA -- "ops mesh,<br/>OPS-INT only" --- FWB["Site B OPNsense pair<br/>routing peers"]
  FWB -- "pf decides" --> VB["Site B VLANs"]

What you need

Steps

1. Plan the overlay and the groups before the first peer

Write the groups down first, because every later step refers to them:

GroupHoldsUsed for
FW-GRP-<site>both firewalls of one siterouting peers for that site's resources
SITE-OPSevery firewall, plus the ops-internal resourcesthe site-to-site mesh
PBS-NODES and similarservice peers such as backup serversone machine flow each
PLATFORM-ADMINS and two narrower rolespeople's devicesadmin access by role
TRAVELLERSopted-in laptopsexit nodes, if you want them

Two rules hold from the start: no policy is ever sourced from "All", so a new peer reaches nothing until someone puts it in a group, and one rule per policy, because the API keeps only the first rule of a multi-rule policy.

2. Build the control plane on one VM

Three containers: Traefik for TLS, the dashboard, and the combined netbird-server image, which runs management, signal, relay and STUN in one process. Pin the versions in one .env beside the compose file:

# /opt/netbird/.env
NETBIRD_VERSION=0.79.0            # server and relay: bare tag, no "v"
NETBIRD_DASHBOARD_VERSION=v2.94.0 # dashboard: tag starts with "v"
TRAEFIK_VERSION=v3.6

The two tag styles are not a typo, and getting them wrong is not loud. A "v" on the server tag fails at pull time, and on a standby that only pulls nightly nobody sees it fail.

The parts of the compose file that matter:

services:
  traefik:
    image: traefik:${TRAEFIK_VERSION}
    networks: { netbird: { ipv4_address: 172.30.0.10 } } # fixed, so it can be the one trusted proxy
    ports: ["443:443", "80:80"]
    command:
      - --providers.docker.exposedbydefault=false
      - --entrypoints.websecure.address=:443
      - --entrypoints.websecure.transport.respondingTimeouts.readTimeout=0 # long-lived gRPC streams
      - --certificatesresolvers.letsencrypt.acme.tlschallenge=true

  dashboard:
    image: netbirdio/dashboard:${NETBIRD_DASHBOARD_VERSION}
    env_file: [./dashboard.env] # OIDC settings; the client secret comes from the secrets store

  netbird-server:
    image: netbirdio/netbird-server:${NETBIRD_VERSION}
    ports: ["3478:3478/udp"] # STUN; everything else goes through Traefik
    volumes:
      - netbird_data:/var/lib/netbird
      - ./config.yaml:/etc/netbird/config.yaml
    command: ["--config", "/etc/netbird/config.yaml"]
    labels:
      # gRPC for signal and management, over h2c to the container
      - traefik.http.routers.nb-grpc.rule=Host(`nb.example.net`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`))
      - traefik.http.services.nb-h2c.loadbalancer.server.scheme=h2c
      # relay, websocket proxy, REST API and the embedded identity provider, over HTTP
      - traefik.http.routers.nb-http.rule=Host(`nb.example.net`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))

And the server's config.yaml, without its secrets:

server:
  listenAddress: ":80"
  exposedAddress: "https://nb.example.net:443"
  stunPorts: [3478]
  dataDir: "/var/lib/netbird"
  auth:
    issuer: "https://nb.example.net/oauth2"
    dashboardRedirectURIs: ["https://nb.example.net/nb-auth", "https://nb.example.net/nb-silent-auth"]
    cliRedirectURIs: ["http://localhost:53000/"]
  reverseProxy:
    trustedHTTPProxies: ["172.30.0.10/32"] # Traefik only
  store:
    engine: "sqlite"
  relay:
    addresses: ["rels://nb.example.net:443", "rels://relay.example.net:443"] # the second relay is the standby

Login uses the embedded identity provider with an upstream OIDC provider and passkeys. Moving it onto the estate's own identity provider is built but not switched over; a small control plane is a reasonable place to defer that.

Open UDP 3478 inbound on the firewall early. Until it was open here, peers that could not find a direct path fell back to the relay and nobody noticed for a while.

3. Ship the database to a warm standby

The data directory holds three SQLite files: store.db (peers, groups, policies: the one that matters), idp.db (the embedded identity provider) and events.db (the audit log). Litestream runs as a systemd service on the primary and ships all three over SFTP:

# /etc/litestream.yml on the control-plane VM
dbs:
  - path: /var/lib/docker/volumes/netbird_data/_data/store.db
    replicas:
      - type: sftp
        host: standby.example.net:22
        user: litestream
        path: /opt/netbird/data/store.db
  # the same block again for idp.db and events.db

With no interval set, Litestream's defaults apply, and the replica trails the primary by about 30 seconds. On the standby:

Read the roll-back section below before relying on that last bullet.

4. Enrol the first peer: your own laptop

Open https://nb.example.net, log in, and enrol your laptop from the dashboard or with netbird up --management-url https://nb.example.net. Put the laptop in PLATFORM-ADMINS. It can reach nothing yet, which is correct: no policy exists.

Set login expiry for people's devices now. Here it is two days, so a lost laptop stops working on its own. Infrastructure peers, which have no person to log in again, are exempt.

5. Make each firewall pair a routing peer

The routing peer how-to has every click. In order:

  1. Install os-netbird on both firewalls and register them with a setup key that drops them into FW-GRP-<site>. Expire the key once both are in.
  2. Assign the wt0 interface on both nodes as NETBIRD, with no address. HA sync does not assign interfaces or install plugins.
  3. Create one NetBird network per site and one resource per VLAN, never a site /16, with FW-GRP-<site> as the routers.
  4. Add one outbound NAT rule per firewall: source the site's 10.<site>.0.0/16, destination the overlay, translated to the wt0 address. The agent's masquerade only covers traffic leaving the overlay; without this rule, anything that starts on the VLAN side leaves with an address the far peer drops.
  5. Start the NETBIRD interface at default deny and add pass rules one flow at a time.

Check the plugin's NetBird version against the server's and against advisories. Here the plugin lagged several releases behind, past a local privilege escalation fixed upstream; until the plugin caught up, tightening the permissions on the agent's socket (chmod 660 /var/run/netbird.sock) closed it.

6. Join the sites

Put every firewall and every ops-internal resource in SITE-OPS, and give that group one bidirectional policy to itself. Management networks and tenant networks are deliberately not in it.

On each firewall, two pass rules carry it, both marked quick:

Do not replace them with one pass for 10.0.0.0/8. It reads tidier and quietly bypasses the block from ops to management.

Machine flows then get their own group and their own narrow policy, each one port-scoped:

The PBS replication how-to is one of those flows from start to finish.

7. Give people access by role

Three human roles cover this estate:

RoleReaches
PLATFORM-ADMINSthe ops and management resources at every site, and NetBird SSH
a single-user lab roleone person's own lab host and nothing else
RESTRICTED-USERSone tenant /24 and DNS

Plus one DNS policy from the role groups to the HQ firewalls on UDP 53, and app groups for single services (TCP 443 to one /32).

A policy here is transport. It lets a laptop build a tunnel to a routing peer. Whether the packet then reaches a host is the packet filter's decision on the NETBIRD interface, and whether the person can log in is the service's (SSH certificates, the secrets store, the application). If removing a NetBird policy is the only thing that would stop someone reaching the storage network, there is one control where there should be two.

Keep the definitions in a script that is dry-run by default and reads back after every write, and run a scheduled audit for the shapes you never want: a policy sourced from "All", a resource wider than a VLAN, a peer whose endpoint sits inside a prefix it advertises.

8. Exit nodes, if you want them

One exit node per country, both firewalls of a site with different metrics, opt-in on each device and offered only to TRAVELLERS. A one-way ICMP policy to the exit-node group is enough for the route to be delivered. Do not add an all-domains nameserver: captive portals in hotels and airports need the local resolver to work.

Verify it worked

  1. From the laptop, netbird status -d lists each routing peer and says whether the path is direct or relayed, and from which endpoint.

  2. On a firewall, routes into the overlay are present. OPNsense's root shell is csh, so wrap it:

    sh -c 'netstat -rn -f inet | grep wt0'
    
  3. Test reachability from a host that does not run a NetBird client, such as an out-of-band management controller on the ops network, and never to a firewall's own address. Both shortcuts test a different path from the one you built.

  4. When something does not pass, capture on wt0 of the sending peer, and work through DNS, then status -d, then the policy, then the packet filter, in that order.

  5. Stop the agent on one firewall of a pair and check the route moves to the other. Start it again.

Roll back

Each layer comes off on its own:

Be honest with yourself about that last one: the promotion script here has not been run for real yet. We did a drill to check the standby, and it uncovered a bad image tag. The server tag had been bumped with a "v", that tag did not exist, and the nightly pull had been failing quietly, so the standby could not have been promoted. The nightly job now checks the images and the restored databases instead of assuming them (step 3), upgrades go to the standby first, and the tooling refuses to upgrade the primary until the standby is on the target version. Run a real promotion on a quiet day before you need one.

Gotchas

FAQ

Why not the hosted service? The control plane holds the map of every site and every peer. That belongs on infrastructure the organisation controls, and a small VM costs less than a seat per person.

Why not an agent on every server? That is a tunnel endpoint, a credential and a control-plane dependency on every server. Routing peers on the firewalls keep it to two devices per site, with the packet filter already in the path. Service peers such as backup servers are the exception, and each has a single narrow policy.

What happens if the control plane disappears for a day? Tunnels already up stay up, so backups, replication and monitoring carry on. Nobody can enrol a device or change a policy, and people start being logged out at the two-day mark. The standby exists to make the outage shorter than that.

Is this zero trust? In the useful sense: nothing is reachable because of where it sits on the network, every flow is a named group to a named resource, and every service still authenticates the person or machine on the other end. The mesh just carries the packets.

Related