Self-hosted NetBird end to end: zero-trust admin access and an encrypted mesh between sites
In short. Run the combined NetBird server, the dashboard and Traefik on one VM, and ship its SQLite databases to a warm standby with Litestream. Enrol one admin laptop. Make each site's OPNsense pair a routing peer that advertises one resource per VLAN, with an outbound NAT rule for traffic entering the overlay. Join the sites with one self-policy over the ops networks, then give people access by role, with no policy sourced from "All". Every policy only opens a tunnel; the firewall's packet filter and the service itself decide what is allowed.
One walkthrough instead of five
NetBird's own documentation and OPNsense's covers each part well: installing the server, the OPNsense plugin, site-to-site routing, site-to-VPN access. I did not find a single walkthrough that joins them in the order you actually build them, for an estate where the same mesh carries both people and machines. This is that order. The design reasons are in Self-hosted NetBird as the operator plane, and the firewall detail is in OPNsense as the routing peer; this post links to both rather than repeating them.
flowchart TB
CP["Control plane VM<br/>Traefik, dashboard,<br/>netbird-server, SQLite"] -- "Litestream,<br/>about 30 s" --> DR["Warm standby<br/>second relay"]
CP -. "login, peers,<br/>policies" .-> OPS["Operator laptop<br/>role group, passkey"]
OPS -- "WireGuard<br/>10.254.0.0/16" --> FWA["Site A OPNsense pair<br/>routing peers,<br/>one resource per VLAN"]
FWA -- "pf decides" --> VA["Site A VLANs"]
FWA -- "ops mesh,<br/>OPS-INT only" --- FWB["Site B OPNsense pair<br/>routing peers"]
FWB -- "pf decides" --> VB["Site B VLANs"]What you need
- A small Linux VM for the control plane (Ubuntu 24.04 here) on an ops VLAN, published through the firewall on TCP 80 and 443 and UDP 3478, with a public name such as
nb.example.net. - A second small host somewhere else for the standby. Here it is a mini PC in a DMZ at a home site, which is cheap and far enough away.
- An OPNsense CARP pair at each site, on a release that has the
os-netbirdplugin. - An address range reserved for the overlay and never used anywhere else. This estate uses
10.254.0.0/16. - A secrets store for the relay secret, the store encryption key, the dashboard's OIDC client secret, setup keys and the management API token. None of them goes in a repository.
- Somewhere to keep the policy definitions as code. The dashboard is fine for looking; scripts that read back what they wrote are better for changing.
Steps
1. Plan the overlay and the groups before the first peer
Write the groups down first, because every later step refers to them:
| Group | Holds | Used for |
|---|---|---|
FW-GRP-<site> | both firewalls of one site | routing peers for that site's resources |
SITE-OPS | every firewall, plus the ops-internal resources | the site-to-site mesh |
PBS-NODES and similar | service peers such as backup servers | one machine flow each |
PLATFORM-ADMINS and two narrower roles | people's devices | admin access by role |
TRAVELLERS | opted-in laptops | exit nodes, if you want them |
Two rules hold from the start: no policy is ever sourced from "All", so a new peer reaches nothing until someone puts it in a group, and one rule per policy, because the API keeps only the first rule of a multi-rule policy.
2. Build the control plane on one VM
Three containers: Traefik for TLS, the dashboard, and the combined netbird-server image, which runs management, signal, relay and STUN in one process. Pin the versions in one .env beside the compose file:
# /opt/netbird/.env
NETBIRD_VERSION=0.79.0 # server and relay: bare tag, no "v"
NETBIRD_DASHBOARD_VERSION=v2.94.0 # dashboard: tag starts with "v"
TRAEFIK_VERSION=v3.6
The two tag styles are not a typo, and getting them wrong is not loud. A "v" on the server tag fails at pull time, and on a standby that only pulls nightly nobody sees it fail.
The parts of the compose file that matter:
services:
traefik:
image: traefik:${TRAEFIK_VERSION}
networks: { netbird: { ipv4_address: 172.30.0.10 } } # fixed, so it can be the one trusted proxy
ports: ["443:443", "80:80"]
command:
- --providers.docker.exposedbydefault=false
- --entrypoints.websecure.address=:443
- --entrypoints.websecure.transport.respondingTimeouts.readTimeout=0 # long-lived gRPC streams
- --certificatesresolvers.letsencrypt.acme.tlschallenge=true
dashboard:
image: netbirdio/dashboard:${NETBIRD_DASHBOARD_VERSION}
env_file: [./dashboard.env] # OIDC settings; the client secret comes from the secrets store
netbird-server:
image: netbirdio/netbird-server:${NETBIRD_VERSION}
ports: ["3478:3478/udp"] # STUN; everything else goes through Traefik
volumes:
- netbird_data:/var/lib/netbird
- ./config.yaml:/etc/netbird/config.yaml
command: ["--config", "/etc/netbird/config.yaml"]
labels:
# gRPC for signal and management, over h2c to the container
- traefik.http.routers.nb-grpc.rule=Host(`nb.example.net`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`))
- traefik.http.services.nb-h2c.loadbalancer.server.scheme=h2c
# relay, websocket proxy, REST API and the embedded identity provider, over HTTP
- traefik.http.routers.nb-http.rule=Host(`nb.example.net`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))
And the server's config.yaml, without its secrets:
server:
listenAddress: ":80"
exposedAddress: "https://nb.example.net:443"
stunPorts: [3478]
dataDir: "/var/lib/netbird"
auth:
issuer: "https://nb.example.net/oauth2"
dashboardRedirectURIs: ["https://nb.example.net/nb-auth", "https://nb.example.net/nb-silent-auth"]
cliRedirectURIs: ["http://localhost:53000/"]
reverseProxy:
trustedHTTPProxies: ["172.30.0.10/32"] # Traefik only
store:
engine: "sqlite"
relay:
addresses: ["rels://nb.example.net:443", "rels://relay.example.net:443"] # the second relay is the standby
Login uses the embedded identity provider with an upstream OIDC provider and passkeys. Moving it onto the estate's own identity provider is built but not switched over; a small control plane is a reasonable place to defer that.
Open UDP 3478 inbound on the firewall early. Until it was open here, peers that could not find a direct path fell back to the relay and nobody noticed for a while.
3. Ship the database to a warm standby
The data directory holds three SQLite files: store.db (peers, groups, policies: the one that matters), idp.db (the embedded identity provider) and events.db (the audit log). Litestream runs as a systemd service on the primary and ships all three over SFTP:
# /etc/litestream.yml on the control-plane VM
dbs:
- path: /var/lib/docker/volumes/netbird_data/_data/store.db
replicas:
- type: sftp
host: standby.example.net:22
user: litestream
path: /opt/netbird/data/store.db
# the same block again for idp.db and events.db
With no interval set, Litestream's defaults apply, and the replica trails the primary by about 30 seconds. On the standby:
- The relay and Traefik always run, so the standby is the second relay every peer already knows about.
- The dashboard and server sit behind a compose profile, so a plain
docker compose pullorstartskips them. Everything that touches them names the profile on purpose. - A nightly job proves the standby is promotable. It restores each database from the replica into a scratch directory with
litestream restoreand runs SQLite's integrity check on it, pulls the pinned images and confirms every one is actually present, and checks its own files against what was last deployed. If any of that fails, it stops without touching the running standby; otherwise it swaps in the fresh databases and restarts the server. Send its exit status to whatever pages you: a check that only writes to a log file catches nothing. - Promotion is a script: stop the nightly job, stop Litestream on the primary if it still answers, restore the three files, swap the standby's hostnames for the primary's in the config and compose files so the identity provider's callbacks match, start the full stack, then point the public name at the standby.
Read the roll-back section below before relying on that last bullet.
4. Enrol the first peer: your own laptop
Open https://nb.example.net, log in, and enrol your laptop from the dashboard or with netbird up --management-url https://nb.example.net. Put the laptop in PLATFORM-ADMINS. It can reach nothing yet, which is correct: no policy exists.
Set login expiry for people's devices now. Here it is two days, so a lost laptop stops working on its own. Infrastructure peers, which have no person to log in again, are exempt.
5. Make each firewall pair a routing peer
The routing peer how-to has every click. In order:
- Install
os-netbirdon both firewalls and register them with a setup key that drops them intoFW-GRP-<site>. Expire the key once both are in. - Assign the
wt0interface on both nodes asNETBIRD, with no address. HA sync does not assign interfaces or install plugins. - Create one NetBird network per site and one resource per VLAN, never a site
/16, withFW-GRP-<site>as the routers. - Add one outbound NAT rule per firewall: source the site's
10.<site>.0.0/16, destination the overlay, translated to thewt0address. The agent's masquerade only covers traffic leaving the overlay; without this rule, anything that starts on the VLAN side leaves with an address the far peer drops. - Start the
NETBIRDinterface at default deny and add pass rules one flow at a time.
Check the plugin's NetBird version against the server's and against advisories. Here the plugin lagged several releases behind, past a local privilege escalation fixed upstream; until the plugin caught up, tightening the permissions on the agent's socket (chmod 660 /var/run/netbird.sock) closed it.
6. Join the sites
Put every firewall and every ops-internal resource in SITE-OPS, and give that group one bidirectional policy to itself. Management networks and tenant networks are deliberately not in it.
On each firewall, two pass rules carry it, both marked quick:
- on
OPS-INT: from the local ops /24 to each remote site's ops /24; - on
NETBIRD: from each remote site's ops /24 to the local ops /24.
Do not replace them with one pass for 10.0.0.0/8. It reads tidier and quietly bypasses the block from ops to management.
Machine flows then get their own group and their own narrow policy, each one port-scoped:
- Backup servers run their own agent, in
PBS-NODES, with a self-policy on TCP 8007. Sync jobs address the far server by its overlay address. Treat a relayed pair as a fault: one pair measured 953 ms relayed against about 1 ms direct. - The secrets store's disaster-recovery voters talk Raft over the overlay on TCP 8200 and 8201 only.
- Monitoring relays at each site send to the central servers over the same mesh.
The PBS replication how-to is one of those flows from start to finish.
7. Give people access by role
Three human roles cover this estate:
| Role | Reaches |
|---|---|
PLATFORM-ADMINS | the ops and management resources at every site, and NetBird SSH |
| a single-user lab role | one person's own lab host and nothing else |
RESTRICTED-USERS | one tenant /24 and DNS |
Plus one DNS policy from the role groups to the HQ firewalls on UDP 53, and app groups for single services (TCP 443 to one /32).
A policy here is transport. It lets a laptop build a tunnel to a routing peer. Whether the packet then reaches a host is the packet filter's decision on the NETBIRD interface, and whether the person can log in is the service's (SSH certificates, the secrets store, the application). If removing a NetBird policy is the only thing that would stop someone reaching the storage network, there is one control where there should be two.
Keep the definitions in a script that is dry-run by default and reads back after every write, and run a scheduled audit for the shapes you never want: a policy sourced from "All", a resource wider than a VLAN, a peer whose endpoint sits inside a prefix it advertises.
8. Exit nodes, if you want them
One exit node per country, both firewalls of a site with different metrics, opt-in on each device and offered only to TRAVELLERS. A one-way ICMP policy to the exit-node group is enough for the route to be delivered. Do not add an all-domains nameserver: captive portals in hotels and airports need the local resolver to work.
Verify it worked
-
From the laptop,
netbird status -dlists each routing peer and says whether the path is direct or relayed, and from which endpoint. -
On a firewall, routes into the overlay are present. OPNsense's root shell is csh, so wrap it:
sh -c 'netstat -rn -f inet | grep wt0' -
Test reachability from a host that does not run a NetBird client, such as an out-of-band management controller on the ops network, and never to a firewall's own address. Both shortcuts test a different path from the one you built.
-
When something does not pass, capture on
wt0of the sending peer, and work through DNS, thenstatus -d, then the policy, then the packet filter, in that order. -
Stop the agent on one firewall of a pair and check the route moves to the other. Start it again.
Roll back
Each layer comes off on its own:
- Exit nodes: disable the routes; devices that opted in fall back to their own internet.
- A machine flow: remove its policy, then its pass rules. Watch what breaks: removing the policy that let the ops mesh reach the secrets store broke it for half an hour before it was restored.
- A site: remove its network resources and its firewalls from
SITE-OPS; the site's own jumphost is the break-glass path and never depended on the mesh. - The control plane: promote the standby and move the public name, then fail back the same way once the primary is rebuilt. Existing tunnels keep working while the control plane is down; what stops is enrolment, changes and, after the two-day expiry, people's logins. Infrastructure peers are exempt and keep running.
Be honest with yourself about that last one: the promotion script here has not been run for real yet. We did a drill to check the standby, and it uncovered a bad image tag. The server tag had been bumped with a "v", that tag did not exist, and the nightly pull had been failing quietly, so the standby could not have been promoted. The nightly job now checks the images and the restored databases instead of assuming them (step 3), upgrades go to the standby first, and the tooling refuses to upgrade the primary until the standby is on the target version. Run a real promotion on a quiet day before you need one.
Gotchas
- A PUT replaces the whole object. Send a group without its resources and the resources are gone. Read back after every write.
- A peer's tunnel endpoint must not sit inside a prefix it advertises. It encapsulates its own tunnel; small packets work and TLS dies at about 1200 bytes. Pin the endpoint to an address outside the advertised ranges.
- Overlay MTU is 1280, the measured path was nearer 1192. If large transfers stall over a relayed path, test with a smaller MSS before blaming anything else.
- Both firewalls of a pair advertise. A hook so that only the CARP master advertises is still on the list. Metrics decide which one peers use.
- A filter rule that exists only on one firewall looks fine until the other one is master. Run the HA sync after every rule change.
- The plugin lags the server. Check its version against advisories, not just against what works.
FAQ
Why not the hosted service? The control plane holds the map of every site and every peer. That belongs on infrastructure the organisation controls, and a small VM costs less than a seat per person.
Why not an agent on every server? That is a tunnel endpoint, a credential and a control-plane dependency on every server. Routing peers on the firewalls keep it to two devices per site, with the packet filter already in the path. Service peers such as backup servers are the exception, and each has a single narrow policy.
What happens if the control plane disappears for a day? Tunnels already up stay up, so backups, replication and monitoring carry on. Nobody can enrol a device or change a policy, and people start being logged out at the two-day mark. The standby exists to make the outage shorter than that.
Is this zero trust? In the useful sense: nothing is reachable because of where it sits on the network, every flow is a named group to a named resource, and every service still authenticates the person or machine on the other end. The mesh just carries the packets.
Related
- Self-hosted NetBird as the operator plane: firewalls as routing peers, policies as transport: the design decisions behind this walkthrough.
- OPNsense as the routing peer for a NetBird mesh: step 5 in full.
- PBS: pull replication between two sites with a read-only token: one machine flow over the mesh.
- A VLAN scheme you can derive in your head: where
10.<site>.<vlan>.0/24comes from.