How-to guides
How-to guides: 33 posts, step-by-step guides with commands, checks and the traps, each linked to the decision behind it.
- OPNsense: announce your own prefix with FRR BGP
The FRR plugin on an OPNsense pair is enough to announce a portable /24 to a transit provider. The settings that matter, the blackhole route without which nothing is advertised, strict filters both ways, and how to confirm the announcement from outside.
- OPNsense: retire a public IP address without a blackhole
An alias nothing referenced was deleted during a cleanup and a routed /28 went dark for about 26 hours. The procedure that came out of it, from the inventory of references to the provider's written confirmation, dual-running the replacement, deleting on both nodes and probing from outside.
- OPNsense as the routing peer for a NetBird mesh
Install the NetBird agent on both firewalls of an HA pair, advertise one route per VLAN with the pair as routing peers, add the outbound NAT the agent's masquerade does not cover, filter on the overlay interface, and avoid the endpoint trap that drops large packets.
- Catalyst 3850: console access from a Linux jumphost and config backups
A USB serial adapter on the stack's console port, the operator in the dialout group, screen at 9600 8N1, and a small expect script that captures the running configuration into the repository. Plus the break-glass path and why management egress never goes over anyone else's WiFi.
- Catalyst 3850: send syslog to a SIEM instead of polling SNMP
Point the switch at the SIEM with logging host, raise the trap level to informational, timestamp with milliseconds, log configuration changes with archive, and let Wazuh decode the rest. Port flaps, LACP changes and failed logins become events next to everything else.
- OPNsense: keep ACME certificates valid on both HA nodes
The ACME client renews on the node it runs on. The backup keeps serving whatever it had, which after ninety days is an expired certificate at the worst possible moment. A daily copy, an import on the backup, and a fingerprint check that tells you when they differ.
- OPNsense: fix HA config sync that stopped syncing
The backup looks healthy and is a month behind. How to tell, what actually triggers the XMLRPC sync, what never syncs at all, and a drift check you can run from a laptop every fifteen minutes.
- PBS: a lot of "verify failed", and the protect-latest trick
When a verify job lights up dozens of snapshots at once, the first job is to stop prune from making it worse. Protect the newest snapshot in every group, read the log for the one chunk that is shared, then decide between re-uploading from the source, re-syncing from the other site, and fixing the disk.
- PBS: fix the fingerprint mismatch after a rename or reinstall
Rename or reinstall a PBS host and it presents a new self-signed certificate. Every PVE storage entry and every remote on the other PBS has the old fingerprint pinned and refuses to connect. Where the pins live, how to read the new fingerprint, and the order to update them in.
- Proxmox: move VMs between two clusters
qm remote-migrate exists and is fussy. The path that works every time is a Proxmox Backup Server backup, already replicated to the other site, restored on the other cluster from a read-only storage entry with a new guest ID.
- PBS: a restore drill for one file, one guest and one site
A backup you have not restored from is a hope. Three rehearsed restores, each with its commands and its evidence: one file from a pxar archive, one guest to a new VMID on the local PBS, and one guest on the other cluster from the replica namespace. What the first drill found, how often to run it, and what to write down.
- Proxmox: SDN VLAN zones with an external IPAM
Two VLAN zones on the VLAN-aware bridge, a VNet per VLAN, and NetBox as the IPAM behind the tenant zone so the subnet Proxmox knows about is the one the address plan allocated. SDN does the plumbing; the firewall still does policy and NAT.
- PBS: back up a directory with proxmox-backup-client from a snapshot
vzdump skips bind mounts, so the files a container serves from CephFS never reach the backup server. A nightly proxmox-backup-client run takes a pxar archive of the latest CephFS snapshot into the tenant's namespace. The token, the command, the timer, and the one-file restore that proves it.
- Proxmox: bind-mount CephFS into an unprivileged container
Create a CephFS subvolume with a quota, scope a cephx key to its path, mount it on every node with a systemd unit, and bind-mount it into an unprivileged container that holds no Ceph key. df inside the container shows the quota, not the cluster.
- Proxmox: fix "cannot migrate local bind mount point"
The error means Proxmox cannot promise the directory behind mp0 exists on the target node. Mount the same path on every node with a systemd unit, check it, and only then add shared=1.
- Proxmox: replace a failed OSD, or reinstall and rejoin a node
Two runbooks that share a shape. For a dead disk: identify, out, stop, destroy, swap, create, watch recovery. For a dead node: drain it, remove its Ceph roles, pvecm delnode, reinstall by PXE, pvecm add, recreate the roles, and keep HA fencing in mind throughout.
- Proxmox: a cluster API address that follows a healthy node
One address for the Proxmox API that moves to a node whose pveproxy actually answers, using keepalived VRRP with a track script. A node with a wedged API still answers ARP, so the check has to ask port 8006, not the kernel.
- Proxmox: a three-node Ceph cluster on used Dell servers
From three second-hand R630-class servers with IT-mode HBAs to a quorate cluster with one replicated RBD pool, HA on every guest and no local storage for guests. The commands, in order, and what to check after each.
- Proxmox: automated install with an answer file and PXE
Build a self-installing Proxmox VE ISO with proxmox-auto-install-assistant, serve the TOML answer from the jumphost, boot it through iDRAC virtual media or PXE, and avoid the two traps that cost me an afternoon each.
- PBS: prune, verify and garbage collection that do not fight
Prune decides what to keep, garbage collection reclaims what nothing references, verify checks that what is kept is still readable. Each is harmless alone and all three get in each other's way when they overlap. A weekly calendar that keeps them apart, and apart from the backup window and the restore drill.
- OPNsense: split-horizon DNS with Unbound forward zones
Internal zones forwarded to the per-site authoritative server, public names overridden to internal addresses for internal clients, and DNS over TLS to the upstream. Plus the test that proves the public answer is right, which is not the one you run from your laptop.
- Proxmox: intra-day snapshots from cron that actually run
A snapshot script that ran from cron three times a day, reported success for months, and never created a snapshot, because cron's PATH does not include /usr/sbin and the output went to /dev/null. The fixed script, and the checks that make it fail loudly.
- PBS: sync job says TASK OK but the datastore is empty
A pull sync that finishes green and copies nothing has one of three causes: the namespaces do not line up, an earlier run is still holding the locks, or the token cannot see anything upstream. One check for each, and a way to see what the job sees before running it.
- PBS: pull replication between two sites with a read-only token
The remote PBS pulls, with a token that can only read one datastore, a pinned fingerprint, and a per-source namespace so guest IDs from two clusters never collide. Every command, in order, then the storage entry that lets the surviving cluster restore.
- OPNsense: manage aliases, rules and NAT from Python
The OPNsense API is good enough to run an estate's firewall policy from a script, as long as the script reads back every write, marks what it owns, and remembers that the HA sync will not run on its own. A working pattern with requests.
- Catalyst 3850: upgrade IOS-XE on a stack in install mode
Confirm install mode, clean the flash, copy and verify the image, install it to every member in one command, reload with a console attached, and keep the old image on flash for the rollback you hope not to need.
- OPNsense: VLANs on a trunk from a Catalyst, with Kea DHCP
Adding a tenant VLAN to an OPNsense pair fed by a Catalyst trunk, from the switch port to the CARP gateway, the firewall rule order that keeps tenants apart, and a Kea DHCP pool for provisioning. Done twice, because there are two nodes.
- Catalyst 3850: port-channel is up but no traffic passes
LACP bundles, the MAC table has entries, and nothing gets through. The port is tagging the wrong way for its host class, and the switch has no reason to tell you.
- Catalyst 3850: cross-stack LACP to a Proxmox bond, both sides
One port on each stack member, a port-channel in LACP active mode, a trunk whose native VLAN does not exist, and an 802.3ad bond under a VLAN-aware bridge on the Proxmox node. Both halves, and how to check they agree.
- Catalyst 3850: trunks with a native VLAN that does not exist
Set every trunk's native VLAN to one you never create, so untagged frames vanish instead of landing in VLAN 1. Then carry the provider's WAN handoff through the stack as a tagged VLAN to both firewalls.
- Catalyst 3850: stack two switches and add a member safely
Cable the stack ring, set priorities in EXEC mode (not config mode), match the software before the new member joins, and power on one switch at a time. The traps are the priority command and auto-upgrade.
- Proxmox: Ubuntu cloud-init templates with static addresses
Turn an Ubuntu cloud image into a Proxmox template, clone it, and give each clone a static address, gateway, nameserver, SSH key, VLAN tag and a MAC derived from the address. Every value comes from the host plan, so nothing is typed twice.
- OPNsense: port forward to a VM behind a CARP pair
A port forward on an HA pair has to land on the shared WAN address, not a node's own, and it has to exist on both nodes. The forward, its filter rule, why NAT reflection is the wrong fix for the inside test, and how to test from outside properly.