Infrastructure engineering on a (very low) budget: start here

In short. Open source gives you everything except someone to blame. This is the series intro: the problems that recur when you build with it at the extreme end of cheap, the house rules that answer them, and a map of the posts that cover each design decision.
The setting
Three sites. Two in production, one in staging. Each has a hypervisor cluster, a firewall pair, a switch stack, a backup server and a jumphost. The compute nodes and switches are ex-lease. The firewalls are second-hand servers running an open source firewall. The backup server at the first site is two generations older than everything else. Two home disaster-recovery boxes are a 2014 Mac Mini and a small desktop with 8GB of RAM.
The software list is open source from the edge in: BGP and firewalling on OPNsense, Proxmox with Ceph and CephFS, Proxmox Backup Server, NetBird for the mesh, Authentik for people, Vault for machines, NetBox as the source of record, Zabbix and Wazuh for monitoring, PowerDNS and Kea per tenant, Stalwart for mail, Nextcloud for files. Nothing in that list has a support contract.
The organisation holds its own address space and announces it. That is the only line item that could not be made free.
What "extreme" actually costs you
The money saved is real. It is paid for in other currencies.
Nobody is coming
With a vendor, a bad day ends with a ticket. Here, a bad day ends when you find it. The firewall's config sync that only runs from the GUI, the snapshot job that reported success for months while doing nothing, the alias that was a provider's next-hop: each was found by looking, not by being told. The rule that came out of this is the most important one in the series: verify outcomes and fail loudly. A job that cannot prove it did the thing did not do the thing.
The GUI and the API are different products
Open source infrastructure tools are usually built GUI-first with an API added later. The gaps are where the surprises live. An API write that does not trigger the HA sync. An API that cannot assign an interface. An API that silently drops fields it does not recognise and wipes the rest on a partial update. Automation has to read back after every write and treat the GUI as the source of truth for anything it cannot verify.
Documentation rots faster than you write it
Every decision here has a record. Several of them were wrong within a month, not because the decision changed but because someone fixed something on the box and the repository did not hear about it. Two rules: fix things in the repository and push them, never on the host, and verify documentation against live state before trusting it. The provisioning that overwrote a hand-patched network interface was doing its job.
Dependency loops hide until the outage
The DNS token lived in the secrets store. Failing over the secrets store needed a DNS change. The secrets store's disaster-recovery voters sat behind the mesh VPN, whose secrets lived in the secrets store. None of this is visible until the day you need all of it at once. The answer is an offline go-bag and a recovery plan that is rehearsed with the normal tools switched off.
Second-hand hardware has second-hand opinions
Out-of-band controllers that reject HTTP/1.0. RAID controllers that cannot be flashed. A home rack that tripped its circuit with two servers and the firewalls on. The budget for cheap hardware includes the afternoons it costs. The rule is to buy chassis used and disks new, and to count amps before cores.
Identity cannot be outsourced to something that might be down
A cloud identity provider is a fine thing until the thing you need to log in to is the switch during an internet outage. Two roots of trust, both self-hosted, with a permanent local admin on every application for the day the identity provider is the broken part.
The dangerous moment is cleanup
Nothing in this series broke during a build. The 26-hour blackhole, the rules drifting onto one firewall, the backup that never left the building: each was discovered or caused during tidying. Retirement needs the same tooling, reference guards and read-backs as creation.
The estate at a glance
flowchart TB
I((Internet)) -- "own /24 via BGP" --> FW["OPNsense pair<br/>CARP, FRR, routing peer for the mesh"]
FW --> SW["Switch stack<br/>poison native VLAN, WAN as a VLAN"]
SW --> PVE["Proxmox cluster, 3 nodes<br/>one RBD pool, CephFS for files"]
SW --> PBS["Proxmox Backup Server<br/>pulls the other site's datastore"]
PVE --> OPS["Ops VLANs<br/>NetBox, Vault, Zabbix, Wazuh, Authentik, NetBird"]
PVE --> T["Tenant VLANs<br/>DNS and DHCP at .2, ingress at .10"]
NB["NetBird mesh<br/>operators, replication, DR voters"] -.- FW
DR["Home DR boxes<br/>warm standby, Vault voters"] -.- NBThe posts, in the order a build happens
Network
- A VLAN scheme you can derive in your head. Six network classes, one addressing rule, fixed host octets, and the per-tenant rule order.
- Second-hand Catalyst, first-class conventions. Why a stack, the poison native VLAN, the WAN as a VLAN, and the tagging mismatch that fails silently.
- OPNsense HA that actually fails over. VHID 1 everywhere, the config sync that only runs when asked, and the drift audit.
- Announce your own /24 from an OPNsense pair. Portable space, one eBGP session per firewall, the Null0 trick, and a dual-run cutover.
Platform
- Three used Dells, one Ceph pool, no local-zfs. Hardware rules, node networking, and the written ban on local storage.
- Proxmox Backup Server that leaves the building. Namespaces, retention by pool, and cross-site pull replication with a read-only token.
- CephFS subvolumes under LXC. File data outside the container, snapshots on a schedule, and the backup that actually sees the files.
Access and identity
- Self-hosted NetBird as the operator plane. Firewalls as routing peers, policies as transport, and a warm standby for the control plane.
- Two roots of trust: Authentik for people, Vault for machines. The protocol ladder, passkey-first login, and the go-bag that breaks the dependency loop.
- What a webshell taught a small hosting shop. Rebuild, never clean. Evidence before anything. The controls that followed.
This site
- An el-cheapo blog on Cloudflare. Static export, a local CMS, no monthly bill, and a deploy that runs from a laptop.
House rules, collected
- Every tool defaults to dry-run;
--applywrites. - Verify outcomes and fail loudly. Silence is not success.
- Fix it in the repository, not on the host.
- Read back after every API write.
- Nothing is "unused" until the provider says so in writing.
- Neither root of trust may depend on something it protects.
- Retirement gets the same tooling as creation.
- Buy chassis used, buy disks new, count amps before cores.
FAQ
Is this a recommendation to do the same? It is a record of what it cost and what it bought. If your organisation can pay for support, pay for it, and spend the engineering time on your own product.
What would you buy first if the budget appeared? A second transit provider at each site, then a dedicated Ceph network at the sites that still tag it on the main stack, then a second person.
Where do the company names go? Nowhere. The decisions are the point; the clients and employers behind them are not named anywhere in this series.