Infrastructure engineering on a (very low) budget: start here

Isometric illustration of a short rack of mismatched second-hand servers next to a small desktop computer on a shelf, one cable leading away.

In short. Open source gives you everything except someone to blame. This is the series intro: the problems that recur when you build with it at the extreme end of cheap, the house rules that answer them, and a map of the posts that cover each design decision.

The setting

Three sites. Two in production, one in staging. Each has a hypervisor cluster, a firewall pair, a switch stack, a backup server and a jumphost. The compute nodes and switches are ex-lease. The firewalls are second-hand servers running an open source firewall. The backup server at the first site is two generations older than everything else. Two home disaster-recovery boxes are a 2014 Mac Mini and a small desktop with 8GB of RAM.

The software list is open source from the edge in: BGP and firewalling on OPNsense, Proxmox with Ceph and CephFS, Proxmox Backup Server, NetBird for the mesh, Authentik for people, Vault for machines, NetBox as the source of record, Zabbix and Wazuh for monitoring, PowerDNS and Kea per tenant, Stalwart for mail, Nextcloud for files. Nothing in that list has a support contract.

The organisation holds its own address space and announces it. That is the only line item that could not be made free.

What "extreme" actually costs you

The money saved is real. It is paid for in other currencies.

Nobody is coming

With a vendor, a bad day ends with a ticket. Here, a bad day ends when you find it. The firewall's config sync that only runs from the GUI, the snapshot job that reported success for months while doing nothing, the alias that was a provider's next-hop: each was found by looking, not by being told. The rule that came out of this is the most important one in the series: verify outcomes and fail loudly. A job that cannot prove it did the thing did not do the thing.

The GUI and the API are different products

Open source infrastructure tools are usually built GUI-first with an API added later. The gaps are where the surprises live. An API write that does not trigger the HA sync. An API that cannot assign an interface. An API that silently drops fields it does not recognise and wipes the rest on a partial update. Automation has to read back after every write and treat the GUI as the source of truth for anything it cannot verify.

Documentation rots faster than you write it

Every decision here has a record. Several of them were wrong within a month, not because the decision changed but because someone fixed something on the box and the repository did not hear about it. Two rules: fix things in the repository and push them, never on the host, and verify documentation against live state before trusting it. The provisioning that overwrote a hand-patched network interface was doing its job.

Dependency loops hide until the outage

The DNS token lived in the secrets store. Failing over the secrets store needed a DNS change. The secrets store's disaster-recovery voters sat behind the mesh VPN, whose secrets lived in the secrets store. None of this is visible until the day you need all of it at once. The answer is an offline go-bag and a recovery plan that is rehearsed with the normal tools switched off.

Second-hand hardware has second-hand opinions

Out-of-band controllers that reject HTTP/1.0. RAID controllers that cannot be flashed. A home rack that tripped its circuit with two servers and the firewalls on. The budget for cheap hardware includes the afternoons it costs. The rule is to buy chassis used and disks new, and to count amps before cores.

Identity cannot be outsourced to something that might be down

A cloud identity provider is a fine thing until the thing you need to log in to is the switch during an internet outage. Two roots of trust, both self-hosted, with a permanent local admin on every application for the day the identity provider is the broken part.

The dangerous moment is cleanup

Nothing in this series broke during a build. The 26-hour blackhole, the rules drifting onto one firewall, the backup that never left the building: each was discovered or caused during tidying. Retirement needs the same tooling, reference guards and read-backs as creation.

The estate at a glance

flowchart TB
  I((Internet)) -- "own /24 via BGP" --> FW["OPNsense pair<br/>CARP, FRR, routing peer for the mesh"]
  FW --> SW["Switch stack<br/>poison native VLAN, WAN as a VLAN"]
  SW --> PVE["Proxmox cluster, 3 nodes<br/>one RBD pool, CephFS for files"]
  SW --> PBS["Proxmox Backup Server<br/>pulls the other site's datastore"]
  PVE --> OPS["Ops VLANs<br/>NetBox, Vault, Zabbix, Wazuh, Authentik, NetBird"]
  PVE --> T["Tenant VLANs<br/>DNS and DHCP at .2, ingress at .10"]
  NB["NetBird mesh<br/>operators, replication, DR voters"] -.- FW
  DR["Home DR boxes<br/>warm standby, Vault voters"] -.- NB

The posts, in the order a build happens

Network

Platform

Access and identity

This site

House rules, collected

  1. Every tool defaults to dry-run; --apply writes.
  2. Verify outcomes and fail loudly. Silence is not success.
  3. Fix it in the repository, not on the host.
  4. Read back after every API write.
  5. Nothing is "unused" until the provider says so in writing.
  6. Neither root of trust may depend on something it protects.
  7. Retirement gets the same tooling as creation.
  8. Buy chassis used, buy disks new, count amps before cores.

FAQ

Is this a recommendation to do the same? It is a record of what it cost and what it bought. If your organisation can pay for support, pay for it, and spend the engineering time on your own product.

What would you buy first if the budget appeared? A second transit provider at each site, then a dedicated Ceph network at the sites that still tag it on the main stack, then a second person.

Where do the company names go? Nowhere. The decisions are the point; the clients and employers behind them are not named anywhere in this series.