A VLAN scheme you can derive in your head

In short. Six network classes, one addressing rule, fixed host octets, and a firewall rule order that keeps tenants apart. Nothing in this scheme needs a spreadsheet to look up. That is the whole point.
Who this is for
You are standing up a first colocation rack, or a homelab that has quietly become production. You know you need VLANs. You have read that "management should be separate". What you have not found is a scheme that is small enough to remember and strict enough to survive the third site.
This is the scheme I use. It has run three sites and a couple of home disaster-recovery boxes, and the parts that changed over that time are listed at the end.
The six network classes
Every site has the same VLAN numbers with the same meaning. Only the site number changes.
| VLAN | Name | What lives there | Internet |
|---|---|---|---|
| 1 | MGMT | Hardware only: out-of-band controllers, hypervisors, backup server, switches, jumphost | Nothing inbound |
| 2 | FW-SYNC | Firewall state and config sync | None |
| 5 | STORAGE | Ceph public and cluster traffic, and all backup traffic. 10G, jumbo frames | None |
| 10 | OPS-INT | Source of record, secrets, monitoring, SIEM | 10/8 only; packages via a proxy |
| 11 | OPS-EXT | Mesh VPN controller, identity provider, ingress | Outbound NAT plus specific inbound |
| 100 + N | Tenant N | One tenant, one VLAN | Allowed |
Two rules make this table work:
- Management is for hardware, not for services. If it has a web UI that people log into to do work, it is not management. It is ops.
- Ops is split by whether it ever faces the internet. The identity provider and the VPN controller must be reachable from outside, so they sit apart from the secrets store and the SIEM, which must not.
The addressing rule
10.<site>.<vlan>.<host>/24 gateway is always .254
Site 1, VLAN 10, host 20 is 10.1.10.20. Site 3, tenant 7 (VLAN 107), its DNS server is 10.3.107.2. Anyone on the team can read an address and know where they are and what they are talking to.
Three consequences follow:
- A region is a /14. Four sites fit in a summary route, which is what the firewalls and the BGP policy use.
- 172 and 192 ranges are never used. A rule that says "10/8" is a rule that says "the whole private estate", and there are no exceptions to remember.
- One range is reserved and never assigned to a VLAN. The mesh VPN overlay has
10.254.0.0/16. It is in the scheme precisely so that nobody allocates it by accident.
IPv6 mirrors IPv4: <prefix>:<site>:<vlan>::/64. The decimal site number is written as if it were hex, so site 12 is :12: in both families and reads the same. A test in the repository enforces this.
Fixed host octets inside a tenant
A tenant VLAN is a /24 with a layout that never changes:
| Host | Role |
|---|---|
| .2 | DNS and DHCP container for that tenant |
| .10 | Ingress: reverse proxy and the identity provider's outpost |
| .11 to .199 | Static guests |
| .200 to .249 | Provisioning DHCP pool |
| .254 | Gateway (the firewall's shared address) |
MAC addresses are derived from the IP: a fixed prefix, then the VLAN, site and host. When a switch shows you a MAC, you know what it is without a lookup. When a DHCP lease looks wrong, you can tell at a glance.
The firewall rule order per tenant
Every tenant gets the same four rules, in this order, generated by a tool rather than typed:
- Block to MGMT
- Block to every other tenant
- Block to both ops VLANs
- Allow to the internet
The order matters because the allow is broad. Put it first and the blocks never match. The tool that creates a tenant writes all four rules in one pass, and a drift audit compares what is on the firewall with what the tool would generate.
flowchart LR
subgraph S["One site: 10.S.V.0/24 per VLAN, gateway .254"]
M["VLAN 1 MGMT<br/>out-of-band, hypervisors, backup, switches"]
Y["VLAN 2 FW-SYNC"]
C["VLAN 5 STORAGE<br/>Ceph and backup, 10G jumbo"]
OI["VLAN 10 OPS-INT<br/>source of record, secrets, monitoring"]
OE["VLAN 11 OPS-EXT<br/>mesh VPN, identity, ingress"]
T1["VLAN 101 tenant 1<br/>.2 dns, .10 ingress, .254 gw"]
T2["VLAN 102 tenant 2"]
end
I((Internet))
P["Package proxy"]
T1 -- "allow" --> I
T2 -- "allow" --> I
OE -- "outbound NAT + specific inbound" --> I
OI -- "10/8 only" --> P --> I
T1 -. "deny" .-> T2
T1 -. "deny" .-> M
T1 -. "deny" .-> OI
OV["10.254/16 reserved<br/>mesh overlay, never a VLAN"]Things that went wrong, and what they taught
- A hypervisor bridge held the gateway address. Someone gave a host bridge the
.254of a VLAN "temporarily". The firewall's shared address and the bridge both answered ARP, and every guest on that VLAN lost its gateway intermittently. Rule now:.254belongs to the firewall pair, full stop, and an audit checks for it. - DHCP on the firewall did not scale. Per-tenant DHCP on the firewall meant per-tenant state on the device that must change least. It moved into the tenant's own container at
.2, next to its DNS. - Split-horizon DNS fooled a test. Testing a "public" URL from inside the estate resolves to the internal address and tests the internal path. Public checks run from outside.
What changed between site one and site three
- The firewall sync VLAN (2) did not exist at the first site, which used a crossover cable. It was added when the second site needed a sync path through the switch stack.
- Storage at the first site runs on a separate unmanaged 10G switch, so VLAN 5 is physical there. Later sites tag it on the main stack.
FAQ
Why not a /16 per site with VLAN as the third octet anyway?
That is exactly what this is. Writing it as 10.<site>.<vlan>.<host> just makes the role of each octet explicit, and the /14 per region is a free consequence.
Why 100 + N for tenants instead of starting at 2? So that no tenant number can collide with an infrastructure VLAN, and so the VLAN number tells you it is a tenant without looking anything up.
What about VLAN 1 being the default VLAN on switches? The switch trunks use a non-existent "poison" native VLAN, so untagged frames go nowhere. VLAN 1 is tagged like any other. The switching post covers that.