Second-hand Catalyst, first-class conventions

Isometric illustration of two stacked network switches with paired cables leaving the front and a small coiled console cable below.

In short. Buy a used stackable pair. Make the native VLAN one that does not exist. Carry the WAN through the stack as a tagged VLAN. Set the same stack priorities at every site. Name the switch after the site so the prompt tells you where you are. Keep a console cable on the jumphost.

Why a stack at all

A single switch is a single point of failure that every other HA decision has to apologise for. Two independent switches give you redundancy but no cross-switch link aggregation, so each server bond becomes active-backup and half your ports idle.

A two-member stack behaves as one switch with two power supplies and two sets of ports. Every server and both firewalls get a port on each member and run proper LACP. Used enterprise stackables cost about what a good prosumer switch costs new. The one I use is a twelve-year-old Catalyst design with 48 multigig ports per member and it has been fine.

Conventions that matter more than the model

The native VLAN does not exist

Every trunk uses native VLAN 998. There is no VLAN 998. An untagged frame arriving on a trunk is dropped, and nothing can accidentally land in VLAN 1 because someone forgot a tag.

The WAN is a VLAN

The transit handoff plugs into the stack, not into a firewall. It is carried as VLAN 999 through the stack and delivered to both firewalls on their trunk LAGs. That gives both firewalls the same WAN without a second handoff, and it means a firewall can be replaced without touching the provider's cable.

Firewall trunks carry everything; node trunks carry what they need

The firewall port-channels carry every VLAN, including sync and storage, because the firewalls are the gateway for all of them. Hypervisor port-channels carry the guest VLANs and management. The backup server is an access port on the storage VLAN, untagged, because it never needs anything else.

Same priorities everywhere, hostname is the FQDN

Stack master priority is 15 on member one and 12 on member two at every site. You always know which member is master without asking. The hostname is the full site FQDN, so the prompt reads as sw.site3.…# and you cannot type a change into the wrong site because you think you are somewhere else.

Config is pushed, not typed

Switch configuration lives in the repository and is pushed from the jumphost with scripts. Hand edits on the console are allowed for break-glass only and must be copied back into the repository the same day. A hand-patched interface on a hypervisor was once overwritten by the next network install; the switch is no different.

Syslog, not SNMP polling

The switches send syslog to the SIEM. Port flaps, LACP changes and login attempts show up as events next to everything else. SNMP polling was never worth its config.

flowchart TB
  WAN["Transit handoff<br/>carried as VLAN 999"]
  subgraph ST["Catalyst stack, two members, native VLAN 998 does not exist"]
    S1["Member 1<br/>priority 15"]
    S2["Member 2<br/>priority 12"]
    S1 === S2
  end
  WAN --> S1
  S1 -- "cross-stack LACP, all VLANs" --> FW1["fw01"]
  S2 -- "cross-stack LACP, all VLANs" --> FW2["fw02"]
  S1 -- "LACP bond0, tagged" --> N1["PVE node 1"]
  S2 -- "LACP bond0, tagged" --> N2["PVE node 2"]
  S1 -- "LACP bond0, tagged" --> N3["PVE node 3"]
  S2 -- "access port, untagged VLAN 5" --> PBS["Backup server"]
  J["Jumphost<br/>serial console = break-glass"] -.-> S1

Two uplinks are two ports, not one LAG

At a site with two independent transit uplinks, each goes to one stack member as its own port in VLAN 999. They are not bundled. Bundling implies one provider on both ends; independent providers need independent ports so a failure on one is visible as a link-down and not hidden inside a half-working LAG.

Tagged versus untagged by host class

This one bites quietly. Compute nodes are tagged (trunk) on the storage VLAN. The backup server is untagged (access). If a port is configured for the wrong class, LACP still bundles and the MAC table still fills from LACP frames, so the port looks healthy. Nothing passes.

The diagnosis is to look at the MAC table for the VLAN that should be busy. If it is empty, the whole layer-2 domain for that VLAN is dead on that port, not one host.

Break-glass is a cable

The jumphost has a USB serial adapter plugged into the stack console port, and the operator account is in the dialout group so no sudo is needed. When the management SVI is unreachable, the path is: mesh VPN to the jumphost, then a serial session. There is no second management network, because that would be a second thing to maintain and a second thing to get wrong.

Do not put management egress on anyone else's WiFi. One jumphost spent ten days unreachable because its only path out was a data centre guest network that someone rotated the password on.

FAQ

Why not a modern 10G or 25G switch? Because the guests do not need it and the storage network has its own path. A used multigig stack for a few hundred dollars leaves budget for disks, which matter more.

Do you really need cross-stack LACP? For the firewalls, yes: a firewall failover that also depends on a switch failover is two problems at once. For the nodes it is nice to have. For the backup server it is unnecessary.

What about the storage network at the first site? A separate unmanaged 10G switch in active-backup. It predates the stack and was never worth migrating. Later sites tag storage on the stack.

Related