Catalyst 3850: upgrade IOS-XE on a stack in install mode
In short.
show versionmust sayINSTALLon every member. Free space withrequest platform software package clean switch all file flash:, copy the.bin, check its MD5, thenrequest platform software package install switch all file flash:<image> new auto-copyand reload. Do it in a window, from the console, and keep the previous image on flash until the stack has run a week on the new one.
I need to upgrade a Catalyst 3850 stack and not end up with one member on each version
Upgrading a stack is one command, and the one command has a few preconditions that are easy to skip. The failure mode when you skip them is a member that boots the old image, or no image, and a stack that is half upgraded. This is the sequence I run. Why the stack exists, and why its members must match before a new one joins, is in the decision post at the end.
What you need
- A WS-C3850-12X48U stack on IOS-XE 16.12 in install mode, both members
Readyinshow switch. - The target image as a single
.binfile, with its published MD5, on a host the switch can reach with SCP or TFTP. In the examples the image iscat3k_caa-universalk9.16.12.xx.SPA.binon the jumphost at10.1.1.20. - A console session on the master from the jumphost. The reload takes the management SVI with it, and the console is where you watch both members come back.
- A maintenance window. Every port-channel in the estate goes down for the duration of the reload, which is around ten minutes on this hardware. Tell the people who will notice.
Steps
- Check the mode and version on every member. The table at the bottom of
show versionhas aModecolumn; both rows must sayINSTALL. Bundle mode is a different procedure and a different post. Also confirm there is room for the new packages, because the install expands the.bininto several files:
sw-a.site.example# show version
sw-a.site.example# dir flash: | include bytes free
sw-a.site.example# dir flash-2: | include bytes free
- Clean out packages from previous upgrades. This removes anything on flash that the current
packages.confdoes not reference, on every member, and asks before it deletes:
sw-a.site.example# request platform software package clean switch all file flash:
- Copy the image to the master and verify it. Copying to
flash:on the master is enough;auto-copyin the install step distributes it to the other member.
sw-a.site.example# copy scp://ops@10.1.1.20/images/cat3k_caa-universalk9.16.12.xx.SPA.bin flash:
sw-a.site.example# verify /md5 flash:cat3k_caa-universalk9.16.12.xx.SPA.bin
Compare the hash to the one published for the release. A truncated copy over a slow link is the second most common cause of a member that will not boot.
- Take a configuration backup now, from the repository's script, so the state before the upgrade is committed. Confirm the boot variable points at
packages.conf, which is what install mode boots from:
sw-a.site.example# show boot
- Install to every member.
newexpands the image into fresh packages and writes a newpackages.conf;auto-copypushes the.binto the other member first. The command runs for several minutes and ends by asking whether to reload.
sw-a.site.example# request platform software package install switch all file flash:cat3k_caa-universalk9.16.12.xx.SPA.bin new auto-copy
Read the output before answering. It lists each member and the packages it wrote. If one member reports an error, do not reload; the old packages.conf is still what the boot variable points at on that member, and you can fix the copy and run the install again.
- Reload the stack from the console with the window open:
sw-a.site.example# write memory
sw-a.site.example# reload
Both members reload together. Watch the console; the master's boot output shows the new version, then the stack forming, then the second member reaching Ready.
- Keep the old
.binon flash. Do not run the clean step again until the stack has run on the new release long enough that you would not want to go back.
Verify it worked
sw-a.site.example# show version
sw-a.site.example# show switch
sw-a.site.example# show etherchannel summary
show version must show the new version and INSTALL on both rows of the table. show switch must show both Ready with the same priorities as before; an upgrade does not change them, but you are checking that nothing else did. show etherchannel summary should show every port-channel SU with its members P again. From a Proxmox node, confirm /proc/net/bonding/bond0 sees both slaves up, and from a VM, ping a gateway. Then run the configuration backup and confirm the diff against the pre-upgrade capture is empty or explainable.
flowchart TB
A["show version: INSTALL on both"] --> B["clean, copy, verify MD5"]
B --> C["install switch all ... new auto-copy"]
C --> D{"every member OK?"}
D -- "yes" --> E["reload from console"]
D -- "no" --> F["fix the copy, run again,<br/>old packages.conf still boots"]
E --> G["show version, show switch,<br/>show etherchannel summary"]
G -- "problem" --> H["install the previous .bin<br/>with the same command"]Gotchas
- Upgrade the stack, not a member. Installing to one member and then another leaves a mismatched stack between the two reloads, and the second member may refuse to join when it comes back.
switch alland one reload. - The install writes packages to flash on every member. If a member's flash is fuller than the master's, the install fails only on that member, and the error is easy to miss in the scrollback. Check free space on both before starting.
- Auto-upgrade of a new member joining on an older release is unreliable with several members. Upgrade a new switch on its own, with this procedure, before it is cabled into the stack.
- The reload drops the management SVI, which drops SSH. If you start the reload from an SSH session you will stare at a dead terminal for ten minutes without knowing which stage the boot is at. The console shows you.
- Rollback is installing the previous image with the same command. The
request platform software packagefamily has arollbackkeyword on some releases, but I have not relied on it. The previous.binon flash and a known-goodpackages.confare the rollback plan I trust, which is why the clean step waits.
FAQ
Can I do this over SSH if the console is not available?
You can issue the commands over SSH; the install step itself does not need the console. The reload does, in the sense that without it you cannot see why a member did not come back. If the jumphost is in the rack, the console is one screen away.
Do I need to upgrade the ROMMON separately? On this platform the install step updates the boot loader when the image requires it and says so in its output. It can add a second reload. Read what it prints rather than assuming.
How long does the whole thing take?
Copying the image, five to fifteen minutes depending on the link. The install step, five to ten. The reload, about ten until both members are Ready and port-channels have bundled. Book an hour and expect to use half.