Catalyst 3850: port-channel is up but no traffic passes
In short. A port-channel that shows
(SU)has only proved that LACPDUs are flowing. A VLAN mismatch, trunk versus access or a VLAN missing from the allowed list, fails silently above that. Look atshow interfaces trunkfor the native and allowed VLANs, then atshow mac address-table vlan Nfor the VLAN that should be busy. If that table is empty on the port, the whole VLAN is dead on it, and the fix is the port's tagging, not the host.
My port-channel says SU and both members are P, but the host cannot reach anything
You bonded a server to a Catalyst 3850 stack. show etherchannel summary looks like the textbook. The server's bond says both slaves are up with the right aggregator. And the server cannot ping its gateway, cannot be pinged, and shows no ARP entries beyond its own. It happened to me with a backup server, and the hour it cost is the reason this post exists.
What you need
- A WS-C3850-12X48U stack on IOS-XE 16.12 with the affected port-channel configured.
- The host on the other end: a Proxmox VE 8.4 node with a tagged
bond0, or the backup server, whose bond is untagged on VLAN 5. The host class matters more than the model. - Console or SSH access to the switch and a session on the host that does not depend on the bond under test.
Steps
-
Confirm the symptom is really above layer 2. Pull one cable and watch
show etherchannel summary: the member goesD, the channel staysSU, the host's bond reports the slave down. If all that works, LACP is fine and the problem is tagging. -
Look at the trunk state of the port-channel, not of the physical ports. This is where the mismatch shows.
sw-a.site.example# show interfaces trunk
Port Mode Encapsulation Status Native vlan
Po3 on 802.1q trunking 998
Po9 on 802.1q trunking 998
Port Vlans allowed on trunk
Po3 1,5,10,11,100-199
Po9 5
Read it against the host class. A Proxmox node (Po3) is a trunk carrying management, storage and the guest VLANs, tagged. The backup server (Po9) must not be in this list at all: its bond is untagged, so its port-channel should be an access port on VLAN 5. If it appears here as a trunk, the server is sending untagged frames into native VLAN 998, which does not exist, and the switch drops them without a counter you will notice.
- Check the MAC table for the VLAN that should be busy. This is the test that distinguishes "one host is unhappy" from "this VLAN is dead on this port".
sw-a.site.example# show mac address-table vlan 5
Mac Address Table
-------------------------------------------
Vlan Mac Address Type Ports
---- ----------- -------- -----
5 0c1e.0105.0b00 DYNAMIC Po3
5 0c1e.0105.0c00 DYNAMIC Po4
The nodes are there. The backup server is not. show mac address-table interface Port-channel9 will probably show you its MAC in a different VLAN, often the port's configured native, or nothing at all. Either way, the switch has never learned a frame from that host in VLAN 5, so no amount of host-side ARP debugging will help.
- Fix the port to match the host class. For an untagged host, the port-channel and both members become access ports on the VLAN:
sw-a.site.example# configure terminal
sw-a.site.example(config)# interface Port-channel9
sw-a.site.example(config-if)# no switchport trunk allowed vlan
sw-a.site.example(config-if)# no switchport trunk native vlan
sw-a.site.example(config-if)# switchport mode access
sw-a.site.example(config-if)# switchport access vlan 5
sw-a.site.example(config-if)# spanning-tree portfast
sw-a.site.example(config-if)# exit
sw-a.site.example(config)# interface range GigabitEthernet1/0/20, GigabitEthernet2/0/20
sw-a.site.example(config-if-range)# switchport mode access
sw-a.site.example(config-if-range)# switchport access vlan 5
sw-a.site.example(config-if-range)# end
sw-a.site.example# write memory
For a tagged host where one VLAN is missing, the fix is the allowed list: switchport trunk allowed vlan add 5 on the port-channel. Use add; without it you replace the whole list and take the other VLANs down with it.
- If the summary showed a member as
s(suspended) orI(individual) rather thanP, the mismatch is between the two bond members, not between switch and host. Compareshow running-config interfacefor both physical ports. One usually has a staleswitchport access vlanor a different allowed list from a previous life. Make them identical to the port-channel and they will bundle.
Verify it worked
The MAC table for the VLAN should now show the port, and the host should answer:
sw-a.site.example# show mac address-table vlan 5 | include Po9
5 0c1e.0105.0c14 DYNAMIC Po9
From the host, ping 10.1.5.254 should succeed, and ip neigh should show the gateway's MAC as reachable. From a Proxmox node, ping the backup server's storage address and confirm the backup job that was failing now connects.
flowchart LR
PVE["pve1<br/>bond0 tagged"] -- "Po3 trunk<br/>allowed 1,5,10,11,100-199" --> ST["Catalyst stack<br/>native 998 does not exist"]
PBS["Backup server<br/>bond untagged"] -- "Po9 access<br/>VLAN 5" --> ST
PBSX["Backup server<br/>if Po9 were a trunk"] -. "untagged frames land in 998<br/>and vanish" .-> STGotchas
- The MAC table is populated by LACPDUs and by CDP or LLDP, which is why it looks alive. Those entries sit in the native VLAN or in no VLAN. Always filter by the VLAN you care about.
switchport trunk allowed vlan 5withoutaddreplaces the list. On a Proxmox node's port-channel that takes every guest off the air at once. The repository holds the full list; push it whole rather than patching on the console.- Changing a port-channel from trunk to access while the members are still trunks causes the members to unbundle. Change the port-channel and then the members in the same session, as in step 4.
show interfaces Po9 switchporttells you the operational mode, which can differ from the administrative one if DTP negotiated something.switchport nonegotiateon every port removes that variable.- A host whose bond is tagged on the wrong VLAN ID looks exactly the same from the switch. If the switch side is right by every check above, read the host's
/etc/network/interfacesfor the VLAN sub-interface number before changing anything else.
FAQ
Why does the switch not log a native VLAN mismatch? It does, but only when CDP or LLDP sees a peer that announces a different native VLAN, which means another switch. A Linux host does not announce one. The frames are dropped as ordinary untagged traffic on a trunk whose native VLAN has no members.
Could I just create VLAN 998 to make it work? You could, and everything untagged in the estate would silently join one broadcast domain. The poison native VLAN exists to turn tagging mistakes into dead ports rather than into cross-tenant leaks. Fix the port instead.
Does show etherchannel summary ever reveal this?
No. The flags describe the aggregation state only. The closest it gets is a member flagged s when the members disagree with each other, which is a different problem and usually a stale per-port setting.