Catalyst 3850: port-channel is up but no traffic passes

In short. A port-channel that shows (SU) has only proved that LACPDUs are flowing. A VLAN mismatch, trunk versus access or a VLAN missing from the allowed list, fails silently above that. Look at show interfaces trunk for the native and allowed VLANs, then at show mac address-table vlan N for the VLAN that should be busy. If that table is empty on the port, the whole VLAN is dead on it, and the fix is the port's tagging, not the host.

My port-channel says SU and both members are P, but the host cannot reach anything

You bonded a server to a Catalyst 3850 stack. show etherchannel summary looks like the textbook. The server's bond says both slaves are up with the right aggregator. And the server cannot ping its gateway, cannot be pinged, and shows no ARP entries beyond its own. It happened to me with a backup server, and the hour it cost is the reason this post exists.

What you need

  1. A WS-C3850-12X48U stack on IOS-XE 16.12 with the affected port-channel configured.
  2. The host on the other end: a Proxmox VE 8.4 node with a tagged bond0, or the backup server, whose bond is untagged on VLAN 5. The host class matters more than the model.
  3. Console or SSH access to the switch and a session on the host that does not depend on the bond under test.

Steps

  1. Confirm the symptom is really above layer 2. Pull one cable and watch show etherchannel summary: the member goes D, the channel stays SU, the host's bond reports the slave down. If all that works, LACP is fine and the problem is tagging.

  2. Look at the trunk state of the port-channel, not of the physical ports. This is where the mismatch shows.

sw-a.site.example# show interfaces trunk

Port        Mode             Encapsulation  Status        Native vlan
Po3         on               802.1q         trunking      998
Po9         on               802.1q         trunking      998

Port        Vlans allowed on trunk
Po3         1,5,10,11,100-199
Po9         5

Read it against the host class. A Proxmox node (Po3) is a trunk carrying management, storage and the guest VLANs, tagged. The backup server (Po9) must not be in this list at all: its bond is untagged, so its port-channel should be an access port on VLAN 5. If it appears here as a trunk, the server is sending untagged frames into native VLAN 998, which does not exist, and the switch drops them without a counter you will notice.

  1. Check the MAC table for the VLAN that should be busy. This is the test that distinguishes "one host is unhappy" from "this VLAN is dead on this port".
sw-a.site.example# show mac address-table vlan 5
          Mac Address Table
-------------------------------------------

Vlan    Mac Address       Type        Ports
----    -----------       --------    -----
   5    0c1e.0105.0b00    DYNAMIC     Po3
   5    0c1e.0105.0c00    DYNAMIC     Po4

The nodes are there. The backup server is not. show mac address-table interface Port-channel9 will probably show you its MAC in a different VLAN, often the port's configured native, or nothing at all. Either way, the switch has never learned a frame from that host in VLAN 5, so no amount of host-side ARP debugging will help.

  1. Fix the port to match the host class. For an untagged host, the port-channel and both members become access ports on the VLAN:
sw-a.site.example# configure terminal
sw-a.site.example(config)# interface Port-channel9
sw-a.site.example(config-if)# no switchport trunk allowed vlan
sw-a.site.example(config-if)# no switchport trunk native vlan
sw-a.site.example(config-if)# switchport mode access
sw-a.site.example(config-if)# switchport access vlan 5
sw-a.site.example(config-if)# spanning-tree portfast
sw-a.site.example(config-if)# exit
sw-a.site.example(config)# interface range GigabitEthernet1/0/20, GigabitEthernet2/0/20
sw-a.site.example(config-if-range)# switchport mode access
sw-a.site.example(config-if-range)# switchport access vlan 5
sw-a.site.example(config-if-range)# end
sw-a.site.example# write memory

For a tagged host where one VLAN is missing, the fix is the allowed list: switchport trunk allowed vlan add 5 on the port-channel. Use add; without it you replace the whole list and take the other VLANs down with it.

  1. If the summary showed a member as s (suspended) or I (individual) rather than P, the mismatch is between the two bond members, not between switch and host. Compare show running-config interface for both physical ports. One usually has a stale switchport access vlan or a different allowed list from a previous life. Make them identical to the port-channel and they will bundle.

Verify it worked

The MAC table for the VLAN should now show the port, and the host should answer:

sw-a.site.example# show mac address-table vlan 5 | include Po9
   5    0c1e.0105.0c14    DYNAMIC     Po9

From the host, ping 10.1.5.254 should succeed, and ip neigh should show the gateway's MAC as reachable. From a Proxmox node, ping the backup server's storage address and confirm the backup job that was failing now connects.

flowchart LR
  PVE["pve1<br/>bond0 tagged"] -- "Po3 trunk<br/>allowed 1,5,10,11,100-199" --> ST["Catalyst stack<br/>native 998 does not exist"]
  PBS["Backup server<br/>bond untagged"] -- "Po9 access<br/>VLAN 5" --> ST
  PBSX["Backup server<br/>if Po9 were a trunk"] -. "untagged frames land in 998<br/>and vanish" .-> ST

Gotchas

FAQ

Why does the switch not log a native VLAN mismatch? It does, but only when CDP or LLDP sees a peer that announces a different native VLAN, which means another switch. A Linux host does not announce one. The frames are dropped as ordinary untagged traffic on a trunk whose native VLAN has no members.

Could I just create VLAN 998 to make it work? You could, and everything untagged in the estate would silently join one broadcast domain. The poison native VLAN exists to turn tagging mistakes into dead ports rather than into cross-tenant leaks. Fix the port instead.

Does show etherchannel summary ever reveal this? No. The flags describe the aggregation state only. The closest it gets is a member flagged s when the members disagree with each other, which is a different problem and usually a stale per-port setting.

Related