Skip to content

Pi 5: onboard Ethernet (RP1/macb/BCM54213PE) drops with "Link is Down" and never recovers; second incident froze the whole system in rp1_adc_read #7653

Description

@tushar-arora1

Summary

Two incidents in ~13 hours on the same Raspberry Pi 5. Both began with macb ... end0: Link is Down after hours of healthy operation. Both left the driver's hardware statistics returning garbage. Only a reboot recovered. Incident 1 affected Ethernet only (WiFi kept working). Incident 2 froze the whole machine, with a CPU stuck spinning in rp1_adc_read.

This looks like the RP1 chip (or the PCIe path to it) wedging, with Ethernet as one symptom. Possibly related to #7345 and #6948, and the forum thread "Pi 5 kernel oops at xhci_halt, intermittent FC0 timeouts on pll_sys_sec/clk_eth, RP1 defect" (forums.raspberrypi.com, p=2375780). I have not confirmed the root cause.

Environment

  • Raspberry Pi 5 Model B Rev 1.0 (Revision: d04170), 64-bit Raspberry Pi OS (Debian 12 bookworm)
  • Incidents on kernel 6.12.96+rpt-rpi-2712 (1:6.12.96-1+rpt1); now running 6.12.109+rpt-rpi-2712 (not yet observed long enough to say whether it helps)
  • Bootloader EEPROM 2026-05-26 (1779807685), up to date per rpi-eeprom-update; firmware 086b83e3 (2026/05/26); raspi-firmware 1.20260907 (now 1.20260915)
  • Power: vcgencmd get_throttled = 0x0 when checked after the fact; no undervoltage messages found in the kernel journal
  • Cmdline includes pci=pcie_bus_safe; arm_boost=1; no Ethernet-related overlays or dtparams
  • The Ethernet link has negotiated 100 Mbps/Full on every boot observed (2026-09-26, 09-27, 09-28), never 1000, although the PHY advertises 1000baseT
  • Workload: Docker (about 16 containers) incl. Jellyfin and a Netdata container with /sys mounted read-only from the host (pid: host); this is a home server
  • WiFi (wlan0) is on 1001100000.mmc (SoC SDIO), not behind RP1

Incident 1: 2026-09-27 17:38, Ethernet only

Sep 26 19:28:42 kernel: macb 1f00100000.ethernet end0: Link is Up - 100Mbps/Full - flow control off
Sep 27 17:38:01 kernel: macb 1f00100000.ethernet end0: Link is Down
  • Roughly 22 h of clean operation before (0 rx/tx errors, gateway ping ~0.5 ms; a watcher script logged snapshots every 5 min).
  • No kernel warning before the drop; only unrelated Docker veth churn earlier at 17:02 and 17:10. SoC temperature 56-58 C.
  • Immediately after the drop the counters were nonsense, e.g. ip -s link RX errors 4294967208 and, five minutes later, ethtool -S end0 tx_frames: 609886253403 while q0_tx_packets stayed at the pre-drop value (855376).
  • ethtool end0: Link detected: no, master-slave status: resolution error.
  • NetworkManager later moved the device to unavailable. WiFi stayed up and the Pi remained reachable over it.
  • Cable re-seated (by the owner): no effect.
  • ip link set end0 down/up: no effect.
  • Unbind/rebind of the driver made it worse:
echo 1f00100000.ethernet > /sys/bus/platform/drivers/macb/unbind    (ok)
echo 1f00100000.ethernet > /sys/bus/platform/drivers/macb/bind      (Invalid argument)

macb 1f00100000.ethernet: gem-ptp-timer ptp clock unregistered.
rp1-clk 1f00018000.clocks: pll_sys_sec: FC0 busy timeout
rp1-clk 1f00018000.clocks: clk_eth: FC0 busy timeout
rp1-clk 1f00018000.clocks: clk_eth_tsu: FC0 busy timeout
macb 1f00100000.ethernet end0: PHY [...] driver [Broadcom BCM54213PE] (irq=POLL)
macb 1f00100000.ethernet: gem-ptp-timer ptp clock registered.
macb 1f00100000.ethernet: gem-ptp-timer ptp clock unregistered.
macb 1f00100000.ethernet: error -ENXIO: IRQ index 1 not found
macb 1f00100000.ethernet: Unable to request IRQ -6 (error -22)
macb 1f00100000.ethernet: probe with driver macb failed with error -22

end0 no longer existed after this. A reboot restored it (link up, 100 Mbps).

Incident 2: 2026-09-28 02:44, whole-system freeze

Healthy until 02:40 (gateway ping 0.5 ms via end0, zero errors). Then:

02:44:10 kernel: macb 1f00100000.ethernet end0: Link is Down
02:44:11 (watcher) ping via end0 FAILED; rx_err=4294967288 tx_err=4294967288   <- same garbage counters
02:44:31 kernel: rcu: INFO: rcu_preempt self-detected stall on CPU
         CPU: 1 UID: 0 PID: 79872 Comm: LIBSENSORS Not tainted 6.12.96+rpt-rpi-2712 #1
         pc : queued_spin_lock_slowpath+0x78/0x448
         lr : _raw_spin_lock+0x64/0x78
         Call trace:
          queued_spin_lock_slowpath
          _raw_spin_lock
          rp1_adc_read+0x34/0x108 [rp1_adc]
          rp1_adc_show+0x44/0xc0 [rp1_adc]
          dev_attr_show
          sysfs_kf_seq_show
          kernfs_seq_show / seq_read_iter / kernfs_fop_read_iter / vfs_read / ksys_read
  • The stall was detected 21 s after the link drop (the default RCU stall timeout), which puts the start of the stall at about the time of the link drop.
  • LIBSENSORS is presumably the Netdata sensors collector reading /sys/class/hwmon/hwmon1 (rp1_adc). I could not confirm the thread name after the reboot.
  • The lock holder is unknown. My hypothesis is that another context was blocked in an MMIO access to an unresponsive RP1, but I have not verified this.
  • The system stayed wedged: the same stall backtrace was still being printed at 06:14, with timestamps stretching apart (minutes between consecutive lines of one backtrace), and a bash watcher loop (ethtool/ip/ping every 15 s) wrote nothing after 02:44:11. Rebooted manually at ~06:20.
  • No AER/PCIe error lines were seen in the previous boot's kernel log around the event (only the normal boot-time messages; I did not search exhaustively).

What I am asking

  1. Is this a known RP1 hang (clock/IRQ/PCIe) with a fix in a newer kernel or firmware? Is there a debug parameter or trace that would show whether RP1 stops responding on PCIe when this happens?
  2. Is the rp1_adc hwmon read path expected to spin on a lock indefinitely when RP1 is unresponsive? If so, can polling hwmon during an RP1 hang take the whole system down?
  3. Should the link-down path in macb on RP1 reset anything else? Garbage statistics right after "Link is Down" suggest the GEM registers are already unreadable at that point.
  4. The link negotiating only 100 Mbps on a gigabit-capable port may point to a marginal cable or a 100 Mbps-only switch port. Could a flapping physical link be what triggers the wedge?

What has been tried

  • Re-seated the cable, ip link toggle, driver unbind/rebind (all failed, see above); reboot fixes it
  • Kernel/firmware update 6.12.96 to 6.12.109 and raspi-firmware 1.20260907 to 1.20260915 (result pending)
  • Enabled the 15 s hardware watchdog via systemd and an automatic reboot after 5 min of unhealthy end0, as a workaround

Attachments available

  • /var/log/end0-watch.log: full ethtool/ethtool -S/ip -s -d link/nmcli/dmesg dumps at each state change and every 5 min, covering both incidents
  • journalctl -b -1 -k for the incident-2 boot (contains the full RCU stall traces)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions