Mysterious boot hang on Graviton: mitigated but never root-caused
I was getting an arm64 build of our Bottlerocket variant ready for production on Graviton EC2 instances. Load tests kept failing at a low rate. Somewhere between 0.3% and 1% of instances froze a few seconds into boot and never came back. The x86_64 build of the same variant has never done this.
I never found the root cause. After a few weeks and many rounds of experiments, two kernel parameter changes took the hang rate to zero: removing quiet and adding pci=realloc. This post walks through what the hang looked like, what got ruled out, and the experiments that led to the mitigation. If you have seen something similar, I would like to hear about it.
What a hang looks like
A hung instance has no journal and no logs shipped anywhere, because nothing got far enough to ship them. The hang happens very early in boot. The EC2 console still reports the instance as running. The only evidence of the hang is the EC2 serial console output, which ends like this:
1Welcome to Bottlerocket OS 1.0.0 (<variant>)!
2
3[ OK ] Created slice Slice /system/modprobe.
4[ OK ] Created slice Slice /system/systemd-networkd-wait-online.
5[ OK ] Created slice Slice /user.
6 Expecting device /dev/disk/by-partlabel/BOTTLEROCKET-DATA...
7 Expecting device /dev/disk/by-partlabel/BOTTLEROCKET-PRIVATE...
8[ OK ] Reached target Path Units.
9[ OK ] Reached target Slice Units.
10[ OK ] Reached target Swaps.
11[ OK ] Listening on Journal Audit Socket.
12[ OK ] Listening on Journal Socket (/dev/log).
13 2025-10-16T18:02:09+00:00
The last line is not from the Linux boot. It is from EC2 itself, which appends a timestamp to the console output when an instance is terminated; a healthy Amazon Linux instance gets the same line on termination. So the console says boot stopped after "Listening on Journal Socket", and about eight minutes later the instance was terminated, in this case by the control plane, which garbage-collects unhealthy instances.
There are also no shutdown messages before the timestamp. On Bottlerocket, the ACPI power button event EC2 sends on stop or terminate is handled by systemd-logind, which asks PID 1 to power off over D-Bus. If boot never got to "Started User Login Management", or if systemd or dbus-broker are stuck, the event is ignored.
Every hang is early, but each stops at a different place
The first theory was a systemd deadlock, some ordering problem in our units. The stop points did not support it. Here is the last line of the serial console output from five hung boots:
1[ OK ] Listening on Journal Socket (/dev/log).
2[ OK ] Mounted CNI Plugin Directory (/opt/cni).
3[ OK ] Finished Enable SELinux permissive labels.
4[ OK ] Started Network Configuration.
5[ OK ] Mounted Containerd content directory (/var/lib/containerd).
Some stopped before the journal was up, some after multi-user.target, some halfway through a burst of containerd plugin-loading messages. A deadlock between units would likely hang the boot at the same place each time. This looked more like the whole machine stopping at a random moment inside a small window. Across 85 hung boots whose console output has kernel timestamps, the last timestamp has a median of 2.5 s:
| Last kernel timestamp | |
|---|---|
| p10 | 1.5 s |
| median | 2.5 s |
| p90 | 4.7 s |
| max | 6.2 s |
Ruling things out
Is it systemd or the journal?
If journald were wedged, the console would go silent while the system kept running. To take systemd and the journal out of the picture, I added a unit that writes straight to /dev/console, with no dependencies, so it starts as early as systemd can start anything:
1[Unit]
2Description=Echo Hello
3DefaultDependencies=no
4
5[Service]
6Type=oneshot
7RemainAfterExit=true
8ExecStart=/bin/bash -c 'for i in {1..20}; do echo "Echo hello attempt $i" > /dev/console; sleep 2; done'
9
10[Install]
11WantedBy=preconfigured.target
A healthy boot prints all 20 lines over 40 seconds. In 12 hangs out of 1,500 boots, the loop got to attempt 1 in eight of them and attempt 2 in four. A bash process doing sleep and write in a loop is about as independent of systemd as a process can be. When it stops too, user space as a whole has stopped getting scheduled, or the kernel can no longer write to the console. Either way the problem is below systemd.
Is it a new kernel?
A kernel regression was the obvious next suspect. So I tried three kernels available at that time:
| Kernel | Hangs |
|---|---|
| 6.1, current kernel kit | 32 / 5,000 (0.64%) |
| 6.1, older kernel kit predating the suspect releases | 13 / 4,995 (0.26%) |
| 6.12 | 8 / 3,000 (0.27%), 23 / 5,000 (0.46%) |
All three hang. The rate moves around between runs, but nothing gets close to zero. So it is not a recent regression.
Is it PCI handling?
Our instances get a second ENI hot-attached a few seconds after launch. On EC2 Nitro, an ENI shows up to the guest as a PCIe device, and our variant has a small unit that rescans the PCI bus to make sure the new device is picked up. Rescanning a bus over and over during early boot is the kind of thing that could upset a driver. So I removed the rescan unit, but there were still 5 hangs out of 1,500.
What I could not remove was the ENI attach itself, the obvious difference between our boot and a vanilla Bottlerocket boot, which doesn't hang. A colleague built a reproducer closer to our setup: launch instances whose user data powers them off at boot, attach an ENI five seconds after launch, and flag any instance that never reaches stopped. It found a handful of m7g.large instances that stopped responding entirely. EC2 reported them as running, they ignored the power-off, and the serial console dropped the connection. That looks like our boot hang, which points at PCI hot-plug as a likely culprit.
A freeze, not a crash
Two observations changed how I think about the hang.
First, one instance came back. Its console has the journal forwarded, so it carries kernel timestamps, and on every other boot of that image the network configuration step runs at about 2.3 to 2.5 seconds. On this one:
1[ OK ] Started Network Configuration.
2[ OK ] Reached target Network.
3[ 603.448833] netdog[955]: Failed to write primary interface to '/var/lib/netdog/primary_interface': No such file or directory (os error 2)
4[ 603.451499] sysctl[956]: kernel.kexec_load_disabled = 1
5...
6[*** ] A start job is running for /dev/dis…EROCKET-DATA (1min 18s / 1min 30s)
The kernel clock jumped about 600 seconds between two steps that normally run back to back. The instance froze for ten minutes, resumed as if nothing happened, and then systemd timed out waiting for the data partition. That last part explained an older symptom: once in a while a boot failed with "timed out waiting for device" on the data partition. That was the same hang, just short enough that the instance woke up before it was terminated.
Second, the watchdogs fired. A Bottlerocket maintainer suggested turning a silent hang into a loud one:
1# kernel command line
2modules_load=softdog nmi_watchdog=panic softlockup_panic=1 hung_task_panic=1 oops=panic
3
4# systemd-system.conf
5RuntimeWatchdogSec=10
systemd pets the watchdog device every few seconds. If PID 1 stops doing that, softdog, a timer inside the kernel, panics the machine. The first run with this setup still produced 7 silent hangs out of 1,500. After softdog was built into the kernel instead of loaded as a module, a later 1,500-instance run had two instances panic from softdog instead of hanging silently. So the kernel's timers kept firing while user space stopped making progress. Whatever this is, it is not the whole VM being paused; it looks more like a lockup that leaves some CPUs or some interrupts alive.
Mysterious mitigation
While chasing the PCI angle I added pci=realloc. In the same series of runs, I also removed quiet from the kernel parameters to let the kernel print everything to the console. Then each change was tested on its own, and together, 5,000 boots each:
| Change | Hangs |
|---|---|
| none (baseline) | 0.3% to 1% across earlier runs |
add pci=realloc |
2 / 2,000 (0.1%), 5 / 5,000 (0.1%) |
remove quiet |
7 / 5,000 (0.14%) |
add pci=realloc, remove quiet |
0 / 1,500, 0 / 5,000 |
Each change on its own cut the rate by about 5x. Together, zero out of 6,500. Is zero real or luck? If the combination were only as good as one change alone, about 0.12%, you would expect 6 hangs in 5,000 boots, and the chance of seeing none is e^-6, about 0.25%. Against the baseline it is far smaller. The reverse holds too: with both changes in place plus the softdog setup, adding quiet back produced the two softdog panics mentioned above. Taking the change away brought the hang back.
I shipped both to the arm64 build, watched a couple of weeks of production traffic on it, and closed the investigation.
Why would these two help?
This is the unsatisfying part. Here is what each parameter does:
quietsets the kernel console log level so that only error and higher messages reach the console. systemd also reads it and turns down its own status output. Removing it means every kernel message during boot goes out the serial console at 115200 baud, about 11 KB/s. On 6.1, printing to a legacy console happens synchronously under the console lock in whichever context called printk.pci=realloc, per the kernel documentation, makes the kernel reallocate PCI bridge resources "if allocations done by BIOS are too small to accommodate resources required by all child devices." That matters for hot-plugged devices, which need to fit into a bridge window the firmware sized before they existed.
My best guess, and it is only a guess: there is a race in early boot on arm64 that a PCIe hot-plug event can trigger. pci=realloc changes how bridge windows and BARs get assigned for the device that shows up a few seconds into boot. Removing quiet changes the timing of nearly everything in early boot, because every printk now waits on a slow serial port. Either one shrinks the window, and both together close it, at least at the rate 5,000-boot runs can measure. That kind of bug is the classic heisenbug: turn on more logging and it goes away.
Why only arm64? I do not know. The arm64 and x86_64 boot paths differ in the interrupt controller, the PCIe host bridge code, the firmware tables, and the memory model, so there is no shortage of candidates. This is also not the first boot-timing problem I have seen on Graviton but not on x86_64; Apply sysctl before systemd was another one, and that one was ours.
Shipping a kernel parameter for one architecture
The hang only happens on Graviton, so the fix should apply to the arm64 build only. However, twoliter, Bottlerocket's build tool, has no way to set kernel parameters per architecture. I filed twoliter#587 for it; it is still open. In the meantime, the build edits the variant's Cargo.toml with a small Go program, modify-variant-kernel-parameters.go, only for aarch64:
1if [[ "${ARCH}" == "aarch64" ]]; then
2 ./modify-variant-kernel-parameters -variant "${VARIANT}" -action add -kernel-parameter "pci=realloc"
3 ./modify-variant-kernel-parameters -variant "${VARIANT}" -action remove -kernel-parameter "quiet"
4fi
5cargo make -e VARIANT="${VARIANT}" -e ARCH="${ARCH}" ...
The program parses the TOML to read the current kernel-parameters array, then rewrites only that array in place rather than re-encoding the whole file, so the diff touches the kernel-parameters array and nothing else.
What I would do sooner next time
-
Turn on the watchdogs on day one. softdog plus the panic parameters converts "the instance is gone and there is no trace" into a panic with a stack and a reboot. I only added them three weeks in, after several rounds of guessing.
-
Write to /dev/console from a dependency-free unit before suspecting systemd. A bash loop said more in one run than days of rearranging unit dependencies.
-
Decide up front how many boots an experiment needs. At a 0.5% base rate, a 1,000-boot run that comes back clean is weak evidence, and it produces too few hangs to show the different ways the bug presents, in this case the different stop points. 5,000 boots per arm gives both: enough hangs to compare symptoms while debugging, and enough clean boots to trust the fix.
-
Take the mitigation, even without the root cause. It is uncomfortable to ship two kernel parameters with only a guess as to why they work. But engineers have to ship. Document the analysis and the data points, and move on.