Booting badc kernels on real hardware: the GPD MicroPC

The VM lanes prove a kernel boots under emulated devices. This box is the other half: a physical machine, reached over a real RS-232 line, that runs the same packages.

What the machine is

   
Model GPD MicroPC, BIOS 4.15
CPU Celeron N4100 (Gemini Lake, Goldmont Plus), 4 cores
Firmware UEFI, Secure Boot disabled
Storage SATA SSD (BIWIN SSD, 119 GB) – no NVMe
Serial 16550A at I/O 0x2f8, IRQ 3, exposed as the DE-9 port
Distribution Fedora 44, kernel 7.1.10-200.fc44

The box is reached over ssh with key authentication and password-less sudo. Its address is site-specific and deliberately not recorded here.

Three of those decide what this box can and cannot test.

No AVX. Goldmont Plus carries SSE4.2 and AES-NI and stops there. The kernel’s AVX2/AVX-512 RAID6 and crypto paths compile but do not execute here, so the inline-asm work they cover still needs the x86_64 Linux box. It also means the EVEX encoding gap is not reachable at runtime on this machine.

Below the baseline. badc assumes x86-64-v3 (native-compilation.md); this CPU stops at x86-64-v2, with none of AVX2, FMA3, BMI1 or BMI2. Its kernels boot because -mno-sse keeps FMA3 out of kernel objects and badc’s integer lowering uses nothing above x86-64-v2. A baseline instruction in that lowering would break the boot here first: the VEX-encoded ones fault, and LZCNT / TZCNT execute as BSR / BSF.

SATA, not NVMe. A boot here exercises ahci/libata, not the nvme path the emulated lanes drive. The two are complementary rather than redundant: the module-autoload defect that hid behind NVMe was a packaging property, and this box reaches it through a different bus.

A battery. Cutting mains power does not reset a laptop, so a networked plug buys nothing. The chipset watchdog is the only unattended recovery the machine has.

The serial port is ttyS1, not ttyS0

Worth stating plainly, because the obvious guess is wrong and produces a console that is silent rather than broken. The board enumerates 32 ttyS* nodes; only two are real:

00:01: ttyS1 at I/O 0x2f8 (irq = 3, base_baud = 115200) is a 16550A
dw-apb-uart.8: ttyS4 at MMIO 0xa1324000 (irq = 4) is a 16550A

ttyS0 at 0x3f8 reports type=0 – no hardware behind it. The dw-apb-uart nodes are the SoC’s own LPSS UARTs, wired to internal peripherals. The physical DE-9 is the legacy-style 16550A at 0x2f8, i.e. COM2, which is why GRUB needs --unit=1 and the kernel needs console=ttyS1.

Confirm on any similar board before trusting a console setting:

sudo dmesg | grep -iE 'ttyS|8250|dw-apb'
for n in 0 1 2 3; do printf 'ttyS%s type=%s port=%s\n' "$n" \
  "$(cat /sys/class/tty/ttyS$n/type)" "$(cat /sys/class/tty/ttyS$n/port)"; done

A type of 4 is a detected 16550A; 0 means the node exists but nothing answers.

Configuration applied

Boot: console, early console, and a bounded panic

/etc/default/grub:

GRUB_TERMINAL="serial console"
GRUB_SERIAL_COMMAND="serial --speed=115200 --unit=1 --word=8 --parity=no --stop=1"
GRUB_CMDLINE_LINUX="console=tty0 console=ttyS1,115200n8 earlycon=uart8250,io,0x2f8,115200n8 panic=30"

then, because Fedora keeps each entry’s command line in its own BLS snippet rather than deriving it from grub.cfg:

grubby --update-kernel=ALL --args="console=tty0 console=ttyS1,115200n8 earlycon=uart8250,io,0x2f8,115200n8 panic=30"
grubby --update-kernel=ALL --remove-args="rhgb quiet"
grub2-mkconfig -o /boot/grub2/grub.cfg

/etc/kernel/cmdline is written from the resulting default entry so a kernel installed later inherits the same arguments – and it must carry root=, which is easy to lose:

sudo grubby --info=DEFAULT          # root= is printed on its own line,
                                    # NOT inside args="..."

grubby --info prints the root device in its own root= field and leaves it out of args=, so an /etc/kernel/cmdline derived from args alone is missing it. The entry kernel-install then writes for a newly installed kernel has no root device; the kernel reaches the initramfs, systemd-gpt-auto-generator looks for a root partition, dev-gpt-auto-root.device times out after 45 s, sysroot.mount fails and the boot parks in an emergency shell. Nothing complains at install time – the installed kernel looks fine and the running one is unaffected. That is how this box was first stranded, and the entry had to be repaired with grubby --update-kernel=... --args="root=UUID=...". Check the file before installing any kernel:

grep -o 'root=[^ ]*' /etc/kernel/cmdline || echo 'MISSING: kernel-install \
  will write an entry with no root device'

Installing a kernel here also silently takes the standing default: rpm -i runs kernel-install, which writes the new entry and points saved_entry at it. Re-assert the fallback after any install:

sudo grubby --set-default /boot/vmlinuz-<the distro kernel>

Each piece earns its place:

A password-less root shell on the port

/etc/systemd/system/serial-getty@ttyS1.service.d/autologin.conf
[Service]
ExecStart=
ExecStart=-/sbin/agetty -o '-p -f -- \\u' --autologin root --keep-baud 115200,57600,38400,9600 ttyS1 $TERM

Debugging a kernel that reaches userspace but misbehaves means running commands at the moment it happens; a login prompt in that situation costs time and sometimes the evidence. The machine is a lab target on a private network with Secure Boot off – it holds no secret that a password would protect.

The -f inside -o is required, and its absence is not obvious: --autologin root alone still leaves login authenticating, so the port answers with a password prompt and the journal records FAILED LOGIN 1 FROM ttyS1 FOR root. -o is the option string handed to login, and -f is what makes it accept the name without a password.

Watchdog: the only unattended recovery

echo iTCO_wdt > /etc/modules-load.d/watchdog.conf
# /etc/systemd/system.conf.d/watchdog.conf
[Manager]
RuntimeWatchdogSec=60
RebootWatchdogSec=2min

The firmware actually provides an ACPI wdat_wdt, which systemd adopts; wdctl reports a 60 s timeout with time left, so it is armed and being petted. A kernel that wedges after systemd starts resets itself within a minute.

Before the real root is mounted nothing pets it, and the case is not hypothetical. A kernel whose boot entry lost its root= reaches the initramfs, fails to mount the root filesystem, and parks in an emergency shell; the watchdog configuration lives on the filesystem that was not mounted, so nothing resets the box.

Magic SysRq is what ends that, and it is this machine’s only remote reset – the battery makes a switched mains outlet useless. A BREAK on the line followed by a command character reaches the kernel, so s, u, b syncs, remounts read-only and reboots:

kernel.sysrq = 1        # persistent; Fedora's default of 16 permits sync alone

tcsendbreak on the descriptor that reads the port sends the BREAK, which is what the harness does after the watchdog has had its chance. It is best-effort: it needs a kernel still servicing interrupts.

A locked root account takes away the other half of the answer. sulogin refuses a console it cannot authenticate on – Cannot open access to console, the root account is locked – which on a box whose only console is a serial line leaves no way in at all:

# /etc/systemd/system/emergency.service.d/override.conf, and the same for
# rescue.service
[Service]
Environment=SYSTEMD_SULOGIN_FORCE=1

Suspend, disabled at every layer that can ask for it

A desktop session on the target will put it to sleep mid-run. Observed on this box as a broadcast from the greeter – The system will suspend now! – which ends the ssh connection and silences the console, and is indistinguishable from a kernel that hung.

systemctl mask sleep.target suspend.target hibernate.target \
  hybrid-sleep.target suspend-then-hibernate.target
/etc/systemd/logind.conf.d/no-sleep.conf
[Login]
HandleLidSwitch=ignore
HandleLidSwitchDocked=ignore
HandleLidSwitchExternalPower=ignore
HandleSuspendKey=ignore
IdleAction=ignore

and the greeter’s own policy, which is what asked here:

sudo -u gdm dbus-run-session -- gsettings set \
  org.gnome.settings-daemon.plugins.power sleep-inactive-ac-type nothing
sudo -u gdm dbus-run-session -- gsettings set \
  org.gnome.settings-daemon.plugins.power sleep-inactive-battery-type nothing
sudo -u gdm dbus-run-session -- gsettings set \
  org.gnome.desktop.session idle-delay 0

Masking the targets is the layer that actually holds: whatever asks – greeter, logind idle, the power button – the request fails rather than being honoured. systemctl suspend now answers Call to Suspend failed: Access denied. The lid settings matter separately, because the machine is a clamshell that will sit closed on a bench; without them, closing it ends the run.

No desktop

systemctl set-default multi-user.target
systemctl isolate multi-user.target        # applies without a reboot

The greeter is what asked to suspend, and a desktop session contributes nothing to a kernel boot test while adding daemons, a compositor and a power policy that can each act on the machine mid-run. Removing it also returns about a gigabyte: the box now sits at 745 MB of 7.7 GB. The masks above still hold whatever runs on top; this removes the layer that kept asking.

A shell in emergency mode, and a way back from a wedge

Both of these were added after a failed boot left the machine unreachable with nothing to do but hold the power button.

# /etc/systemd/system/{emergency,rescue}.service.d/sulogin-force.conf
[Service]
Environment=SYSTEMD_SULOGIN_FORCE=1
# /etc/sysctl.d/99-sysrq.conf
kernel.sysrq = 1

The autologin getty covers a boot that reaches userspace; SYSTEMD_SULOGIN_FORCE=1 covers the emergency and rescue shells a boot that fails drops to, past the locked root account above.

kernel.sysrq=1 makes a serial BREAK followed by a key reach the kernel, so a wedged box can be synced and reset over the wire (BREAK then s, then b) – including where the systemd watchdog cannot help, with systemd running but stuck at a prompt.

What a failed boot leaves behind

Nothing on disk. A boot that ends in emergency mode does not get far enough to flush the journal, so journalctl -b -1 has no record of it – verified after exactly that failure. The serial console is the only evidence, which means the capture has to be running before the reboot is issued and stay open across it. Opening the port afterwards catches whatever is still in flight and nothing that came before.

One case now leaves more than that. With nmi_watchdog=panic softlockup_panic=1 on the badc entry, a lockup the NMI watchdog detects panics rather than sitting there, so the trace reaches the console and, where hwprep.py arm has enabled pstore, survives the reboot. A hang the detector cannot catch is unchanged: the console holds whatever was printed, and the chipset watchdog is what ends it.

Booting a badc kernel

Do not make one the default. Install it, select it for exactly one boot, and let any failure fall back:

sudo dnf install -y ./kernel-7.1.10-*.x86_64.rpm    # or rpm -i
sudo grubby --info=ALL | grep -E '^(index|title)'   # find its index
sudo grub2-reboot <index>                            # one shot only
sudo systemctl reboot

GRUB_DEFAULT=saved keeps the distro kernel as the standing choice, so a kernel that hangs is one power-button press away from a working system, and an unattended failure that trips the watchdog comes back on the distro kernel by itself.

Installing a kernel moves the standing default to it, as above, which removes the fallback the one-shot scheme depends on. Put it back before rebooting:

sudo grubby --set-default /boot/vmlinuz-7.1.10-200.fc44.x86_64

rpm -i refuses a kernel whose version-release orders below the one the box already runs – the pinned release built as -1 against Fedora’s -200.fc44 – with package kernel-7.1.10-200.fc44.x86_64 (which is newer than ...) is already installed. Kernels are install-only, so the version ordering is not meaningful here; --oldpackage is what gets past it, and dnf install applies the same semantics on its own.

The harness lane

packages.py --phases hw runs that sequence unattended, with the same probes, dmesg scanners and exercise stage the qemu lanes use. The README documents the phase; what it needs from this box is what the sections above configure:

python3 demos/linux/packages.py --arch x86_64 --distro fedora --phases hw \
    --release <kernel release> --package <kernel rpm> \
    --hw-host <host> --hw-serial /dev/cu.usbserial-XXXX \
    --workdir <scratch> --report hw-x86_64.json

Four of this box’s properties are the lane’s load-bearing assumptions:

The lane leaves the machine on the standing default and clears any pending one-shot selection on every exit path, including the failing ones. Like the qemu lanes it turns on core capture, which writes 99-badc-* drop-ins under /etc/sysctl.d, /etc/security/limits.d and /etc/systemd/system.conf.d and sets kernel.core_pattern to a file pattern; those persist on the machine.

Watching from the Mac

The DE-9 goes to a USB adapter on the Mac. Use the cu.* node, not tty.* – the latter blocks waiting for carrier detect:

ls /dev/cu.usbserial-* /dev/cu.SLAB_USBtoUART 2>/dev/null
picocom -b 115200 /dev/cu.usbserial-XXXX      # interactive; C-a C-x to quit

Each open() of the port resets its termios on macOS, so setting the speed with stty -f in one command and reading in the next gets the default rate, not 115200 – the line then delivers a few bytes of plausible-looking garbage rather than silence, which reads like a wiring fault and is not one. Set the speed in the same descriptor that does the reading. picocom does this; so does the harness, using termios from the standard library rather than a pyserial dependency, which also keeps the capture identical on the Linux lanes.

FTDI and CP210x adapters work with the drivers macOS ships; CH340 clones usually need a kext.

Still open

Undoing all of it

Every change above is reversible, and none of it touches the distribution kernel or the bootloader binaries. /etc/default/grub was backed up before the first edit.

# 1. Boot arguments and the GRUB console
sudo cp /etc/default/grub.badc-backup /etc/default/grub
sudo grubby --update-kernel=ALL --remove-args=\
"console=tty0 console=ttyS1,115200n8 earlycon=uart8250,io,0x2f8,115200n8 panic=30"
sudo grubby --update-kernel=ALL --args="rhgb quiet"
sudo rm -f /etc/kernel/cmdline          # regenerated on the next kernel install
sudo grub2-mkconfig -o /boot/grub2/grub.cfg

# 2. The serial root shell
sudo systemctl disable --now serial-getty@ttyS1.service
sudo rm -rf /etc/systemd/system/serial-getty@ttyS1.service.d

# 3. The watchdog
sudo rm -f /etc/modules-load.d/watchdog.conf /etc/systemd/system.conf.d/watchdog.conf
sudo systemctl daemon-reexec

# 4. Sleep
sudo systemctl unmask sleep.target suspend.target hibernate.target \
  hybrid-sleep.target suspend-then-hibernate.target
sudo rm -f /etc/systemd/logind.conf.d/no-sleep.conf
sudo systemctl restart systemd-logind
for k in sleep-inactive-ac-type sleep-inactive-battery-type; do
  sudo -u gdm dbus-run-session -- gsettings reset \
    org.gnome.settings-daemon.plugins.power $k
done
sudo -u gdm dbus-run-session -- gsettings reset org.gnome.desktop.session idle-delay

# 5. The desktop
sudo systemctl set-default graphical.target
sudo systemctl isolate graphical.target

# 6. Core capture, if the lane's probes ran
sudo rm -f /etc/sysctl.d/99-badc-cores.conf /etc/security/limits.d/99-badc-core.conf \
           /etc/systemd/system.conf.d/99-badc-core.conf
sudo rm -rf /var/crash
sudo systemctl daemon-reexec

# 7. Rescue-shell access and sysrq
sudo rm -rf /etc/systemd/system/emergency.service.d/sulogin-force.conf \
            /etc/systemd/system/rescue.service.d/sulogin-force.conf
sudo rm -f /etc/sysctl.d/99-sysrq.conf
sudo systemctl daemon-reload

One caveat on reversing the boot arguments: grubby --remove-args edits the BLS entries that exist at that moment, so a kernel installed while the serial configuration was in place keeps the arguments until it is removed or its entry is edited too. grubby --info=ALL shows what each entry currently carries.

^ To the top