PPS capture latency: polling vs interrupt, same pulse

BPI-R64 (Cortex-A53 @ 1.35GHz, 12.5MHz arch counter), u-blox GNSS 1Hz timepulse. A differential measurement: for every pulse, CPU1 busy-polls the GPIO data register directly (~480ns per MMIO read, counter reads bracketing the detecting read) while CPU0 takes the EINT interrupt with the arch counter value captured at exception entry (entry.S) and again in the pps-gpio hardirq handler. All three captures are raw counter values for the same physical edge, so the pulse's own timing and the pin's input path are common-mode; the deltas isolate the capture mechanisms. The clock was free-running (no hardpps, no NTP) — nothing disciplined, nothing to perturb.

Run on a live Fedora 44 system (full userspace), not an idle test bench. Load segments: a dnf upgrade (network, mmc I/O, rpm scriptlets) and loops of dracut --force (compression pipeline). cpufreq governor as marked; this SoC has no cpuidle driver (idle = WFI at the governor's frequency).

Delay after the polled edge (ns)

segmentcapturemeanp50p95 p99p100σ
idle / performance
118 pulses, 0 excluded
entry.S +170 +160 +400 +400 +400 133
hardirq handler +7361 +7120 +8880 +9280 +10080 796
idle / schedutil
60 pulses, 0 excluded
entry.S +585 +640 +800 +880 +960 180
hardirq handler +13652 +13360 +17440 +18080 +18640 1458
idle / schedutil (post-load)
700 pulses, 3 excluded
entry.S +584 +560 +800 +960 +7520 360
hardirq handler +14441 +13600 +18720 +19840 +35680 2538
load (dracut) / performance
1599 pulses, 17 excluded
entry.S +413 +240 +880 +2080 +66560 2094
hardirq handler +11107 +8880 +20080 +25280 +78160 5010
load (dracut) / schedutil
1722 pulses, 11 excluded
entry.S +311 +240 +800 +1520 +14160 520
hardirq handler +11423 +9440 +21040 +24080 +58000 4456
load (dnf) / schedutil
1086 pulses, 14 excluded
entry.S +833 +320 +1520 +7760 +174880 5810
hardirq handler +13079 +12720 +20640 +24480 +187360 7008

Exclusions: ~1% of load-segment pulses are excluded where the poll shows the entry stamp predating the pulse by more than 800ns. Those are pulses which arrived while CPU0 was already in an exception taken for an earlier interrupt (the timer tick): the interrupt is serviced from that exception without re-entry, so the entry.S stamp belongs to the tick, not the pulse. The error decays at the crystal-vs-GPS beat rate (~110 ticks/pulse at +8.6ppm) and recurs every ~467s (HZ=250), matching observation. A production entry.S backdate needs a sanity bound against the handler-time snapshot to reject these; the excluded pulses are listed in the run log.

Distributions

The entry.S body barely moves under load (p50 240–320ns) but grows a smooth heavy tail — there is no knee, so no natural threshold to clip at. The handler timestamp is microseconds at best and its whole body shifts with frequency and load.

Can filtering recover the entry.S tail?

The jitter of the polling method has a physical floor: the edge lands uniformly within one GPIO read period, σ = T/√12 ≈ 140ns for the 480ns direct-MMIO loop (~1.5× that via gpiolib). The loop period stretches ~10–15% under load (bus contention on the MMIO read, visible per-pulse as the bracket width) and with CPU frequency, but that is the poll's entire load sensitivity — it holds the CPU with interrupts off, so a pulse is either captured at floor quality or (if the wake loses the race) simply missed. Its jitter is its floor. The entry.S capture has no such floor — idle, its intrinsic jitter measures below the poll's quantisation — but under load it grows a heavy tail of delayed and mis-attributed stamps.

Method: a smoothed ideal pulse train, built retrospectively from the poll's captures only (leave-one-out local linear fit, ±128 pulses, worst-10% trimmed), plus a short local median of the poll residuals subtracted from both methods to remove crystal thermal wander (strongly autocorrelated, lag-1 r≈0.97 under load: frequency motion the servo tracks, not capture jitter — without this step it inflates both methods' σ several-fold and masks the real comparison). The reference inherits a few tens of ns of the poll's own quantisation, which lands on the entry residuals in quadrature: the construction slightly flatters the poll, not the irq path. Mis-attributed stamps (entry predating the polled edge by >800ns; the tick-shadow episodes above) are gated per-sample — they arrive in runs at the crystal-vs-tick beat, which no fixed-length median survives, but each is individually detectable in a real deployment from the entry→handler gap. A median-of-5 then eats the residual isolated delays:

segmentcaptureσp50 p95p99worst
idle / performance
118 pulses, 0 gated
poll (spin) 179 +11 +337 +360 375
entry.S raw 80 +175 +295 +360 372
entry.S, gated + median-of-5 49 +174 +262 +264 264
load (dracut) / schedutil
1722 pulses, 11 gated
poll (spin) 154 +2 +249 +347 597
entry.S raw 511 +239 +755 +1520 14273
entry.S, gated + median-of-5 123 +238 +507 +680 1051
load (dnf) / schedutil
1086 pulses, 14 gated
poll (spin) 155 -2 +252 +338 444
entry.S raw 5814 +316 +1613 +8442 175019
entry.S, gated + median-of-5 280 +311 +740 +969 3632

The poll measures σ ≈ 135–155ns in every segment — load-invariant at its quantisation floor, as the structural argument predicts. The entry.S capture spans σ ≈ 50–70ns (idle, raw — genuinely below the poll's floor) through ≈120ns (moderate load, filtered) to ≈280ns (irq-heavy load, filtered — the dnf tail is dense enough that some of it survives a 5-tap median). So: better than polling when the system is quiet, comparable under moderate load, and ~2× worse under irq-heavy load even after gating and filtering — and it requires the per-sample attribution gate to be usable at all.

The run-structured mis-attribution class would disappear structurally if the counter were stamped per-interrupt at the GIC acknowledge (IAR) read rather than once per exception entry — the dispatch loop can service several interrupts from one exception, and the stamp's granularity is currently coarser than the attribution the hardware provides.

Method notes