Repository navigation
Terminal browsing - #6
Merged
Merged
Conversation
`runner` hardcoded `-bios /usr/share/OVMF/OVMF_CODE.fd`, a path Ubuntu 24.04 renamed. QEMU therefore refused to start, and "this machine has no emulator, every test costs a power cycle" hardened into a project-wide assumption. Both qemu-system-x86_64 and ovmf were installed the whole time. Search a candidate list instead, honour $NYX_OVMF, and pass $NYX_QEMU_ARGS through so a headless serial-capturing boot needs no edit to this file. On a missing firmware, say so and name the paths tried rather than dying in a way that looks like a missing emulator. Note for anyone reinstating this: the `OVMF_CODE_4M.fd` variants are pflash-only and produce ZERO serial output with `-bios`, which looks identical to "QEMU does not work here". The combined /usr/share/ovmf/OVMF.fd is the one that boots. Co-Authored-By: Claude <noreply@anthropic.com>
…ad here Three faults compounding, none of which anything could observe. `init_timer` programmed a hardcoded APIC count of 0xA000 with a comment claiming "a fast 1ms tick rate". Nothing ever measured it. Measured on this laptop: the LAPIC runs off the 24 MHz core crystal, so 0xA000 at divide-16 is **27.3 ms**. That was the scheduling quantum, and it is why hover highlights lagged while the pointer stayed smooth -- the cursor comes off the mouse IRQ via the hardware cursor plane, but a highlight needs the window server scheduled, and it could not run more than ~37 times a second. `calibrate_tsc` had been silently failing since it was written: this machine's PIT channel 2 gate never asserts, so it fell back to a hardcoded 2000 MHz and said so through `serial_println!` -- on a laptop with no serial port. The tell was `sched` reporting TSC exactly 2000. That value also feeds iwlwifi's firmware delays and the SMP INIT/SIPI timing, so those were scaled by real_MHz/2000. Derive the TSC from CPUID instead, which needs no timer: leaf 0x15 (crystal x ratio) cross-checked against leaf 0x16 (base frequency -- an Intel invariant TSC ticks at exactly that). Intel leaves ECX=0 on client parts so the 24 MHz crystal is an ASSUMPTION, trusted only when 0x16 agrees within 10%. PIT remains the last resort. `TSC_SOURCE` publishes which path won, so "never calibrated" can no longer masquerade as a measurement. `init_timer` now computes its count from the calibration to hit a real 1 ms, and falls back to the historical constant if calibration could not run -- an uncalibratable machine behaves exactly as before. Measured after: tick period 1000 us, TSC 2496 MHz, 27.3ms -> 1.0ms quantum. Co-Authored-By: Claude <noreply@anthropic.com>
The largest source of interactive latency on this machine was the GPU, and it took three failed attempts to find because every wait in the driver was bounded by an ITERATION COUNT. An iteration count means a different duration in every loop body -- 20,000,000 of `clflush + mfence + read` is ~23.5 ms here, 1,000,000 of a bare register poll is well under a millisecond -- so the numbers were mutually incomparable, said nothing about time, and could not be reasoned about from the source. Each fix patched a loop found by inference while the one actually stalling sat elsewhere. The tell every time was a dead-constant duration (23,541-23,549 us, a 9 us spread) that did not divide by whatever timeout had just been changed. The full inventory was nine: BLT ring-full, forcewake ack, reset-ready, GDRST pulse, the selftest fence, the RCS ring-full, and THREE 30M fence waits in pipeline.rs -- which is what `draw_scene` actually goes through, and so was the real 23.5 ms all along. All nine now use `SpinDeadline`, a TSC-based budget in microseconds. Re-grep before believing this is done; the fix is complete when the inventory is empty, not when one site looks right. What the measurement then showed: the RCS fence NEVER signals on this hardware. Every composite ran its wait to the end, returned EngineHang, and the compositor silently fell back to software -- after burning 23.5 ms with interrupts masked, ~4 times a second. Invisible for the life of the project because the timeout message goes to a serial port this laptop does not have. So: latch the path off after 8 consecutive failures rather than re-learning it every frame. Counted per COMPOSITE, not per fence wait -- `draw_scene` performs two waits, so a first attempt at counting them individually flapped 0-1-0-1 and could never reach the limit. These timeouts are deadlines for declaring hardware DEAD, never throughput knobs. `BLT_FENCE_TIMEOUT_US` briefly sat at 2 ms and abandoned real blits mid-flight, producing a half-drawn shifted screen at boot. Shortening a wait for work that WILL complete does not remove a stall, it corrupts output. GL keeps its own strike counter, cleared by `gl_init`. Sharing the compositor's latch stopped glcube ever starting: the desktop fails 8 composites within a second of boot and the shared latch then refused GL before it touched hardware. `gl_init` re-programs forcewake and MOCS -- exactly the state this engine is known to lose after RC6 -- so a fresh GL session can succeed where the compositor's cached state cannot, and must be allowed to try. Measured: 536 fell from ~4/s at 23.5 ms to exactly 8 occurrences total; 537 (text, gated once it surfaced as the next consumer of the same dead fence) from 34 to 4. Co-Authored-By: Claude <noreply@anthropic.com>
Hardware named `execve` as the longest interrupts-off window on the machine: 102,702 us inside a single syscall. The cause was the filesystem bridge. For EVERY 512-byte sector, `nyx_nvme_read_block` allocated `vec![0u8; 8192]`, zeroed it, used it, and freed it -- purely to hand the driver a 4 KiB-aligned slice the driver never needed, since `read_block` DMAs into its own aligned `DATA_BUF` and copies out. Reading a 1.4 MB binary meant ~2,900 of those: some 23 MB of memset to deliver 1.4 MB, all under the allocator's `without_interrupts`. Worse, each sector was a separate NVMe submit/doorbell/poll round trip. Since SYSCALL masks interrupts and the driver polls, that serialised ~2,900 device latencies with the timer dead. Static buffers instead of per-sector allocation, plus a single-entry 4 KiB chunk cache and `read_blocks` (cdw12 NLB is 0-BASED; capped at 8 because PRP1 covers one page and PRP2 is left zero). Nothing in this stack ever reads the namespace's LBA format -- 512 is assumed everywhere. It is almost certainly right (the GPT header is found at LBA 1, which could not happen at a 4096-byte LBA) but "almost certainly" is the wrong standard for a value that tells a controller how much to DMA: too large and it overruns DATA_BUF and corrupts the disk holding the system. So `verify_multiblock` proves it on the device at boot -- an 8-block read must equal eight 1-block reads -- and the fast path is gated on that. The write path invalidates the chunk unconditionally, including on failure paths. A stale chunk would serve pre-write bytes and silently corrupt the filesystem. Measured (QEMU, warm boot, same syscall): 390,716 us -> 63,974 us. On hardware 102,702 us -> 29,180 us. Verified on a fresh disk too, which exercises the full installer write path. Co-Authored-By: Claude <noreply@anthropic.com>
Nothing in this kernel could answer "how long is a tick", "how long did that task wait to run", or "how long was this core unable to take an interrupt". `CONTEXT_SWITCHES` was the entire budget: one global cumulative counter, exposed by syscall 523, read by nobody. Every performance claim about the scheduler was therefore unfalsifiable, and several turned out to be wrong. Per-CPU `SchedStats` in `PerCpu`: plain u64, no atomics (only the owning core writes, always at IF=0), `align(64)` against false sharing. Deliberately not more global atomics -- `UPTIME_MS` and `CONTEXT_SWITCHES` are bumped by every core on every tick and are themselves a measurable SMP cost. Interrupts-off time is measured WITHOUT instrumenting cli/sti. The APIC timer is periodic, so the TSC gap between consecutive timer interrupts IS one tick period unless the core could not be interrupted. One rdtsc per tick. Attribution took two attempts. `cur_syscall` alone cannot blame a syscall for the window it caused: SYSCALL masks interrupts and `sysretq` restores IF, so the suppressed tick only arrives after the syscall has returned and the marker is already cleared -- it reported "not in a syscall" for exactly the case it was built for. `max_sys_cycles`, timed INSIDE the call, cannot be fooled; the two now cross-check (390,828 us gap vs 390,716 us syscall, agreeing to 0.03%). A 16-entry ring of recent stalls, plus a cumulative per-syscall tally with means -- a single worst-case sample named only the boot-time execve, which says nothing about steady state, and a recurring cause has to be able to name itself.⚠️ `sys_hist` measures syscall WALL time, not masked time. A 1,000,202 us reading for syscall 525 is sleep() sleeping for a second. Cross-check the tick gap before reacting. `sched` / `sched hist` in apps/terminal surfaces it -- the first process visibility this system has had; there is no ps/top/uptime. F12 dumps the same report, from the THERMAL GOVERNOR rather than the keyboard ISR: `serial_println!` is a byte-at-a-time UART spin at IF=0 and would forge the very latency being reported. The report ends with the kernel build stamp, because a stale flash makes a fixed kernel and an unfixed one produce identical numbers -- which cost a debugging cycle here, exactly as acpi.rs already warned it would. All of it compiles out without the `sched_stats` feature. Co-Authored-By: Claude <noreply@anthropic.com>
Input. The keyboard and mouse ISRs woke EVERY task with a finite `wake_tsc`, because `Blocked` carried no reason and a waker had to guess. PS/2 emits one IRQ per byte, three or four per motion event, so a moving pointer ran a whole-task scan plus a full context switch plus a 1 KiB FPU save/restore per byte. It also broke sleep() system-wide: `sys_sleep_ms` treats a cleared `wake_tsc` as a legal early return, so while the mouse moved, every app's 16 ms frame sleep, wifiagent's 500 ms and init's 1000 ms all returned immediately. `WaitReason` lets a waker wake exactly what is waiting for the thing that happened. The ISRs now wake only `Input` waiters, and only reschedule when they actually woke something -- they used to tail-call the scheduler on every byte, which measured as 32% of all schedule() calls changing nothing during pointer motion.⚠️ The herd was accidentally MASKING the real problem: it dragged apps out of their frame sleep on every key release, which is why typing felt better than a 16 ms poll loop should allow. Removing it alone would have regressed input latency, so `ipc_recv` gains a deadline mode and `ipc_send` now wakes any IPC waiter -- previously it only woke one parked forever, so a receiver waiting with a timeout was never woken by the message it was waiting for. libs/gui blocks on that instead of `sleep(16)` + poll, which cost up to a full frame per event. SMP. `fork`/`clone`/`spawn_thread` pushed onto the CALLING core unconditionally, and since everything descends from init on core 0, that is where the whole system ran: measured over 42 s on 8 cores, cores 1-7 took ~42,000 timer interrupts EACH and switched task zero times. `place_task` picks by live load, biased to the caller's core. The old balancer (syscall 58 only) ranked by `tasks.len()`, which counts the idle task and every tombstone, so a core that had reaped processes looked permanently busier.⚠️ Handover goes through a per-core INBOX, never a direct push into another core's `tasks`. `ipc_send`/`futex`/`wait4`/`sysinfo` all iterate other cores' task lists; a remote push that reallocated the Vec would leave them following a freed pointer. `tasks` is now pre-reserved and slots are RECYCLED -- entries cannot simply be removed because `core_task_idx` holds an index into it. That also finally bounds the per-tick O(n) scans, which walked every process that had ever existed. `alloc_slot` hands a task BACK on a full core rather than dropping it; a dropped Process is a process that silently ceases to exist. Clock. `UPTIME_MS` was `fetch_add(1)` per tick on EVERY core, so its rate was cores/tick_period and changed as cores came online -- the per-core tick counts summed exactly to it. Now derived from the TSC: monotonic, identical on every core, and immune to ticks lost to interrupts-off windows. This had to land with the quantum change or a 27x faster tick would have made the clock 27x worse. syscall 503 submits under the GPU lock and then waits WITHOUT it, interrupts enabled. Those are real blits completing (~9.7 ms, ~2/s) so the wait cannot be shortened, but it need not be uninterruptible and need not block every other core's GPU syscall for the duration. Measured: quantum 27.3ms -> 1.0ms, clock 1.00x real time, scheduler work during pointer motion -46%, voluntary yields -81%. Co-Authored-By: Claude <noreply@anthropic.com>
`usb.rs` enumerated devices, reset ports, addressed them, configured interrupt endpoints, set boot protocol and SET_IDLE -- and then nothing ever collected a report. `poll_all_mice` was called from NOWHERE, and its only output was a `serial_println!` to a port this laptop does not have. All the hard work existed; what was missing was something to turn the handle. Three changes. Enumeration never recorded WHAT a device was. The interface descriptor's bInterfaceClass/bInterfaceProtocol were skipped entirely, so every report was parsed as a mouse -- a USB keyboard's modifier and keycode bytes were fed to the cursor as dx/dy, which is the only thing it could ever have done. Keyboard boot reports are now parsed and translated in `shell::handle_hid_key`, deliberately into the SAME `KEY_QUEUE` and the SAME Private-Use-Area encoding the 8042 path uses -- a USB keyboard and the built-in one must be indistinguishable above that line, or every app would need to know which keyboard a keystroke came from. Ctrl/Alt are dropped to match `HandleControl::Ignore`: Nyx has no modifier chords, and introducing them on one input device only would be worse than not having them. The boot protocol reports keys HELD rather than transitions, so presses are detected by diffing against the previous report; Super is an edge on the modifier byte, which has no usage ID of its own. A dedicated kernel task polls at 8 ms (the interval a full-speed HID interrupt endpoint is specified at, so faster would re-read the same report). NOT the timer ISR: a USB transfer is far too much work for a handler that runs with interrupts masked, and putting it there would recreate exactly the stalls the preceding commits removed. It `try_lock`s the controller, never blocks -- the USB syscalls take that mutex at IF=0, and blocking on it from a preemptible task is the preemption-boundary deadlock net/mod.rs documents.⚠️ The task is built from `new_idle_ap`'s RESERVED PID range, not `Process::new()`. The boot daemons hold the well-known low numbers and apps hardcode COMPOSITOR_PID=4 for window IPC; consuming an ordinary PID here would shift the compositor off 4 and break every app's ability to open a window. That regression has happened before, which is why the reserved range exists. `kernel_sleep_ms` moved to scheduler.rs and is now shared with the thermal governor -- two copies would drift, and the thermal one already lacked the `WaitReason` tag that keeps the input ISRs from waking it.⚠️ NOT verified end to end. Under QEMU's qemu-xhci the command ring never answers ("NoOp Command Failed"), so enumeration stops after port detection and no HID report is ever produced there. The mouse byte offsets are left exactly as found (b1/b2/b3, not the textbook b0/b1/b2) because that is the only part of this path with hardware evidence behind it. Real hardware is the test. Co-Authored-By: Claude <noreply@anthropic.com>
…or keys Two separate reasons the keyboard felt slow, neither of them the scheduler. The 8042's typematic rate was never programmed, so it sat at the power-on default -- which is the SLOWEST the hardware offers: 500 ms before a held key repeats, then ~10.9 characters per second. Arrowing through a file or holding backspace was limited by that and nothing else. Now set to 250 ms / 30.0 cps, the fastest the PS/2 protocol can express.⚠️ Sent to the data port directly, NOT via `write_mouse`. The 0xD4 prefix that helper adds targets the AUX device, so this would have gone to the trackpad -- which has no such command -- and left the keyboard exactly as it was. The window server polled for keystrokes on a timer. While UPTIME_MS ran ~5.5x fast its `sleep(2)` was really ~0.36 ms, so it happened to poll at ~2,700 Hz; making the clock honest turned that into a real 2 ms and therefore made every keystroke up to 2 ms later. Correcting the clock silently regressed input. Syscall 576 blocks on `WaitReason::Input` instead. That mechanism has existed since the input thundering herd was removed -- the ISRs have been waking only `Input` waiters this whole time and NOTHING ever registered as one. Now the keyboard IRQ wakes the shell directly: the poll interval leaves the latency path, and ~500 syscalls a second go with it.⚠️ Always with a timeout, never indefinitely. The shell IS the window server; a missed wake on a blocking call there stops the desktop rather than one app. On timeout it behaves exactly as the old sleep did. The key it collects is carried into the next pass rather than handled at the wait site, so `process_input` remains the single place that interprets a keystroke. Co-Authored-By: Claude <noreply@anthropic.com>
First of three drivers for the PNP0C50 precision touchpad. This one does no I/O -- it answers "where is the device and how do I talk to it", which nothing could answer before. None of it can be hardcoded, and that is the whole reason this phase exists. This DSDT declares the SAME touch-device slot on four separate I2C buses and patches its _HID and slave address at _INI from an NVS variable: the one slot becomes WCOM4831@0x0A, ALPS0000@0x2C, ELAN2097@0x10, NTRG0001@0x07, SYNA2393 or DLL077A depending on which panel the factory fitted, and _STA decides which of the four is real. `_CRS` is a Method whose result additionally depends on OSYS and SDM0, so even the resource template cannot be read statically. `acpi_find_i2c_hid` already walked to these devices and threw everything away -- it retrieved the full pathname and immediately freed it unread. It now extracts, per present device: slave address and bus speed (from the I2cSerialBusV2 in _CRS), the GPIO interrupt pin (from GpioInt), the HID descriptor register (from _DSM with the standard HID-I2C UUID, function 1), and the controller's PCI bus/device/function (by resolving the _CRS ResourceSource path and evaluating that controller's _ADR). Uses AcpiWalkResources, which was vendored, compiled and proven by the EC's port discovery but had exactly one caller. Fixed stack buffers throughout rather than ACPI_ALLOCATE_BUFFER: the results are bounded, and not allocating means not owning a free.⚠️ Exposed as `acpi probe 13`, NOT run at boot. The comment on PROBE is blunt about why -- "three boots died on the governor's automatic first pass and each guess at which call was responsible cost a power cycle" -- so every new ACPI call in this tree is opt-in first. Breadcrumbs 70/71/72 narrow a hang to _CRS, _DSM or neither.⚠️ Syscall 577 copies from the published cache and evaluates NOTHING. AML in a syscall is the preemption-boundary deadlock that wedged this machine on `panel` and `battery`: SYSCALL runs with IF=0, and AcpiEvaluateObject takes the interpreter mutex, allocates, and here can end in a firmware SMI. The governor (IF=1) evaluates; the syscall copies scalars.⚠️ The probe-step range check in the terminal is extended to 13. It carries a warning that step 7 was silently rejected for a while and its dump never ran -- same trap, one number later. An entry is only published when the _CRS walk actually found an I2C descriptor. A slave address of zero makes every other field meaningless, and reporting it would aim the driver at an address that does not exist. QEMU has no PNP0C50 device, so `touchpad` there correctly reports none; that proves the plumbing and nothing else. The numbers have to come from hardware. Co-Authored-By: Claude <noreply@anthropic.com>
… is a GSI Two fixes, both found by running the previous commit on real hardware. THE MOUSE REGRESSION. `9aba0a7` added the keyboard's typematic-rate handshake to the end of the 8042 init -- AFTER `write_mouse(0xF4)`. That command is "enable data reporting": the instant it is acked the mouse begins streaming 3-byte packets into the SAME output buffer the handshake reads from. And `wait_for_read` only tests status bit 0 (output buffer full), never bit 5 (AUX data), so it cannot tell a keyboard ACK from a mouse byte. The four blind reads swallowed mouse bytes, desynchronised the packet state machine, and left the pointer dead. Moved ahead of `0xF6`/`0xF4`, so nothing is in flight while it runs and both sequences are clean.⚠️ QEMU never reproduced this -- its 8042 emulation tolerates the interleaving and clicks kept working all the way through. The fix is reasoned from the ordering, not from a reproduction. THE INTERRUPT. The hardware probe returned a blank GPIO pin, which looked like "no interrupt" but is not. This device's `_CRS` is a Method returning EITHER a GpioInt OR a plain Interrupt, selected by OSYS and SDM0 -- and on this board it returns the plain Interrupt. The resource callback only understood SERIAL_BUS and GPIO, so an APIC GSI fell on the floor. Now captures EXTENDED_IRQ/IRQ as well, and `touchpad` says which kind it got. If it reports a GSI, the IOAPIC can route the touchpad exactly like the keyboard and mouse lines, and the GPIO controller driver -- the hardest of the three this project was scoped around -- is not needed at all. Hardware result from the previous commit, for the record: \_SB_.PCI0.I2C1.TPD0 on PCI 00:15.1, address 0x2c @ 400 kHz, HID descriptor register 0x0020, _STA 0xf. That is an ALPS0000 (BADR 0x2C in the DSDT template) and it is everything the I2C controller driver needs. Co-Authored-By: Claude <noreply@anthropic.com>
Two attempts, neither fixed it, and a dead pointer on a laptop is not worth a faster key repeat. mouse.rs is now behaviourally identical to `9aba0a7^`, the last state where the trackpad is known to have worked (the remaining diff is whitespace only). What was tried, and why it was wrong: The first attempt put the handshake AFTER `write_mouse(0xF4)` — "enable data reporting" — so the mouse was already streaming 3-byte packets into the same output buffer the handshake reads from. `wait_for_read` only tests status bit 0 (output buffer full), never bit 5 (AUX data), so it cannot tell a keyboard ACK from a mouse byte; the blind reads desynchronised the packet state machine. That explanation was at best incomplete: moving the handshake BEFORE `0xF6`/ `0xF4`, with nothing else in flight, did not bring the pointer back. So the real mechanism is still unidentified, and guessing a third time with the machine's only pointing device broken is not a reasonable thing to do.⚠️ QEMU reproduces NEITHER failure. Its 8042 tolerates the interleaving and clicks kept working through both broken builds, which is why this shipped twice. This cannot be developed here. Reintroducing it needs, at minimum: drain the output buffer first, VERIFY each response is 0xFA rather than reading blindly, and treat a missing ACK as "skip the whole thing" instead of continuing into the mouse sequence. The comment left in place records that. The blocking key read (syscall 576) is untouched and stays — it is a separate change, in software, and it is what actually removed the keystroke latency that the typematic rate was only a small part of. Co-Authored-By: Claude <noreply@anthropic.com>
The pointer died after 5db7eac, and reverting the typematic change did not bring it back, because that was never the cause. 5db7eac made the boot-time PNP0C50 scan share the full discovery callback, which evaluates the HID _DSM on every boot. On this firmware that call is not a query: TPD0._DSM: If (DRDY == 0 && Arg0 == HIDG) { DRDY = 1; EV5 (0, 0) } EV5 -> ECDV.EDPE: If (PMED == 0) { PMED = 1; EISC (0x81, 0x20, 0) } PMED is "PS/2 Mouse Emulation Disable": the first HIDG _DSM tells the EC an I2C-HID driver has arrived, and the EC moves the touchpad off the 8042 AUX port. With no I2C-HID driver yet, that left no pointer at all. - The boot scan counts devices again and evaluates nothing. - Discovery reads the HID descriptor register from the firmware's HID2 Name instead of _DSM. Same value (0x20), no side effects. - _DSM now belongs to the I2C-HID driver, at the moment it takes over. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every keystroke in the Command called mark_full, so each one was a full-screen recomposite plus a CPU scrim blend over every pixel. The scrim itself called alpha_blend per pixel: a bounds check and three integer divisions, about a million times a frame at 1280x800. - rebuild_command and Up/Down damage only the panel, old and new extent. Open, close and the 1 Hz heartbeat still repaint everything. - The constant-colour blend precomputes the colour's share of each channel and divides by 255 with the exact shift identity. Checked exhaustively against alpha_blend over every alpha and channel pair: 0 mismatches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
paint_text falls back to nyx_gui's CPU font whenever sys_gpu_draw_text fails (no Intel GPU, or the render engine latched off after hangs), but measure() kept using the GPU atlas, whose 20px light is wider. The Command's caret drifted further from the query with every letter, and right-aligned and centred labels were off the same way: "esc to dismiss" ran past the panel edge and "System nominal" overlapped its icon. paint_text now records which path drew the frame, and measure() uses that path's widths. The switch lags one frame. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Phase 2 of the I2C-HID touchpad. A polled Designware I2C master for the
LPSS controller ACPI discovery found (\_SB.PCI0.I2C1 = PCI 00:15.1 here):
PCI D0 via the PM capability (what the firmware's own _PS0 does), memory
+ bus master, BAR0 mapped, LPSS function released from reset, identity
checked against IC_COMP_TYPE, then master mode with the board's own SCL
timings. Every wait is a SpinDeadline in microseconds.
The timings come from the firmware's per-bus NVS Names (SSHn/SSLn/SSDn,
FMHn/FMLn/FMDn), read alongside the rest of discovery. Plain Names, no
side effects. Without them it falls back to standard mode computed for a
216 MHz clock, the fastest this family uses, so a wrong guess only runs
the bus slower.
The probe reads the 30-byte HID descriptor and checks wHIDDescLength and
bcdVersion. It runs in the usb-hid kernel task at IF=1, because bring-up
sleeps through the D3->D0 transition; syscall 578 only raises the request
and copies the published result. `touchpad` now ends with it, and
`touchpad i2c` runs it alone.
This does NOT hand the touchpad over from PS/2. That is the device's
_DSM, which nothing here calls, so the pointer is unaffected. An address
NACK would mean the part only answers on I2C after the handover, which is
worth knowing before Phase 3.
QEMU has no LPSS controller: verified there only as far as the request,
kernel-task, publish and copy round trip ("no ACPI data").
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware: `PCI 8086:06e9 PMCSR 0xb BAR0 0x0`. The Serial IO controller is in PCI mode but left in D3 with BAR0 unassigned; the OS is expected to allocate it. The first probe took the zero at face value, mapped physical page 0 and read the real-mode IVT as IC_COMP_TYPE (0x02000100). Mapping page 0 is itself a hazard: a kernel null dereference stops faulting. - A BAR below 1 MiB or not 4K-aligned now counts as unassigned and is never mapped. - An unassigned BAR0 gets 0x40_1000_0000 (BAR1 the next page), where Linux puts these controllers on this platform family. There is no general allocator, and PCI0._CRS cannot be trusted here (it reads host-bridge PCI_Config, which our OS layer answers with all-ones), so the address must be shown to be free: above TOUUD, within the CPU's physical address width, above every BAR base and bridge prefetchable window any other function has, and unmapped at that virtual address. Only this controller is sized, with its decode off; live devices are only read. - Memory decode is enabled only once BAR0 holds a real address. - The probe reports every fact the decision used, so a refusal says which check refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware refused 0x40_1000_0000, correctly: another device already had a BAR at 0x60_01B2_8000, so this firmware's 64-bit devices live near 384 GiB and 256 GiB was never shown to be in the window. The "nothing above 256 GiB" rule was too coarse to find a real address. - The root bridge's 64-bit window is readable after all. PCI0._CRS takes it from M64B/M64L, plain Names in the SANV SystemMemory region (unlike the PCI_Config fields that made _CRS itself untrustworthy). Discovery reads them alongside the timing variables. - The candidate is the first 1 MiB boundary inside that window that is also above every existing claim. Live BARs are still only read: a BAR's extent is bounded by its alignment (PCI BARs are size-aligned), and only the highest base needs the bound, since anything lower extending past it would overlap it. Bridge windows count by limit. - The probe reports the candidate and the window even when it refuses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Phase 2 worked on hardware: BAR0 at 0x60_01C0_0000, Designware identity confirmed, fast mode on the firmware's timings, and a valid HID descriptor. The touchpad is ELAN 04f3:30cb (not ALPS, as the address had suggested), and it answers on I2C with no _DSM handover. Phase 3a, read-only: - Transfers larger than the 64-deep FIFOs (the 381-byte report descriptor) are fed as space allows. Interrupts are masked only per refill. If the TX FIFO runs dry and the controller ends the message, that is reported as Underrun and retried, never returned as a short read. - hid_desc.rs: a report descriptor parser that flattens every input field to (report id, bit offset, size, usage, app collection), and finds the relative mouse layout. Pure, and host-tested against a mouse + precision touchpad descriptor: 6 tests. - The probe reads and parses the descriptor, then samples the input register for 3 s, counting reports per ID and decoding them as a mouse where the layout allows. Findings come back as text via syscall 578 op 2. Deliberately no RESET, SET_POWER, Set Feature or _DSM: the touchpad is the machine's pointer over PS/2, and any of those could move it off that path before this driver can replace it. Nothing is fed to the cursor yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware, Phase 3a: the 381-byte report descriptor reads cleanly and parses into mouse (id 1: buttons bit 0, X bits 8-15, Y bits 16-23), touch pad (id 4), device config and a vendor collection (id 92). But 750 reads of the input register over 3 s of finger movement returned 0 reports: the device stays silent on I2C until the host initialises it. `touchpad on` (syscall 578 op 3) sends SET_POWER(ON) and RESET, clears the reset sentinel, and samples for 3 s. Only if mouse reports actually arrive does the I2C driver take the pointer: `poll` drains the input register every 4 ms from the usb-hid task and moves the cursor through the new mouse::update_relative, at the same x2 gain as PS/2. Phase 4: while POINTER_ACTIVE is set, the PS/2 AUX handler drops its bytes, so a touchpad reporting on both paths cannot move the cursor twice. 50 consecutive I2C read failures clear the flag and hand the pointer back to PS/2. `touchpad off` (op 4) does the same on request. `touchpad handover` runs the firmware _DSM as `acpi probe 14` (breadcrumbs 73/74), for the case where initialisation alone produces no reports. It is the call that makes the EC stop PS/2 mouse emulation, so it stays separate and deliberate, never automatic. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… tunable Hardware: `touchpad on` worked. SET_POWER and RESET succeeded, the reset sentinel read back as 00 00, then 453 mouse reports in 3 s (~150/s) with 0 bus errors. I2C now drives the pointer, with no firmware handover needed. It was "very fast, almost too sensitive". The fixed x2 gain was copied from the PS/2 path to keep the feel the same, but over I2C the ELAN part reports in finer counts than its PS/2 emulation did, so the same multiplier overshoots. - Speed is now a percentage, default 100 (half the old speed), set with `touchpad speed <10-400>` (syscall 578 op 5; 0 only reads it). - Scaled motion keeps its sub-pixel remainder between reports, so below 100% a slow finger still moves the pointer instead of truncating to 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware: the pointer ran off to the edge of the screen on its own after moving. The driver read the input register every 4 ms regardless, and without an asserted interrupt this ELAN part answers with its LAST report again, so the last motion was re-applied on every poll. The first sampling run showed it: 453 "reports" from 750 reads, mostly identical. - GSI 82 (the _CRS Interrupt: Level, ActiveLow) is routed through the IOAPIC to vector 0x32 after checking the IOAPIC has that many entries. The handler only masks, flags and EOIs; the usb-hid task reads exactly one report per assertion, then unmasks. A second queued report re-asserts the line as soon as it is unmasked. - The line is routed before RESET, so reset completion is observed on it rather than timed. That also proves the routing: if it never fires, `touchpad on` refuses the pointer instead of polling blind. - mouse::update_relative now holds MOUSE_STATE with interrupts masked. Its callers are kernel tasks at IF=1, and syscall 505 takes the same lock at IF=0. A task preempted while holding it would have left that syscall spinning forever on its core. - The sampler shows reports that carry motion or buttons, and counts left and right presses separately, so a click is visible in the output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware confirmed the interrupt-driven path: the IRQ fired ~5 ms after RESET, and 326 reads gave 326 real reports (none replayed), with 113 left-button reports. The user reports the speed is right, no drift, and tap-to-click opens apps. - The usb-hid task now enables it itself, as a state machine in service() so USB polling is never blocked. At 8 s of uptime (when `touchpad` used to be typed by hand, never the governor's automatic first pass) it requests `acpi probe 13`, re-asks every 3 s in case the single probe slot was overwritten, and gives up after 15 s. - At boot nobody touches the pad, so the pointer is taken on the interrupt firing at RESET instead of on sampled reports. - The fallback for that weaker proof: every I2C report resets a counter, and every PS/2 AUX byte dropped while I2C is active increments it. If PS/2 delivers ~20 packets while I2C stays silent, PS/2 gets the pointer back. - `touchpad status` (syscall 578 op 6) shows which path drives the pointer, whether it fell back, probes completed and interrupts taken. In QEMU a clean boot reads "probes run: 1" with no command typed, which proves the boot path ran, and it correctly declined (no touchpad there). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The user confirmed on hardware that both physical buttons report, and that the pointer "initially is slow": the boot-time enable waited 8 s for the governor, leaving the old PS/2 path in charge until then. The wait was caution about ACPI on the governor's automatic first pass, which once killed three boots. But there is a safer place than the governor for this: scan_for_modern_inputs already walks PNP0C50 on every boot, right after ACPICA starts, before the APs and the scheduler exist. Being single-threaded there removes the hazard that sends every RUNTIME evaluation through the governor (two contexts inside the interpreter at once). The evaluations are exactly probe 13's, proven on hardware across many runs, and still never the _DSM. So discovery now fills the cache during boot, and the usb-hid task enables on its very first pass, before the desktop is up. The 8 s governor path stays as the fallback when boot discovery finds nothing. Verified in QEMU with an SSDT declaring a fake TPD0 shaped like the laptop's (0x2C on I2C1, GSI 82 level/active-low, HID2, M64B/M64L): "I2C-HID discovery at boot: 1 usable device(s)", the enable ran straight away, and with no LPSS controller behind it the pointer correctly stayed on PS/2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`touchpad ptp` switches the ELAN from its firmware mouse emulation to precision mode. After that the device reports every contact, and interpreting them becomes the OS's job: 1 finger moves pointer 1 finger taps left click (held 60 ms so the desktop sees it) 2 fingers move scroll the window under the pointer 2 fingers tap right click 3 fingers swipe up opens the Command (Super), down dismisses (Esc) clickpad press left, or right with two fingers down - hid_desc: Feature items are parsed, with offsets per (kind, report ID). ptp_layout groups the touch pad's fields into contact slots by order (each Tip Switch starts a finger), and input_mode finds the Device Configuration feature that switches modes. - gesture.rs: the recognizer. Pure, with the caller passing the time, so every rule is tested on the host. Motion is taken only while the same fingers stay down, so a finger landing never throws the pointer. - i2c_hid: SET_REPORT on the Input Mode feature (command register, then data register, then a length that counts itself). Hybrid frames are assembled across reports by contact count, the device's confidence bit rejects palms, and a full pad width is 1.5 screen widths at 100% speed. - Scroll: the kernel accumulates it (578 op 7, a non-blocking swap), and the shell applies it through the scrollbar's existing MSG_SCROLL. So every app that scrolls by its bar now scrolls by touch. The shell builds on the last target it sent for 150 ms, because the app's header lags it. - push_key: keystrokes queued from kernel tasks (USB HID and now gestures) take KEY_QUEUE with interrupts masked; the keyboard IRQ takes the same lock. Opt-in (`touchpad ptp`, back with `touchpad mouse`) until proven on the hardware: a mis-parsed layout would freeze the pointer, and because the touchpad keeps reporting, the PS/2 fallback would never trigger. Host: 20 parser + gesture tests. QEMU (fake TPD0 SSDT): mode-switch round trip answers "not active", desktop and PS/2 unaffected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ports Hardware: on `touchpad ptp` the pointer went instantly far from the centre. The device's real precision-mode layout is not known yet, so this makes the next run diagnosable and fixes what can be fixed blind. - Jump rejection (gesture.rs): no finger moves a quarter of the pad between two frames ~7 ms apart. Such a delta is a tracking glitch and is dropped rather than applied, as libinput does. Motion resumes from the new position on the next frame. - ptp_layout grouping: a field the finger being assembled already has now starts the next finger. Grouping on every Tip Switch broke on ELAN descriptors, whose fingers declare Confidence BEFORE Tip Switch as one two-bit field, which gave each finger's confidence to the previous one. - `touchpad ptp` refuses when the X/Y range parsed below 64, since the scaling divides by it and would explode into jumps. - `touchpad` prints finger 1's fields in full (offset+size, range), and `touchpad log` (578 op 9) shows the last 8 precision-mode reports raw beside what was decoded from each, so a wrong layout is visible. Host: 22 parser + gesture tests, including the ELAN ordering and a jump. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware showed the precision layout is sane (1 slot per report, X
0..3211, Y 0..2431, contact count in byte 7, button in byte 8) and
exposed two bugs of mine.
- Running the `touchpad` diagnostic tore the live I2C driver down and
never brought it back, silently returning the pointer to PS/2; then
`touchpad ptp` found "nothing to switch". Every probe now re-takes the
pointer on the interrupt-at-RESET proof (the boot semantics), and turns
precision mode back on if it was on, since RESET drops it.
- `touchpad on` still required mouse reports within a 3 s window, so it
refused whenever nobody touched the pad then ("0 reads"). The reset
interrupt has proven itself on every run; it is the proof now.
The likeliest cause of the pointer being thrown: while I2C drives the
pointer, PS/2 bytes are dropped and the 3-byte decoder freezes
mid-packet. Whenever PS/2 took back over (the fallback, or the teardown
above), its first packet was decoded misaligned, which throws the
pointer. PS/2 is now framed by time: a packet's bytes arrive back to back
and packets are milliseconds apart, so a byte after a >5 ms gap starts a
packet. After a suppression, nothing is decoded until such a boundary.
The bit-3 check alone could not find it, since a motion byte can have
bit 3 set too.
QEMU (PS/2 is its only pointer): clicks land exactly where sent.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware: after `touchpad ptp` the pointer zig-zagged, ran to the bottom and vanished. `touchpad log` was EMPTY, so not one I2C report arrived after the switch and the precision decoding never ran. Status showed the PS/2 fallback had fired: the motion came from PS/2, which started emitting garbage the moment the touchpad changed mode. The working theory: the EC's PS/2 emulation is still attached (the _DSM handover was never sent). In precision mode it takes the touchpad's output and translates it as if it were mouse reports. Windows and Linux send the handover before enabling precision mode. - The PS/2 fallback, while precision mode is on, now reverts the device to I2C mouse mode (proven) and keeps ignoring PS/2, instead of handing the pointer to PS/2's garbage. Reported as PTP_FAILED (status bit 6). - Every read in precision mode is logged, empty and oversized ones included, so "no reports" can be told from "reads that returned nothing". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware: `touchpad handover` panicked the machine. EXCEPTION: GPF IP 0x17b3971 = AcpiExSystemMemorySpaceHandler+0x181: movaps %xmm0,0x10(%rsp) movaps faults on any address that is not 16-byte aligned. The SysV ABI requires RSP = 8 (mod 16) at a function's first instruction, which is the state right after a `call` pushes its return address. Every kernel task started through a crafted iretq frame (thermal governor, BSP and AP idle tasks, usb-hid) got RSP = kernel_stack_top, which is 16-aligned. So each ran its whole life misaligned by 8. Most code never notices. This path does: the handover's _DSM goes EV5 -> EDPE -> EISC -> GENS -> SMBF, Dell's SMI mailbox, which creates SystemMemory regions on the fly. The first access to a new region allocates its context and spills an XMM register with an aligned store. The same _DSM ran fine at boot in 5db7eac, on the aligned boot stack, which is what pointed at the stack rather than the AML. The fix: RSP = top - 8, at all four sites. The slot plays the return address; these tasks never return. Reproduced and verified in QEMU. The fake-touchpad SSDT's _DSM creates a SystemMemory region per call (read-only, BIOS area): `touchpad handover` panics the unfixed build and not the fixed one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…host bridge Hardware: after the handover (which no longer panics), `touchpad ptp` left the pointer dead, and nothing brought it back but typing `touchpad mouse`. The PS/2-based recovery could not fire, because the handover is the call that switches PS/2 emulation off. - Precision mode is a TRIAL: if no touching finger is decoded within 10 s, the device goes back to mouse mode by itself. - `touchpad handover` re-initialises straight afterwards (SET_POWER, RESET, interrupt check, take the pointer), as Windows and Linux do. The EC may reset the touchpad when it releases it, and going from handover straight to ptp skipped the re-init. QEMU found a real hazard while testing this. An unresolved controller path left ctrl_adr 0, which decodes as device 0 function 0, and the probe landed on the HOST BRIDGE (8086:1237). The BAR checks refused before anything was written, but two guards now make sure: - ctrl_adr 0 is refused outright (NO_CONTROLLER). - bring-up refuses any function that is not a serial-bus controller (class 0x0C; LPSS I2C is 0x0C80) before touching power, BARs or registers. Verified by aiming the fake SSDT at QEMU's ISA bridge (8086:7000): refused, nothing touched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware: after `touchpad handover` the re-init succeeded (the RESET interrupt fired at ~5 ms), yet even MOUSE mode then delivered no motion. Before the handover the same re-init worked. The leading hypothesis: once the firmware knows a real driver is present, the device reports in precision mode (touch pad report ID 4), and the driver in mouse mode silently dropped every such report as "not the mouse ID". - A touch pad report arriving in mouse mode now switches the driver's decoding to precision mode (ADOPTED_PTP, status bit 7) instead of being dropped. - Every report is logged in every mode, and counted by kind (mouse, touch pad, other ID, empty/oversized; 578 op 10). `touchpad status` shows the counts, so the next hardware run says exactly what the device sends after a handover. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware: after the handover, even a second re-initialisation got only
zero-length reads (00 00 then filler) from the touchpad. It answered
every command and raised its interrupt, but never produced input.
(ELAN's harmless empty interrupts are a known Linux quirk,
I2C_HID_QUIRK_BOGUS_IRQ.) Also corrected: the device ACPI reports is
\_SB.PCI0.I2C1.TPDA ("DELL09E1"), not TPD0. Its _DSM runs the same
EV5 -> PMED handover.
Setting Input Mode was all this driver did. hid-multitouch does more:
1. reads the Win8 certification blob (vendor feature 0xFF00:0xC5)
once, "to enable some devices";
2. sets Input Mode = 3;
3. sets Selective Reporting: Surface Switch and Button Switch = 1,
read-modify-write so other fields in that report keep their values.
`touchpad ptp` now does all three, and prints which steps the device has
and whether each succeeded. General GET_REPORT/SET_REPORT feature helpers
replace the one-off Input Mode writer. hid_desc gains find_feature, and a
host test for a 4-byte-usage blob plus the switch pair.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ad end here Hardware: with the full hid-multitouch enable sequence after the handover (Win8 blob read: done; selective reporting: done), the touchpad still sent only zero-length reads over I2C, while PS/2 data kept arriving, so the fallback correctly returned the pointer to PS/2. On this Dell the handover leaves the touchpad on the EC's path rather than freeing it for I2C. Precision mode over I2C stays parked until there is ground truth from Linux on this machine. What does work: in mouse mode (no handover) the ELAN firmware does its own gestures, and two-finger scroll comes out as a wheel. Its mouse reports are 11 bytes, of which buttons and X/Y use 3. - MouseLayout finds the wheel (GD 0x38) and AC Pan (Consumer 0x238). Host test with a wheel + pan mouse collection. - Mouse-mode wheel notches feed SCROLL_ACCUM (48 px per notch, up moves the view up), so the shell's existing MSG_SCROLL path scrolls the window under the pointer. No precision mode, no handover. - `touchpad` prints where the wheel and pan are; `help` marks the handover as experimental and harmful on this machine. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ground truth from Fedora on the same laptop:
- /proc/bus/input/devices: the touchpad runs over I2C, as "DELL09E1:00
04F3:30CB Touchpad" under hid-multitouch + i2c_hid_acpi.
- /proc/interrupts: IRQ 82 IR-IO-APIC fasteoi (level), 6005 interrupts;
i8042 IRQ 12 only 137. Linux's i2c_hid_acpi always evaluates the _DSM,
so after the handover I2C is meant to work. The silence Nyx saw was
Nyx's bug, not the laptop.
- i2c-hid debug, one finger:
0e 00 | 04 | 03 | cc 03 | 45 06 | 8c e0 | 01 | 80 | 19 33
len id tip+conf X=972 Y=1605 scan count btn extra
This matches the layout the parser derived (conf bit 0, tip bit 1, id
bits 4-7, X/Y 16-bit, count byte 7, button byte 8). A lift is 04 01.
The difference: Linux reads a report with a plain I2C read
(i2c_hid_get_input -> i2c_master_recv). Nyx wrote the input register
(03 00) first, then read. That worked in the default mode; after the
handover the device answered such reads with 00 00. All three input
reads (poll, reset sentinel, probe sampling) are now plain reads.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…screens Hardware, after 132b725 (plain reads): handover, re-init, `touchpad ptp`, and the pointer is driven by the precision touchpad. "It works, but the speed is slightly fast." At 100% a full pad width now moves the pointer 1.2 screen widths, down from 1.5. Scroll keeps its own scale, which was not reported as fast; the two share nothing but the pad size. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The user confirmed every gesture on the hardware: tap, two-finger tap,
two-finger scroll, three-finger swipe, and the clickpad press.
The boot state machine now carries on past the first enable:
1. discovery (boot-time, single-threaded) -> enable in mouse mode;
2. ONLY if I2C took the pointer: the _DSM handover (acpi probe 14,
re-asked every 3 s, up to 10 s). If I2C never came up, the handover
is skipped and PS/2 stays the pointer, because the handover
switches PS/2 off for the whole power cycle;
3. re-initialise, as Windows and Linux do after the handover (retried
twice, since PS/2 is gone by then);
4. precision mode WITHOUT the 10 s trial. Nobody touches the pad
during boot, so the trial would always revert it.
HANDED_OVER is set by every handover path (boot and `touchpad
handover`). After it, stray PS/2 bytes are dropped uncounted: there is
no working PS/2 path to fall back to, and a "fallback" would only move a
working precision touchpad into a dead mode.
AUTO_PRECISION = true switches this on; false boots into mouse mode.
QEMU (fake TPD0, no LPSS controller): the first enable fails, the
handover is correctly skipped, PS/2 keeps working.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On the hardware, launching an app "looked stuck": fork, execve, and the app loading its binary from NVMe all happen before it asks for a window, and until then a click had no visible answer. While a launch is in flight: - the pointer is the design's Working cursor (defined, never used until now); - the app's dock slot shows three dots stepping every 150 ms, in place of its running dot. It ends when the forked PID asks for its first window (execve keeps the PID, so the create-window request is the natural finish line), or after 15 s, so an app that fails to exec or crashes first cannot leave the desktop "working" forever. `launch` now returns the child PID; the dock and the Command both go through start_launch. QEMU caught a bug before it shipped. The main loop reads `now` once, BEFORE input is processed, so a launch started in the same pass stamps `since` AFTER `now`. `wrapping_sub` made that a huge elapsed time and the timeout cleared the launch in the pass that began it. It is saturating now. Verified with a throwaway build that never clears on window arrival (three dots, the lit one accent-blue, and the Working ring cursor), and with the real build (a single running dot once the window opens). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Command's caret gap was traced to the shell measuring in the GPU atlas while drawing in the CPU bitmap font. That fallback happens when the kernel refuses GPU text. Whether the hardware desktop was on it was left as a question for the user to answer by eye. This makes the machine answer it. - Syscall 537 counts text batches drawn and refused, and draw_text counts refusals caused by the render-engine latch (8 consecutive hung composites) separately. - Syscall 575 op 3 returns a GpuHealth snapshot: hang count and limit, whether latched, GL hangs, the text counts, and whether the Intel render engine initialised at all (try_lock; busy means present). - `gpu` prints a verdict: working (smooth Meridian font), mostly working, never succeeded, latched off after hangs, or no Intel engine. QEMU: "no Intel render engine", text drawn 0 / refused 53, as expected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
250 ms delay and 30 characters/s instead of the 8042's power-on 500 ms and ~11/s. Reverted twice before (339fa61) because it killed the then-PS/2 touchpad, and the mechanism was never pinned down. The likeliest one: both attempts read the controller's replies blindly, one byte per command. At that point the output buffer can hold stale bytes from the firmware. One stale byte shifts every "ACK" after it, and the mouse's own 0xFA is left for the packet decoder, where 0xFA (bit 3 set) looks like a packet start. Moving the handshake earlier could not fix that, which matches what was observed. Now: drain; send 0xF3 and wait (20 ms SpinDeadline) for a reply the status register does NOT mark as AUX, requiring 0xFA; the same for the rate byte; on a missing ACK, give up (re-enabling the keyboard if it is left waiting for a parameter); drain again, so the mouse sequence (still the known-good one) starts from an empty buffer. Interrupts are still off at this point in boot, so no handler can take a reply first. The pointer no longer depends on PS/2 (the I2C touchpad takes it at boot), and PS/2 reads are now framed by time, so even a stray byte cannot desynchronise it for long. The outcome is recorded (575 op 4) and `keyboard` reports it. QEMU: "250 ms delay, 30 characters/s (set at boot, acknowledged)". QEMU's 8042 never reproduced the old failure, so the hardware has the final word. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hardware (`gpu`): render engine latched off, GPU text drawn 0, refused 233. It has never completed a batch on this laptop (the notes record the same hangs back in July). Every even-length submission this driver makes worked on the hardware: the 4-dword MI_STORE_DATA_IMM self-test and the 6-dword PIPE_CONTROL fence. The one odd-length submission, the 3-dword MI_BATCH_BUFFER_START, hung, and it is how every composite and every GPU text draw is issued. On Gen8+ the ring TAIL must be qword-aligned; i915 pads every emission with MI_NOOP for this. A misaligned tail leaves the command streamer undefined, and every submission after it misaligned too. The boot code had put this down to "BB_END doesn't return on this HW" and skipped batches in its self-tests. The batch was probably never entered correctly. - rcs_submit pads to an even dword count with MI_NOOP (and counts the pad in the free-space check). exec_batch pads the batch after BB_END too, as i915 does. The BLT ring was already written with even lengths, which is why it always worked. - The batch self-test runs at boot again, and every self-test result is recorded (BOOT_TESTS). - The first fence timeout of a boot captures the engine state i915's error capture uses: HEAD/TAIL/CTL, ACTHD (where it is executing), IPEHR (the command it choked on), IPEIR, INSTDONE, MI_MODE, EIR, FAULT, ERROR_GEN6, forcewake ack. Before this it went only to a serial port the laptop does not have. - `gpu` prints the boot tests and the first-hang registers (575 op 3, GpuHealth now 104 bytes). Untested on hardware: QEMU has no Intel GPU. The next boot's `gpu` either shows "batch pass" and GPU text drawn, or the registers of what is next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The fallback for refused GPU text was nyx_gui's bitmap font, a different typeface with different metrics. On the test laptop that was the desktop's ONLY font, since the render engine never completed a batch. The font atlas needs no GPU: it is coverage in ordinary memory, built at shell start either way. Atlas::draw_cpu blends each glyph quad from it, so the fallback keeps the real Meridian typography and the exact metrics `measure` uses (TEXT_ON_GPU now means "atlas text"). It is also better than the GPU path in one way: the batched shader carries one luminance per quad, while this blends the full colour. The bitmap font remains only for the no-atlas case. - The first QEMU run drew nothing: the shell's label colours do not keep the alpha byte meaningful (the GPU shader ignores it; fades lerp the colour). Opacity now comes from coverage alone. - Font fallback (nyx_gui font): a character a face lacks is taken from the default DejaVu face before rasterize substitutes '?'. The Meridian faces lack the return arrow, so "↵ open" read "? open". QEMU, which always takes this path: dock icons, captions, the Command's tracked uppercase heading and 20px light query all in Meridian, the caret snug after the text, and "↵ open" correct. 675 host tests pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lift On the hardware, tap-then-hold to drag a window came out as a double click, which on a caption maximises the window. The tap sent press+release at the lift, so tap + touch was two presses in quick succession. Now, as libinput does, a one-finger tap presses at the lift and holds the RELEASE for 180 ms. A finger landing in that window keeps the button down: moving drags, lifting drops. If that second touch is itself a tap, the first click is released and a second whole click follows after a visible 40 ms gap (the desktop samples button state), so double-tap is still a double click. gesture::Engine::tick closes the window when the pad goes quiet; poll() calls it. 15 host tests, including the drag that used to maximise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e engine Hardware showed `render hangs 8 of 8 (latched)` alongside "no render hang this boot": FIRST_HANG is only taken in wait_fence_value, but composites, text and GL wait in pipeline.rs finish_submit_and_wait, which has its own spin. Every strike came from a path that recorded nothing. SCENE_HANG captures the engine registers there (before reset_render wipes them), plus the last PROGRESS MARKER the stream wrote, which says which stage the engine stalled at (prologue step, or which mesh), the stream length, and whether the ring refused the submit or the fence never came. `gpu` prints it decoded. GpuHealth grows to 168 bytes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the device ID The first SCENE_HANG from hardware: every composite stalls on the PIPE_CONTROL after mesh 0's 3DPRIMITIVE (IPEHR 0x7a000004, last marker 0x20), with INSTDONE_1 = 0xffdfffff — every geometry unit DONE, only CS Done clear. The draw hangs in the pixel stage (dispatch, sampler, or RT write), which INSTDONE_1 does not cover. No boot test has ever run a pixel shader, and the 3D engine was brought up on a Comet Lake-H (0x9BC4); the test laptop is a different SKU. `gpu retry solid` swaps the window quads' PS for a four-mov constant magenta with no sampler message and clears the latch; `tex` uses the plain textured PS; `normal` restores. GPU text stands aside during a test so it cannot muddy the result. `gpu` now shows the PCI device ID and the active shader. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… dead GPU text July's idle-heat fix (383e7f4, the same commit that added GPU text) made park_gpu release the RENDER forcewake after 1 s idle so the GT can reach RC6, on the stated promise that "ensure_ready() runs before every render". Only the GL path called it. The compositor and text drew on a GT that had been through RC6 — and with no hardware context, RC6 wipes the MOCS tables while the ring registers survive. ensure_ready compounded it: it restored MOCS only when the ring was lost, and the hardware snapshot shows CTL still 0x1. Render-target writes with undefined caching never retire — `program_mocs` documents exactly that — which matches the hardware: every composite hangs after the 3DPRIMITIVE with every geometry unit idle (INSTDONE_1 0xffdfffff), even with a constant-colour shader, on the same 0x9BC4 the cube once rendered on. Now ensure_ready checks MOCS entry 0 on its own (one MMIO read when intact) and re-programs it; compositor and text call ensure_ready before drawing. `gpu` shows how many restores happened, which confirms or refutes this. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ith an IPI The keyboard and mouse IRQs are routed to the BSP, and their wake loop scanned only the BSP's tasks. Since place_task spread processes across cores, the window server can live elsewhere — and then a keystroke never woke it: it saw the key only when its 2 ms read_key_wait timeout expired and its own core's next tick ran it. ipc_send (the shell forwarding a key to the app) did mark a remote receiver Ready, but that core still only noticed on its next tick. - scheduler::wake_input_waiters: wakes Input waiters on all active cores (the same cross-core pattern ipc_send and futex wake use); keyboard and mouse handlers share it instead of two copies of a per-core loop. - A reschedule IPI (vector 0x42, apic::send_ipi) to any remote core that gained a Ready task, from both the input wake and ipc_send. - `sched` shows reschedule IPIs received per core. QEMU, 4 cores, typing into the terminal: cpu1 received 15 — each one a keystroke-path wake that used to wait for a timeout or a tick. SchedStats grows to 1344 bytes (both ABI asserts updated). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Content-unchanged moves, kept separate from the edits that follow so `git log --follow` keeps each file's history: Readme.md -> README.md NYX-Evolution.txt -> docs/archive/userspace-evolution.md SYNTAX.md, CLI.md, CHANGELOG.md -> docs/qclang/ IMPROVEMENT.TXT (an 11-line Intel GPU phase list, every phase long complete, referenced nowhere) is deleted; its phases are recorded in docs/ROADMAP.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
New documentation written against the source, not the old README: - ARCHITECTURE: system diagram, design decisions, repository map - KERNEL: boot sequence, memory, SMP scheduling + reschedule IPIs, the real interrupt table (incl. 0x31/0x32/0x42), the syscall table taken from the dispatcher (65 Linux-numbered, 501-578 native, next free 579) - GRAPHICS: GPU bring-up, GGTT layout, blitter, display, render engine, consumers and CPU fallbacks, debugging, and the resolved-bugs table - UI: Meridian, frame and keystroke paths, window protocol, apps - ROADMAP: POSIX floor -> musl -> libc++ -> gate 3, and the real history - BUILD: toolchain, Build.sh, what a QEMU boot actually needs (an NVMe image — plain ./Build.sh stops at "No NVMe Drive Detected!"), host tests, CI - README: the docs hub Nine Mermaid diagrams, all validated with mermaid.parse. Anything that could not be checked is marked "Verification required" (Wi-Fi MSI vector 0x31 with no IDT entry, Gen11/12 support, USB HID on hardware, the runner's OVMF search order). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README had drifted badly: its project tree listed files that no longer exist (window.rs, executor.rs, tarfs.rs, apps/compositor, apps/network) and missed the Meridian shell, most apps and half of libs/; it said syscalls end at 574, Wi-Fi was a "driver skeleton", input was PS/2-only with no scroll, the UI font was DejaVu, claimed capability-based permissions (there are none), and its Quick Start ran ./runner/run-qemu.sh, which does not exist. Now a concise front page — status table, architecture sketch, hardware, build, repo map — linking into docs/ for detail. CONTRIBUTING drops the nonexistent nyx-entityd and points at the real build/test/syscall rules. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- archive/userspace-evolution.md: status header recording what shipped
(image viewer, std port, on-device qclang) — verified in the tree
- qclang/{SYNTAX,CLI,CHANGELOG}.md: audit headers. A Bell program in the
spec's style compiles with the host qclang; every CLI subcommand exists
(plus an undocumented `version`); spec version v0.6.0 vs crate 0.2.2 is
flagged, not resolved
- linux-cross-reference, quantum/simulator: QEMU does exist now
- network-architecture: "next free 573" -> points at KERNEL.md (579)
- terminal-browser: two-finger touchpad scrolling now reaches the terminal
- quantum/security: re-point its reference to the old README line
- gpu/intel/mod.rs: comment-only — it cited a "device table" in
NYX-Evolution.txt that never existed
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.