Skip to content

Terminal browsing - #4

Merged
Asmodeus14 merged 7 commits into
masterfrom
terminal-browsing
Sep 9, 2026
Merged

Asmodeus14 merged 7 commits into
masterfrom
terminal-browsing

Conversation

@Asmodeus14

Copy link
Copy Markdown
Owner

No description provided.

Asmodeus14 and others added 7 commits September 9, 2026 13:49
… make the timer honest

`acpi probe 9` died at mark 55 and I read that as "the attach itself, not
_REG". Wrong: AcpiInstallAddressSpaceHandlerInternal attaches the handler and
then calls AcpiEvExecuteRegMethods before returning, so the one pair of marks
around the whole call bracketed both candidates and distinguished nothing.
A breadcrumb pair only bisects if the two things it brackets are actually
sequential at that level.

ACPICA supports the split and its own header recommends it, so:

  probe 9  -> AcpiInstallAddressSpaceHandlerNo_Reg, marks 55 -> 56 (attach)
  probe 10 -> AcpiExecuteRegMethods, marks 57 -> 58 (the AML half)

_REG is survivable to skip, so a registered handler is worth having even if
_REG has to stay off permanently. On hardware probe 9 now reports INSTALLED --
the first EmbeddedControl handler this kernel has ever registered. probe 5
(_BIF/_BST) still panics with GPF error 0; that is the next thing to chase,
not a claim this fixed it.

AcpiOsGetTimer returned tsc / 100, which hardcodes a 1 GHz TSC. ACPICA's
contract is 100-ns units, so the divisor must be TSC_MHZ / 10 (~200 here) and
every ACPICA delay was running about half as long as asked. Fixing it fixes
AcpiOsStall and AcpiOsSleep together -- their own arithmetic was always right
given a conforming timer.

An earlier note in this project put that error at ~250x. It was wrong in the
direction that stops you looking: 250x reads as a smoking gun, 2x reads as
ignorable, and nobody had done the arithmetic either way. Recompute a
suspicious constant, do not restate it.

EC_TIMEOUT 100000 -> 10000 follows from the same fix. It is an iteration count
written when a "1 us" stall really took ~0.5 us, so with the timer honest the
same constant silently became 100 ms per wait, three waits deep, on the 1 Hz
governor tick. An iteration count is not a timeout unless the per-iteration
delay is real.

Also: `acpi probe` accepted only 1..=8, so 9 and 10 were silently rejected.
Step 7 was lost the same way once already.

Co-Authored-By: Claude <noreply@anthropic.com>
…ting the same pages twice

Three fixes for the illegible red screen so far, all guesses, all short. The
reason they were guesses is that nothing on screen said what geometry the
report was using, so every theory had to be tested with a power cycle.

So the report now prints its own targets, right under the title, where the
screen has stayed legible in every photograph of this failure:

  surface 0: base 0x… 1920x1080 stride 1920 bpp 4 tiling 0
  surface 1: none

Wrong tiling scrambles, wrong stride shears, a shared base means the same
pages were painted twice through different swizzles.

It has already paid for itself. The first photo with this line shows
tiling 0 -- TILE_LINEAR -- which means the swizzle-invariant full-buffer fill
added earlier never ran at all (it is gated on tiling != LINEAR) and the plain
linear path was correct the whole time. Tiling was never the cause. That kills
the theory the last two attempts were built on.

`Screen::all` now drops the second target when both describe the same base.
SCANOUT and FIRMWARE can be the same pages -- that is what happens when P1a's
plane scan or allocation fell back and left the plane on the firmware buffer --
and they are registered with different tiling, so painting both writes one
swizzle over the other. Also not the cause here: the photo shows two distinct
bases. Kept anyway, because it is right and it removes a class of failure that
only appears on the one path where nobody can debug it.

Neither of these makes the screen legible. What is left is another writer
putting stale desktop content back after the fill lands, and the artifacts are
about 16 px wide -- one 64-byte cache line at 32bpp -- which points at the GPU
render cache evicting dirty lines from the last composite. Not yet tested.

Until it is, do not trust the graphical panic screen for a breadcrumb: use the
CMOS [USERMARK] line and hold a key at power-on for the boot log.

Co-Authored-By: Claude <noreply@anthropic.com>
…ler's frame after AML

`acpi probe 9 -> 11 -> 5` now reports `present=1 bif=1 bst=1` with design 4474 /
last-full 2131 / 11400 mV. `bif`/`bst` mean the packages PARSED, so this is the
real control-method battery, not the raw-port map that happens to print the
same numbers.

The bigger result is not the battery -- the raw path already worked. It is that
an EmbeddedControl address-space handler is installed and AML can reach it,
which is the precondition for everything else this firmware routes through the
EC: thermal zones, AC adapter, lid.

★ CORRECTION TO 24a042a, EARLIER IN THIS BRANCH.

That commit message says "_REG is survivable to skip -- plenty of AML reads work
without it". True in general, FALSE for this firmware, and it cost a boot.
`ECR1`, the primitive under every EC read, branches on `ECRD`:

    Method (ECR1, 1, Serialized) {
        If ((ECRD == Zero)) { Return (EISC (0x80, Arg0, Zero)) }   // SMI mailbox
        ... Local0 = EC00 / EC01 / ... / EC06 ...                  // the EC region
    }

`Name (ECRD, Zero)` defaults to zero and the ONLY assignment of `ECRD = One` in
the whole DSDT is inside the EC device's `_REG`. So attaching a handler without
`_REG` does not leave the firmware merely uninformed -- it leaves our handler
NEVER CONSULTED, while AML detours through `EISC` -> `GENS` -> `SMBF`, which
builds a SystemMemory region from an NVS base and fires an SMI. That wild access
was the #GP at breadcrumb 61. A general ACPI truth is not a fact about your
firmware; the DSDT is, and it is free to read.

`_REG` itself is fatal here (probe 10, mark 57, never returns) because it
tail-calls `ECIN()`: a `Notify`, a cross-tree call into `GFX0.GLID`, and two SMI
round trips. But `ECIN` is LID/AC notification housekeeping, not data-path init.
So `acpi probe 11` sets `\ECRD` directly -- root-scope integer, via
`AcpiNsGetAttachedObject`, no AML executed, verified by reading back.

Justification is not hope: the raw-port battery has read these same registers
correctly, hardware-verified against Fedora, with no `_REG` for the life of the
project. The EC does not need `ECIN()` to answer. What we skip is notification,
and if something later depends on that handshake it will not work.

Other fixes here:

- The probe 5 guard read `nyx_ec_installed` ("the ports are known") when its own
  doc comment states the precondition is `nyx_ec_handler_ok` ("AML can reach the
  EC"). So every `acpi probe 5` ever run evaluated battery AML with no handler
  attached -- the exact crash the guard existed to prevent. When two flags mean
  subtly different things, the bug does not stay in the comments.

- Gate on `ECRD` itself, never on "did we call _REG". Two flags in this subsystem
  have already drifted from what they claimed; the firmware's own variable cannot.

- `acpi_battery_read` now lands its result in a volatile static and
  `acpi_battery_fetch` copies it out afterwards. Writing through the caller's
  pointer while interpreter frames unwound beneath it faulted on three boots with
  a pointer that was demonstrably non-NULL, canonical, unchanged across the call
  and high-half, on the stack we were executing on, with every AcpiOs* wait a
  bounded spin. No theory survived the evidence. This is an avoidance, not a root
  cause, and is labelled as one in the code.

- `panel` prints `last AML battery read: present/bif/bst/...`. The line that was
  meant to answer "AML or raw ports?" went to `vga_println!` -- the kernel boot
  log, invisible once the desktop is up. That is the same mistake as "serial
  still has it" on a laptop with no serial port. `battery` cannot distinguish the
  two paths, because they print identical numbers.

Breadcrumbs: 59 no handler, 60 handle, 61/62 _STA, 63/64 _BIF, 65/66 _BST,
68 ECRD still 0, 69/71 ECRD write, 72/73 unpack, 74/75 Rust side.

Co-Authored-By: Claude <noreply@anthropic.com>
`acpi probe 12` evaluates `_TMP` on all four sensors. Hardware-verified:
CPU 42 / MEM 39 / SKN 42 / M.2 37 C -- four distinct, plausible, moving values.
The sensors are `\_SB.PCI0.B0D4` and `ECDV.{TMEM,TSKN,NGFF}`, all declared in
ssdt7 and none of them in the DSDT. `KDRT(n)` writes the sensor index to EC 0x33
and reads degrees C from EC 0x34; `_TMP` returns deci-Kelvin, so C = (dK-2732)/10.

★★ THIS IS WHY `ECRD` MATTERED BEYOND THE BATTERY.

    Method (_TMP, 0, Serialized) {
        If (\ECRD) { Local0 = ECDV.KDRT (n); Return ((0x0AAC + (Local0 * 0x0A))) }
        Else       { Return (0x0BB8) }
    }

0x0BB8 = 3000 deci-Kelvin = 26.85 C. With `ECRD` clear -- which is every boot
this project has ever had until now -- EVERY thermal sensor returns that
hardcoded constant. It is not an error code and it does not look like one: it is
a plausible idle temperature, on a laptop with a documented overheating history.
A thermal stack built against it would have looked like it worked forever.

So `panel` flags it explicitly when all sensors read 3000 rather than printing
the number. Same principle already applied to Graphics % and signal strength in
the Meridian views: a convincing wrong number is worse than a gap.

AC adapter is a RAW EC READ, no AML, and that is deliberate. The DSDT gives it
away directly -- `_PSR` is `ECG5() & 1`, `BAT0._STA` is `ECG5() & 2`, and
`ECG5()` is `ECRB(0x06)` -- so EC register 0x06 bit0 is AC-online and bit1 is
battery-present, one read, no interpreter and no handler needed. No bank select:
0x03 selects the battery window only, and ECG5 does not touch it. Evaluating
`_PSR` would additionally fire `PNOT()` whenever the state differs from `PWRS`,
a Serialized method that notifies every power consumer -- real work we do not
want on a poll for a bit a register read already gives.

Results land in statics collected by `acpi_thermal_fetch`, the same shape as the
battery, for the reason recorded there: writing through a caller's pointer while
interpreter frames unwind beneath it faulted on three consecutive boots.

Marks 80..83 per sensor, 84 on completion -- one each, because a single pair
around all four would repeat the mark-55 mistake of bracketing several things
and distinguishing none.

⚠️ NOT CHANGED: the Entity's battery is still the Dell-specific raw register map.
`acpi_tick` overwrites the cache with `ec_battery()` every second, BEFORE the
probe dispatch, so probe 5's vendor-neutral values survive at most one tick.
`_BIF`/`_BST` are a proven capability, not the live source. Promoting them means
running attach+ECRD on the governor tick -- the same automatic path that panicked
the box at ring 3 -- so it wants a permanent fall-back to raw, not a swap, and
that decision is still open.

Co-Authored-By: Claude <noreply@anthropic.com>
…isarms itself

`acpi_tick` called `ec_battery()` unconditionally -- the register map decoded
from this one Dell's DSDT (0x03 bank select, 0x10 state, 0x1E last-full, ...).
Correct here, hardware-verified against Fedora, and true of no other machine.

It now calls `battery_current()`, which prefers `_BIF`/`_BST` once they have
proven themselves this boot. HW-confirmed: `panel` reports
`battery source: AML _BIF/_BST (vendor-neutral)` with no probes run.

`battery_source_init()` runs once, on the governor:

    85  attach the EC address-space handler   fail -> latch raw
    86  set \ECRD                             fail -> latch raw
    87  read _BIF/_BST
    88  require present && have_bif && have_bst && last_full_cap > 0
                                              -> AML,  else latch raw

`_REG`, the one call that is actually fatal on this firmware, is never involved.
`have_bif`/`have_bst` mean the PACKAGES PARSED -- the methods merely returning is
a different claim, and those two were conflated in this subsystem once already.
The check is repeated every tick, not just at promotion, so a source degrading to
zeroes latches back to raw instead of showing a plausible 0%.

★★ THIS PUTS ACPI BACK ON THE ONCE-A-SECOND PATH, WHICH HAS BRICKED THIS BOX.

`acpi_ec_install` briefly did the handler install on this same tick and the
machine panicked the instant it reached ring 3: no serial console, no opt-out,
unbootable without a reflash. The panic was not the real damage -- the
irrecoverability was.

So this arms itself only after checking it did not kill the last boot.
`postmortem::prev_user_mark()` is the CMOS byte that survives a power cycle; if
it names 85..=89 the previous boot stopped inside this code and we stay on raw
and say so. Worst case becomes one bad boot, then raw forever, self-healing.

⚠️ Worth reusing: `prev_user_mark()` plus a reserved mark range is a general
auto-disarm for any risky call on an automatic path. It is the difference between
a bad experiment and a brick, and it costs about ten lines.

`panel` also prints the live source, because the promotion announces itself with
`vga_println!` -- the kernel boot log, invisible once the desktop is up. That is
the same mistake as "serial still has it" on a laptop with no serial port, made
twice in one session now.

The raw map is not deleted and is not deprecated. It is the fallback, it is still
hardware-verified, and on this machine both paths print identical numbers
(2131 / 4474 / 11400) -- which is exactly why `battery` alone can never tell you
which one is live, and why `panel` now states it.

Co-Authored-By: Claude <noreply@anthropic.com>
Thermal showed one number, `current_temp`, from `get_intel_silicon_temp()` -- an
MSR read of the package sensor. The EC has three more that the MSR has no access
to at all, and now that `_TMP` works through the ACPI EC handler they are real:

    Thermal    42 C
               Within envelope . Memory 39 . Skin 42 . M.2 37

HW-confirmed: memory 39, skin 42, M.2 37.

★ CPU is deliberately NOT in the EC list. `current_temp` is already the CPU from
the MSR, and the EC's own CPU sensor reads slightly differently. Printing two
numbers for the same part only invites the reader to work out which one is lying.
What the EC adds is the parts the MSR cannot reach.

★★ 0 MEANS NO READING, NEVER A COLD PART. `thermal_refresh` maps the 0x0BB8
stub (3000 dK = 26.85 C) to 0, and `ec_sensors()` skips zeroes, so a sensor that
is not answering shows as a GAP. This firmware returns a *convincing* idle
temperature when it has nothing to say -- the one failure mode that never looks
like one. Same convention `current_temp` already used, and the same reasoning
that left Graphics % and signal strength off the other Meridian views.

The governor refresh gets the same auto-disarm as the battery promotion: `_TMP`
is AML on the once-a-second tick, so `thermal_refresh` checks
`postmortem::prev_user_mark()` against marks 80..=84 and permanently disables the
sensors if the previous boot died inside them. It also re-reads only every 4th
tick -- four AML evaluations a second, each two EC transactions with timeouts, is
real work to spend on numbers a human reads at 1 Hz at best. And it does nothing
at all unless `ECRD` is set, because without it every sensor is the stub.

⚠️ `ec_temps` is appended at the END of `SystemInfo`, in BOTH
`nyx-kernel/src/interrupts.rs` and `libs/api/src/lib.rs`. It is `#[repr(C)]`
written through a raw pointer by the kernel, so a field inserted anywhere else on
one side only silently shifts every field after it on the other. Both
declarations carry that warning.

Co-Authored-By: Claude <noreply@anthropic.com>
…ible

HW-verified with `fault panic`: the report comes up clean and solid for the first
time since it was written.

★ THE EVIDENCE WAS ON SCREEN FOR FOUR ATTEMPTS AND I KEPT WALKING PAST IT.

The desktop showed through the report as short dark runs, and the runs were about
SIXTEEN PIXELS wide. At 32bpp that is 64 bytes — one cache line. Cache-line
granularity is not something a drawing bug produces. It is the signature of
memory that has been written but not written BACK.

We paint with the CPU; the display engine scans physical RAM. With a write-back
mapping our stores sit in L1/L2 and reach RAM whenever the hardware decides, so
the panel shows new content where lines have been evicted and the old desktop
where they are still dirty — in exactly 64-byte runs. Write-combining fails the
same way with partially filled WC buffers.

Fix, after each paint pass:

    sfence      // drain the write-combining buffers
    wbinvd      // write back and invalidate every cache line

Neither takes a lock, which is the whole requirement here: the panicking core may
already hold any lock in the kernel, which is why this module is lock-free
throughout. `wbinvd` costs milliseconds and flushes the entire hierarchy; on a
machine that is not coming back that is not a cost.

★★ WHY THE THREE EARLIER FIXES ALL FAILED, AND WHY THEY ALL MEASURED CLEAN.

They were correct. The geometry was never wrong:

  - full-buffer swizzle-invariant fill  -> never ran; `tiling 0` means LINEAR
  - drop FIRMWARE when it aliases SCANOUT -> never fired; the bases differ
  - paint the whole report twice          -> changed the TIMING of eviction
                                             without ever forcing it, which is
                                             exactly why it helped inconsistently

I kept asking what we drew, when the problem was that it never left the cache.
The surface-parameter line added in d1b806d is what made this findable: it
printed `tiling 0` and killed the swizzle theory in one photograph, which is the
only reason the search moved on to the pixels' journey rather than their layout.

⚠️ `wbinvd` is applied in `fault_banner` too, which runs per dying process rather
than once at the end of the world. A process fault is an exceptional event, not a
hot path, and a banner half-resident in L2 is exactly as useless as a report that
is. If something ever crash-loops hard enough for the flush to matter, the flush
is not the problem worth fixing.

The graphical panic screen can be trusted for a breadcrumb again.

Co-Authored-By: Claude <noreply@anthropic.com>
@Asmodeus14
Asmodeus14 merged commit 0793919 into master Sep 9, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant