Skip to content

Terminal browsing - #1

Merged
Asmodeus14 merged 8 commits into
masterfrom
terminal-browsing
Sep 3, 2026
Merged

Asmodeus14 merged 8 commits into
masterfrom
terminal-browsing

Conversation

@Asmodeus14

Copy link
Copy Markdown
Owner

No description provided.

Asmodeus14 and others added 8 commits August 23, 2026 11:31
The engine could parse, cascade, lay out and paint; this is what lets a page
CHANGE after load, which is what most of the modern web depends on.

boa_engine 0.19 with default-features = false. The defaults include `intl`,
which drags in ICU and its locale data; the browser binary is already large
and Nyx has no OOM killer, so paying tens of MB for Intl.* before a single
page script runs is the wrong trade. Turn it on when something needs
locale-aware formatting.

The DOM lives in a thread-local while scripts run. Native functions handed to
boa are plain fn pointers and anything they capture must be traceable by
boa's GC, which Rc<RefCell<Dom>> is not. Making the DOM a GC-managed object
is invasive and wrong (it outlives any one script), and threading it as a
capture fights the GC for no benefit. The thread-local is honest here for a
reason specific to this engine: script execution is a BOUNDED PHASE ON ONE
THREAD. run() installs the DOM, runs every script and takes it back out, so
no binding can outlive it and two runtimes cannot disagree about which DOM is
current. run() takes the Dom by value and returns it, because a script can
restructure the tree and a borrow would have to be held across arbitrary JS.

Scripts run BEFORE the cascade. A script that rewrites text or flips a class
changes what the selectors match, so cascading first would style the
pre-script document and then paint the post-script one -- which shows up as a
page that is styled almost right, the hardest kind of bug to see.

Bindings: console.log/warn/error, document.getElementById, document.title,
and per-element tagName, id, textContent (get and set), getAttribute and
setAttribute. Elements carry their NodeId in a non-enumerable, read-only
__nid so a script cannot repoint an element at another node, and so
`for (k in el)` does not show engine plumbing.

Setting textContent replaces the children with one text node per spec, and
leaves the old nodes in the arena rather than compacting: NodeId is an index,
and compacting would invalidate every id a script still holds.

A script that throws does not stop the ones after it -- that is what browsers
do, and a page whose third banner script fails should still render. Errors go
to the serial log, never to the modal error dialog: a page whose analytics
script throws has not failed to load, and a browser that claims otherwise
trains you to dismiss the dialog that matters.

Deliberately absent rather than stubbed, because a stub that accepts a call
and does nothing makes a page render the wrong thing silently: innerHTML (it
re-enters the parser mid-tree and can restructure the arena under a live
NodeId), event listeners (no event loop in the engine yet), timers, and
querySelector (the selector engine exists in css.rs; wiring it in is its own
change with its own tests). Scripts with a non-JavaScript `type` and external
`src=` scripts are skipped -- fetching belongs to the embedder, which owns
the network.

Verified: 63 host tests green (10 new), including that a script's mutation is
visible in the DOM the caller gets back, which is what the cascade then runs
over. boa compiles for x86_64-unknown-nyx; browser binary 3.4 MB -> 11.5 MB.

Co-Authored-By: Claude <noreply@anthropic.com>
… in-repo

A git clone ... ~/foo run from the Windows session shell (which does not
expand ~) makes a real directory named ~ inside the project. git add -A then
commits it as a stray gitlink pointing at a repo nobody else can fetch.

Co-Authored-By: Claude <noreply@anthropic.com>
Rules inside @media were skipped wholesale. Wikipedia alone has 80 blocks, so
a large share of a real page's CSS never reached the cascade -- which is why
sites rendered under-styled rather than unstyled, the harder symptom to read.

Media queries are resolved at CASCADE time, not parse time. The viewport
changes on every resize, and re-parsing 200 KB of CSS to answer "is it still
wider than 600px" would make resizing cost more than loading. So a StyleRule
carries the @media lists that enclose it and `applies(ctx)` is asked during
RuleIndex::build.

A Vec of lists rather than one flattened condition, because nesting means AND
and commas mean OR, and OR does not distribute over AND -- flattening would
quietly change which rules apply.

The filter is applied when bucketing, but `declarations` still advances for
every rule including the filtered ones. Source order breaks cascade ties and
is defined over the whole sheet; skipping the count for a non-matching rule
would shift every later rule's order and change which declaration wins. That
bug would only ever appear on pages that use @media, which is to say on real
pages and not in any existing test.

Supported: media types (screen/all match, print does not), `and`, comma-OR,
`not`, `only`, and min/max-width in px and em. em resolves against the
initial 16px because media queries resolve em against the initial font size,
never the element's -- so there is no cascade dependency to chase.

An unknown feature makes the query false, per spec, and that is also the safe
direction: matching an unrecognised query would apply a print or narrow-screen
sheet to the window.

A resize now re-cascades rather than only re-laying-out, since the width
decides which rules apply and not just where lines break; dragging across a
breakpoint otherwise keeps the styles from the width the page was loaded at.
The concatenated CSS text is kept on the Browser so that path re-parses
without re-fetching.

compute() keeps its signature and defaults to a desktop viewport, so the 53
existing tests are untouched; compute_media() is the one the browser calls.

One existing test changed meaning rather than being deleted:
at_rules_are_skipped_without_eating_what_follows asserted the rule inside
@media was dropped. The property that always mattered -- the rule AFTER an
at-rule survives -- is still asserted, now against @supports, which is still
skipped wholesale.

Verified: 77 host tests green (14 new).

Co-Authored-By: Claude <noreply@anthropic.com>
.gitattributes — force LF on *.sh. Windows git defaults to core.autocrlf=true, so a
plain `git checkout` rewrites Build.sh with CRLF, the shebang becomes `#!/bin/bash\r`,
and bash reports "cannot execute: required file not found". The failure lands in the
wrapper, so the build still exits 0 while writing a one-line log and leaving the
PREVIOUS .efi.img in place. A stale image that looks like a clean build is the worst
possible outcome on a machine where every test costs a power cycle; only the artifact
timestamp gave it away.

.gitignore — ignore a literal `~` directory. A shell that does not expand `~`
(PowerShell, cmd) turns `git clone ... ~/foo` into a real folder named `~` inside the
repo, which `git add -A` then stages as an embedded git repository. Twice now. Also
ignore build_*.log.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeJAyn7ps35LUadXG4Faq2
Chasing an intermittent hard freeze exposed that every channel for observing this
machine dies with it: BOOT_LOG is read through System Monitor (a userspace app), there
is no serial port, and the screen is a GPU-owned scanout buffer the kernel cannot paint
into. A starvation freeze never panics, so the death record stayed empty too. The
result was several power cycles that each eliminated one guess.

CMOS survives all of it.

  * Liveness heartbeat, once a second from the scheduler: tick, per-core task counts,
    and how long since ANY core scheduled a non-idle task.
  * Per-core alive bitmask. This is the field that matters. `PerCpu` owns its Scheduler
    BY VALUE, so tasks/runnable/current describe only the core that wrote the beat — the
    first version reported `tasks=1 runnable=0` and concluded "every userspace task was
    Blocked" when that is just what a secondary core looks like. A core spinning with
    interrupts masked takes no timer, never reaches schedule(), and goes dark in the mask
    while the others keep beating. That distinguishes "the kernel died" from "one core
    wedged and stranded its tasks", which nothing else could.
  * WHY_KERNEL_OOM with the requested size, so heap exhaustion stops being filed as a
    generic panic with detail 0.
  * Panics record file basename + line, stored as raw ASCII so the report reads
    "panicked at memo...:412" with no lookup table to get wrong at the worst moment.
  * user_mark(): a byte userspace can set (syscall 551), so a freeze inside an ordinary
    program says which step it died on.
  * The report now prints to the SCREEN as well as the log, and holds ~10s. Printing it
    only into BOOT_LOG made it unreadable in exactly the case it exists for.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeJAyn7ps35LUadXG4Faq2
…inal

Deletes apps/browser, libs/web (html5ever + cssparser + cascade + layout) and the
Network Suite, and moves browsing to commands over the same libs/net transport. The
image drops 26MB -> 21MB. All three are recoverable from history.

  get <url>    fetch and render as text; a bare host gets https, not http
  links        list the last page's links
  open <n>     follow one, resolved via Url::join against the POST-redirect URL
  dns <host>   resolve with timing, and an interpretation of the failure

libs/htmltext does the rendering: zero dependencies, 19 host tests. A workspace member
specifically so those tests run on the host — untestable string scanning is how a tag
stripper quietly starts printing a page's JavaScript at the user. It is a scanner, not a
parser, deliberately: libs/web had a real tree builder, but linking it into the terminal
would drag the whole cascade and layout stack in to do a tag strip.

`get` guesses https for a bare host because guessing http costs an entire extra request
and DNS lookup on any site that redirects, and Fetch::finish restarts the state machine
at Stage::Resolve for a redirect. A connection-level failure walks the guess back to
http once; a scheme the user typed is never touched.

--- Clock -------------------------------------------------------------------------

Every https:// fetch failed with "certificate not valid yet" because this machine's RTC
reads 2024. rustls checks validity windows against SystemTime::now(), so a wrong clock
presents as a TLS bug and is neither.

  date              RTC and SystemTime side by side; if both are wrong the hardware is
                    wrong, if they disagree the conversion is
  time sync [url]   set the clock from an HTTP Date: header
  tz [+5:30]        DISPLAY timezone

NTP would be the right answer and is not available: UdpSocket in the PAL is
`UdpSocket(!)`, every method unsupported. An HTTP Date: header gives the same bootstrap
over transport that works, and crucially arrives over plain http:// — you cannot fetch
the time over https when the wrong time is what breaks https. It is unauthenticated;
so is plain SNTP. Good enough to make a hobby OS work, not to rely on.

syscall 552 writes the RTC so the fix survives a reboot, following the MC146818
sequence: SET bit first to halt the update cycle, values in whatever format status B
advertises (writing binary into a BCD clock silently stores a plausible WRONG date),
SET cleared last. It refuses implausible values, because the number needed to repair a
bad write is the one it would have destroyed. 550 remains as a software-only fallback.

The timezone is display-only and stored in minutes. The clock never leaves UTC —
localising it would re-break TLS windows and every file timestamp, and storing local
time in the RTC is why Windows/Linux dual-boots fight over it. Minutes because +5:30,
+5:45 and +12:45 are all real. The taskbar clock now goes through sys_get_rtc_local(),
replacing a recompile-to-change whole-hour constant that could not express +5:30 and
shifted only `hour`, leaving the DATE a day behind whenever the shift crossed midnight.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeJAyn7ps35LUadXG4Faq2
`cores alive: 0b11111110` from the boot report finally located the freeze: bit 0 clear,
bits 1-7 set. Core 0 stopped reaching schedule() while every other core kept beating —
so it was spinning with interrupts masked, taking no timer, stranding every userspace
task (they all live on core 0). Four seconds later the rest piled on and the beat
stopped too. No panic, no fault: not a crash, a spin.

That says WHICH CORE but not WHICH LOCK, so instrument the blocking acquisitions
reachable from the network path with interrupts masked: the five in
poll_network_locked, WIFI_SOCKETS, WIFI_IFACE, and MEMORY_MANAGER via virt_to_phys.

★ `lock_watched` does NOT change control flow. It spins exactly as `.lock()` did; after
~20M spins it writes one byte to CMOS naming the site and keeps spinning. This is the
corrected form of a mistake made earlier in the same investigation: giving fifteen
MEMORY_MANAGER acquisitions a watchdog that ended in `loop { hlt() }` converted "wait
for the other core" into "kill the machine if the holder is slow" and hard-froze the box
at boot. The diagnosis was right and the remedy was not. A byte cannot introduce a new
failure mode; a slow-but-healthy acquire is unaffected.

virt_to_phys is included because MEMORY_MANAGER is a bare spin lock with no interrupt
masking and is reached from the NIC, NVMe, AHCI and USB drivers — i.e. from interrupts-
off context. It is the prime remaining suspect, and it is instrumented, not watchdogged.

Ruled out by reading rather than by a boot: serial::_print already wraps SERIAL1 in
without_interrupts, so it cannot be the preemptible holder.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeJAyn7ps35LUadXG4Faq2
master carried three commits that never left this machine — the boa JavaScript engine
and DOM bindings (f02171d), the @media cascade evaluation (abc3495), and a .gitignore
fix (cf69a2e). All of them touch apps/browser and libs/web, which this branch deleted
when browsing moved to the terminal.

Resolved every modify/delete conflict in favour of the deletion. The resulting tree is
byte-identical to c3af65f: this merge changes no code, it only stops that work from
being stranded on a branch nobody would look at again. The engine, the cascade and the
layout code stay reachable in history if the graphical browser is ever revived.

cf69a2e is superseded by 31e8e39, which added the same `~/` rule with the reasoning;
the duplicate the auto-merge produced is removed.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeJAyn7ps35LUadXG4Faq2
@Asmodeus14
Asmodeus14 merged commit 43f5965 into master Sep 3, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant