Skip to content

Terminal browsing - #5

Merged
Asmodeus14 merged 9 commits into
masterfrom
terminal-browsing
Sep 22, 2026
Merged

Asmodeus14 merged 9 commits into
masterfrom
terminal-browsing

Conversation

@Asmodeus14

Copy link
Copy Markdown
Owner

No description provided.

Asmodeus14 and others added 9 commits September 11, 2026 09:54
…imits

Found by auditing rather than by a failure, which is the point — none of these announced
themselves.

★★ smoltcp was seeded with zero. `Config::random_seed` was never set on ANY of the three
interface constructions, leaving it at its Default. smoltcp derives TCP initial sequence
numbers and DNS query IDs from that seed, so both were identical on every boot of every Nyx
machine. The DNS consequence is the serious one: the query ID is the primary anti-spoofing
token in a UDP exchange, and a predictable ID reduces cache poisoning from guessing a number
to winning a race. The kernel already had a real RNG and it simply was not wired here.

★★ getrandom reported success while returning non-cryptographic bytes. `random::fill` returns
false when it fell back to the TSC-seeded xorshift; syscall 318 discarded that boolean and
returned the byte count regardless. Its caller checks only the count, so weak entropy would
reach rustls — which draws ECDHE private keys, the client random and GCM nonces from exactly
this call — with nothing in the chain able to notice. `random::is_cryptographic()` was written
to guard this and had ZERO callers: the design saw the risk and the wiring was never finished.

Now a NYX_GRND_CRYPTO flag, so key material fails closed while plain callers keep the
always-succeed behaviour. That asymmetry is deliberate: std seeds every HashMap through
getrandom during startup and cannot cope with an error, so making the default path fail would
make a DRNG-less machine unbootable.

sys_socket scanned `3..32` while FD_MAX is 256 and every other allocator walks all of it. A
process could hold 29 sockets, and one whose low fds were already files got EMFILE with 224
slots free. The browser opens a connection per request, so it can reach this.

NEXT_LOCAL_PORT was a bare fetch_add on a u16: after ~16k sockets it wrapped to 0 and walked
up through 22, 80, 443. Now folded into the RFC 6335 dynamic range.

The RTL8168 ISR printed a line per packet interrupt, with interrupts masked, to an ~11 KB/s
UART this laptop has no cable for. Under real inbound traffic the machine would spend all its
time in the handler. Replaced by an atomic counter.

sys_connect gained a fourth argument, a caller-supplied handshake deadline, because a fixed
10 s is the wrong answer once a client walks a name's several addresses looking for one that
answers — four unreachable candidates cost forty seconds. 0 keeps the old default.

Also the wired driver's `max_burst_size = Some(1)`. In smoltcp that is not a transmit hint, it
CLAMPS the advertised TCP receive window to one segment; the same defect was measured at
~5 KB/s on the WiFi path and the fix never reached here.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
The parser hand-computed the 802.11 header length — 24, plus 6 for a 4-address frame, plus 2
for QoS — and then SCANNED up to 20 bytes for the LLC/SNAP signature to absorb whatever the
arithmetic could not model. The device has been reporting the exact answer all along:

    header_len = (mac_flags2 & 0x1f) * 2      in 2-byte words
    pad        = (mac_flags2 & 0x20) ? 2 : 0

Offsets confirmed two ways: from Linux's fw/api/rx.h (facts only — Nyx is Apache-2.0 and Linux
is GPL-2.0, see docs/linux-cross-reference.md), and from `mpdu_len` at p+8 having worked since
this driver was written, which anchors the whole layout. `descsz=48` measured on hardware
confirms the v1 variant: a 20-byte common header plus a 28-byte tail, with every field used
here living in the common part, so the v1/v3 distinction does not matter.

★ The first attempt matched ZERO frames, and the design is why that cost one boot instead of an
investigation: the descriptor offset is TRIED, the signature is CHECKED there, and the old scan
is kept as a fallback — so a wrong assumption costs a counter rather than lost traffic. The
counter then named the error. `claimed=34` against `true=28`, a difference of exactly 8: the
descriptor's header length ALREADY includes the security header, and adding `ccmp` on top
double-counted it.

Decryption status is in the descriptor too (STATUS_DECRYPTED, bit 11). Every frame previously
dropped as NOSNAP was a group-addressed one encrypted under the GTK that we never decrypt — and
we discovered that by hunting for a signature that could not possibly be there. One bit-test
replaces the search, and the counter now states the reason instead of a symptom of it.

Also: the bridge-tunnel LLC/SNAP variant (aa aa 03 00 00 f8) is a legal encapsulation we were
rejecting outright, having only recognised the RFC1042 form.

DHCP options 51 (lease), 58 (T1) and 59 (T2) are finally parsed, big-endian as network byte
order requires, with the acquisition time recorded. T1/T2 fall back to RFC 2131 4.4.5's
defaults when the server omits them. Nothing here renews — that is the smoltcp socket added in
the next commit — but a lease that cannot expire silently is the precondition for it.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
The lease was taken once at sys_wifi_connect and held forever, because option 51 was never
parsed. When the server expired it, the link silently stopped routing — the least debuggable
failure this machine can produce.

The WiFi socket set now carries a smoltcp dhcpv4::Socket, driven from poll_wifi exactly as the
wired path drives its own. smoltcp owns T1/T2 internally and re-requests on schedule.

★ The two DHCP clients are SEQUENTIAL, never concurrent, which is what makes this safe. The
driver still performs the first exchange over raw 802.11 frames, and that is not redundant: it
is what proves the encrypted link carries traffic AND what calibrates rx_desc_size, both of
which must happen before an IP interface can exist at all. Then smoltcp takes over. The
objection recorded in the old comment — "two DHCP clients bidding for the same MAC" — applies
to running them at the same time, which this does not do; the server returns the same address
to the same MAC, so the handover is invisible.

★★ And the part that panicked the box. socket/dns.rs:588 dispatches with
`cx.get_source_address(dst).unwrap()`, carrying upstream's own "TODO remove unwrap", which
returns None whenever the interface holds no address in the destination's family. That is a
KERNEL PANIC reached from ordinary idle-task polling, and it fired by two different routes: a
lookup issued before DHCP granted a lease, and — introduced by the work above — a freshly
created dhcpv4::Socket emitting `Deconfigured` on its FIRST poll. It is announcing "no lease
from me yet", not reporting one lost, and obeying it tore down the driver's working bootstrap
lease while the boot-proof query was still in flight.

Guarding the syscall was never sufficient, and that is the lesson worth keeping: the syscall is
not the only thing that starts a query. init_wifi_iface starts one, so does the retry inside
poll_wifi, and any future caller would inherit the trap. So gate the SOCKET, not the callers:
with no usable IPv4 address, update_servers(&[]) makes smoltcp's own dispatch take the branch
immediately above the unwrap — `if pq.server_idx >= servers.len() { set Failure; continue }` —
failing pending queries cleanly through a path it already supports.

Applied to the wired stack too, which has the identical hole and had simply never been observed
to panic.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
connect_timeout() was unsupported() on the grounds that the kernel would not take a deadline.
It does now (syscall 42 gained a fourth argument), and that matters for one reason: a client
walking a name's several addresses needs to give up on each quickly. At the fixed 10 s, four
unreachable candidates cost forty seconds of blank window.

sys4 passes the fourth argument in r10, NOT rcx — the syscall instruction clobbers rcx with the
return address.

The rest is diagnostics, and they exist because this laptop has no serial console: every
serial_println! in the drivers is written and never read. Syscalls 569-572 put the numbers
where a person can see them — the RX parser's counters, the NOSNAP frame dump, the ring state,
and a resolver that returns up to four addresses instead of one.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
…f five

★ "every url with get should work, even links" — and the worst of these was silent. Url::join
tested only for "://", so mailto:, tel: and data: (which have no slashes) fell into the
RELATIVE-PATH branch: `mailto:a@b` became the path /mailto:a@b, requested from whatever host
the reader happened to be on. libs/htmltext numbers those links, so `open <n>` hit it
routinely. Now RFC 3986 scheme syntax, reported by name. A colon in a later path segment is
still a path, and ./ still escapes a first-segment colon, matching browsers.

The request target was interpolated raw into `GET {path} HTTP/1.1`. That target is
space-delimited, so one space produced a malformed request line — and a CR/LF in a crafted link
could inject headers. Now percent-encoded, conservatively: already-encoded triplets pass
through, because re-encoding %20 into %2520 asks for a different file.

Also: userinfo mis-parsed as a port, port 0 accepted, host not case-folded (splitting the DNS
cache in two).

Errors now differ. Every failure used to arrive as Io — a name that does not resolve, a host
that will not answer, and a certificate that is not valid yet are three different problems with
three different fixes, and "Network error" is exactly the outcome this taxonomy prevents.

★★ Connection reuse. Every link click previously paid a fresh DNS lookup, TCP handshake and
FULL TLS handshake — the last of which is a software certificate verification on this machine.
Three rules keep it safe, and rule 1 is the one that matters: only a FRAMED response may be
reused (Content-Length or chunked, never read-to-EOF). A body delimited only by the close has
no length, so "the body ended" and "the peer paused" are the same observation, and guessing
wrong hands the remainder to the NEXT request as if it were its response. That rule is
mutation-tested, because its failure corrupts a different request than the one that broke.

Multi-address resolution, held in the Fetch rather than re-resolved: a connect failure drops
only the address that failed. Going back to DNS for the next candidate cost a round trip and,
worse, a transient resolver failure then reported "cannot resolve" for a name that had already
resolved.

gzip (2.8-5.3x measured), charset decoding, and a silent-peer watchdog for a host that accepts
a connection and then sends nothing — measured against a Google frontend: 43 reads, zero bytes,
45 seconds.

99 host tests, up from 23. The HTTP framing is driven through a reader that returns n bytes per
call, swept down to n=1, because a terminator straddling two reads is where framing bugs live.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
Browsing: back/forward/reload/history, open <url> as well as open <n>, find/next over the PAGE
rather than the scrollback, reader mode, stop, and up/down command history — the arrows
previously fell through to the input buffer and typed a literal box glyph into the command.

The scrollback was one String, re-split and re-allocated into a fresh Vec<String> EVERY FRAME,
with a comment claiming it was capped. It was not: MAX_PAGE_CHARS caps one page, and nothing
capped the total. Now a bounded line buffer with the wrap cached against a revision — and the
cache deliberately split at the boundary between "enormous and rarely changes" (the scrollback)
and "two lines, changes every frame" (status and prompt), because folding the live status in
re-wrapped the whole history sixty times a second to animate a millisecond counter.

Rendering gained headings, bullets, verbatim <pre>, blockquotes, table cells, <base href> and
word wrap with hanging indents. A page used to arrive as a wall of lines where a heading and a
paragraph were indistinguishable, which is readable in the sense that the words are present and
unreadable in the sense that you cannot skim it.

time sync now reads the RTC BACK and reports what it holds, rather than asserting success — a 0
from the syscall means the write was issued, not that the clock kept it, and the old message
claimed the second while knowing only the first. On this machine the RTC does not survive a
power cycle, so that reassurance was the only thing standing between the user and the truth.
It also refuses a time before the image was built: the value arrives over plain HTTP and is
unauthenticated by construction, and a clock set BACKWARDS revives expired certificates.

docs/: the audit with every finding classified and cited, the architecture map, an HTTPS
document, a terminal-browser document, and docs/linux-cross-reference.md — which exists because
the RX ring bug cost most of a session when one diff against the reference implementation would
have found it. It records the licence boundary that makes that legitimate: read Linux for facts,
never transcribe code.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
poll checked for a pending signal; the four socket loops never did. A process parked in a
socket read could not be interrupted by anything — kill it and nothing happened until the fd
happened to wake up. It is also the stated blocker on the POSIX floor's Ladybird gate 3, since
an event loop is built precisely on being woken by a signal.

signal_would_interrupt() already existed and was already right — it honours DISPOSITION, so an
ignored signal does not interrupt. That subtlety matters: SIGCHLD defaults to ignore and arrives
whenever a child exits, so the naive `sigpending & !sigmask` test would make every blocking call
return a spurious EINTR in any program that spawns processes, which reads as random I/O failure.
The helper simply had two callers instead of six.

★ Where the check goes is the interesting part. It sits immediately before the yield, and both
the read and write loops return on the FIRST successful transfer — so at that point the
transferred count is zero and EINTR cannot lose data. POSIX requires that a call which has
already moved bytes report the count rather than the interruption, and this placement satisfies
it by construction rather than by handling a partial-transfer case.

Two cases are not clean, and are documented rather than papered over. connect() leaves the
handshake RUNNING — POSIX says so and smoltcp agrees, the socket stays in SynSent — so the
caller must poll for writability rather than re-connect; it is reported anyway, because a
connect that cannot be interrupted is a process that cannot be killed for the kernel's full 10 s
deadline. And sys_dns_resolve returns a packed address with 0 for failure and has no errno
channel, so it cancels the query and reports failure; cancelling is not optional, since a query
left Pending is not merely a leaked slot — smoltcp retransmits it forever and every poll walks
the list.

⚠️ This made ErrorKind::Interrupted reachable in userspace for the first time, and libs/net would
have treated it as fatal — turning any signal into a failed page load. would_block() and the
five blocking read loops in http.rs now retry on it, which is what every Read implementation
does for exactly this reason.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
`SockStats` is created in `Transport::connect` and lives with the
connection, so a kept-alive socket carries the previous request's
totals. `silent_too_long` tested `read_bytes == 0`, which on a reused
connection is false forever — the ten-second watchdog was disabled on
exactly the connections most likely to be half-dead. A server that
closed its end while idle then produced no bytes and no EOF, and the
fetch ran to the 45 s whole-request deadline instead of being abandoned
after ten and retried on a fresh socket. That is the first of the three
back-to-back fetches timing out at exactly 45 s.

Record `(read, write)` when a request installs a transport — both the
fresh path and the `take_idle` path — and difference against it. The
same baseline makes `diagnostic()` describe this request rather than
everything the connection has ever served, and the line now says
REUSED, which nothing did before: three fetches produced no evidence
at all about the feature under suspicion.

Add `keepalive on|off`. There is no QEMU here, so settling whether
reuse helps by building it twice costs two power cycles; a runtime
switch makes it four commands in one boot. Switching off also drops
the held connection, so it takes effect now rather than one request
late.

Also drop `connect_retries`, dead since candidate addresses replaced
re-resolution.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016a5boZV2ZD4iBvGkt4PgH9
…al hardware

Nyx modelled two compute substrates implicitly — the CPU and the Intel GPU — but
neither was a *resource*; each was a bespoke `static Mutex<Option<Driver>>` with
hand-written syscalls. This adds the QPU as a discoverable one, and then actually
runs a circuit on a physical quantum processor.

The audit that started this found the thing that shaped everything after it: Nyx
ALREADY had a complete quantum stack. `tools/compiler` is a real compiler (.ql →
QIR → OpenQASM) with a dependency-free state-vector evaluator already running
on-device in apps/qcstudio, and libs/meridian has host-tested circuit-diagram
geometry. So the circuit representation and the simulator existed. What did not
exist was the OS model. tools/compiler is not modified by any of this.

  libs/quantum      no_std, ZERO deps: device model, circuit IR, backend trait,
                    state-vector sim (a port of statevector.rs), credential parsing
  libs/json         minimal reader/writer, no_std, zero deps, explicit depth cap
  libs/quantum-rt   QpuSession, qclang adapter, IonQ + IBM providers
  nyx-kernel        quantum.rs: registry + PCI probe + syscalls 573/574
  libs/net          Request/request_once — POST, headers, bodies
  terminal          `quantum` commands + a reusable Picker component
  settings          a Quantum pane; sysmon a Quantum disclosure row

HONESTY IS A TYPE, NOT A CONVENTION

QpuStatus::Hardware is produced in exactly one place in the tree — the kernel's
PCI probe, against a device table that is deliberately EMPTY. SimBackend::info()
hardcodes Simulator with no setter. QpuSession::backend_for has no Hardware arm,
so an attached-but-driverless QPU refuses rather than falling back to a simulator.

`Remote` alone would have been a lie: providers serve cloud simulators through the
same API, lifecycle and JSON as real processors, so remote_is_simulator exists and
REMOTE/hw vs REMOTE/sim is rendered distinctly everywhere. Likewise Readout is an
enum with no conversion between Counts and Probabilities — IonQ never says how many
shots produced each bucket, and 0.5 × 1024 = "512" would invent a measurement.

The port is cross-checked against qclang_compiler::statevector on every cargo test:
same program, both evaluators, agreement to 1e-12. The original has been producing
histograms on real hardware, so if they disagree the port is wrong.

FOUR PRE-EXISTING NETWORK BUGS, FOUND BY DRIVING A LIVE ENDPOINT

  * smoltcp's DNS_MAX_RESULT_COUNT defaults to 1, so the multi-address fallback
    (syscall 572) had NEVER worked — the resolver discarded every alternative
    before the kernel saw it.
  * DNS_MAX_SERVER_COUNT defaults to 1 and update_servers PANICS above capacity.
    Most home routers advertise two DNS servers. Now raised AND clamped at both
    call sites: raising a limit is not a fix when the count comes from the network.
  * resolve_host took candidates[0] with no fallback on the one-shot path.
  * DNS is UDP: a lost datagram is recovered by sending another, not by waiting.
    Retries with a fresh query (3 x 4s) instead of one long timeout.

And one of mine: quantum remote run polled to completion inside on_key, freezing
the window for the whole job while the progress messages were never drawn — the
exact mistake the `Loading` doc comment warns about. Now pumped from update().

IBM: OpenQASM 2.0 (not 3 — their loader rejects 3, and QASM2 has no $N hardware
qubit syntax, so a transpiled circuit is qreg q[full-width] indexed physically),
Primitives V2 (the version field is separate from the already-V2-shaped pubs, and
its absence silently selects V1), IAM token exchange, Service-CRN, and job
re-attach so a link failure cannot strand a result already paid for.

Credentials are baked into the image at build time from a gitignored file: Nyx has
no clipboard and cannot express a paste chord, and an IBM CRN is ~120 characters.

Verified: 609 host tests. On hardware, `quantum run bell` gives 525/499 across 00
and 11 with nothing in 01/10; and a Bell pair executed on ibm_fez.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M751uJzRrhAm7uatcHr7hr
@Asmodeus14
Asmodeus14 merged commit 0a3de6a into master Sep 22, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant