Firmware that makes an RP2350-based One ROM board behave like a Soviet 1801RE2 mask ROM — a chip that sits on the MPI bus, the domestic equivalent of DEC's Q-bus, where address and data share one set of sixteen lines.
Target machine: Elektronika MS 0511 (UKNC), which uses four of them.
Status: working. An Elektronika MS 0511 boots with all four of its 1801RE2 mask ROMs replaced by a single One ROM Fire 24 in the DS4 socket. The image conversion tooling is finished and tested. Read Before you plug anything in first, and Prior art before deciding this is the right project at all — someone has already built a purpose-made board for this job.
One ROM is built around one shape of chip: address lines in, data lines out,
chip select gating the drivers. The v0.7 firmware makes that explicit — serving
is decomposed into three families of PIO algorithm (CS*, ADDR*, DATA*),
selected per ROM type by the Rust pre-processor, and every one of them assumes
address and data occupy separate pins. ADDR0 reads a block of address pins and
uses the value to index a table; DATA0 puts a word from a FIFO onto the data
pins. There is nothing to configure that makes those two sets of pins the same
set of pins.
The plugin API is not a way in either. Plugins run on the CPU cores alongside
PIO serving, and the API (firmware/ora/api.h) exposes logging, memory, USB,
IRQs, ROM slot access and an address monitor — but no way to claim the socket
GPIOs or replace the serving state machines. A plugin can watch the bus; it
cannot drive it.
So this is custom firmware that treats the Fire 24 as a well-documented RP2350 carrier board with a known socket-to-GPIO map, not an extension of One ROM. Upstreaming it later as a fourth algorithm family would be reasonable; starting there would not.
The 1801RE2 is a 4K × 16 mask ROM. Internally it latches a 12-bit word address
from nAD1–nAD12 and compares nAD13–nAD15 against a mask-programmed three-bit
value — the "code" — that fixes which 8 KB window of the 64 KB address space it
answers for. It drives data and asserts nRPLY only on a match. Everything on the
bus is inverted, hence the n prefixes.
| Code | Window (octal) |
|---|---|
| 000 | 160000–177777 |
| 001 | 140000–157777 |
| 010 | 120000–137777 |
| 011 | 100000–117777 |
| 100 | 060000–077777 |
| 101 | 040000–057777 |
| 110 | 020000–037777 |
| 111 | 000000–017777 |
A read cycle runs: the host puts the address on the AD lines and asserts nSYNC; the slave latches it; the host releases the AD lines and asserts nDIN; the slave drives data and pulls nRPLY low; the host takes the data and releases nDIN; the slave releases the bus; nSYNC releases.
From the KR1801RE2 datasheet, table 11.26 and figure 11.30. Power follows JEDEC and matches what the Fire 24 hard-wires, so the board drops in unmodified.
| pin | signal | pin | signal | pin | signal | ||
|---|---|---|---|---|---|---|---|
| 1 | RD (nDIN) | 9 | AD9 | 17 | AD12 | ||
| 2 | AN (nRPLY) | 10 | AD10 | 18 | AD13 | ||
| 3 | SYN (nSYNC) | 11 | AD11 | 19 | AD14 | ||
| 4 | AD4 | 12 | GND | 20 | AD15 | ||
| 5 | AD5 | 13 | AD3 | 21 | n/c | ||
| 6 | AD6 | 14 | AD2 | 22 | n/c | ||
| 7 | AD7 | 15 | AD1 | 23 | CS | ||
| 8 | AD8 | 16 | AD0 | 24 | Ucc |
Everything is active low, which is why the k1801 RTL names these nAD, nSYNC, nDIN and nRPLY. Two things are worth pulling out:
There is a chip select on pin 23, which is how the UKNC banks a window out — see below. The design had assumed such a pin had to exist; it does.
AN on pin 2 is the reply — nRPLY, the signal the ROM asserts to complete a transfer. Table 11.26's "вход" against it is a misprint; figure 11.30 draws it on the output side. It is driven here, open-drain: the output value is always low and only the direction is toggled, so the line is either pulled down or released to the bus pull-up, never driven high.
The MS 0511 has two K1801VM2 processors — a central one at 8 MHz and a peripheral one at 6.25 MHz — and the ROMs are on the peripheral processor's bus, not the central one's. That is where any scope probing has to happen.
The PP address space, and where the four chips land in it:
| Range (octal) | Contents | Chip |
|---|---|---|
| 000000–077777 | PP RAM, 32 KB | |
| 100000–117777 | banked window | 1801RE2-205, code 011 |
| 120000–137777 | system ROM | 1801RE2-206, code 010 |
| 140000–157777 | system ROM | 1801RE2-207, code 001 |
| 160000–176777 | system ROM | 1801RE2-208, code 000 |
| 177000–177777 | I/O page |
Two consequences fall straight out of that table, and both are the difference between working firmware and a board that fights the machine for its own bus.
Every window can be banked away. Port 177054 decides, per window, whether the PP sees ROM or RAM: bit 5 for 120000, bit 6 for 140000, bit 7 for 160000, and bits 0–4 for the 100000 window, which can also be switched to one of six banks of external cartridge ROM. A mask ROM has no logic to do that itself, so something external must gate it. Sheet 1 of the schematic shows how, and it is not what you would guess:
- Pin 1 is fed by EDIN, not by the raw K1DIN net. EDIN comes off the output side of the CGM (D10, pin 53) — a read strobe already qualified by the banking state. A window switched to RAM simply never strobes its ROM. This is the real per-window gate, and it means the gating arrives for free: the firmware waits on the read strobe and never hears one it should not answer.
- CS on pin 23 is strapped to ground on DS1, DS2 and DS3 — permanently selected. Only DS4, the 205 covering the switchable 100000 window, has it driven, from the CGM's CE0. So CS arbitrates one window, not four. It is also what settles the polarity question: grounded means selected, so active low.
The consequence for the firmware is in Serving, and it is the reason the response machine is re-armed every cycle.
The code 0 chip overlaps the I/O page. Its window is 160000–177777 but only
160000–176777 is ROM; the top 512 bytes belong to the machine's registers.
gen_rom_images.py therefore serves only 3840 of the 4096 words for a code 0
image by default, and --full-window overrides that if you ever need it.
Because the chip decodes its own window from nAD13–nAD15, one board in any of the four sockets can answer for all four. Every socket carries the same 1AD bus, the same K1SYNC, the same EDIN and the same RPLY. Only CS differs, and that decides how much care is needed:
- DS1, DS2 or DS3 — CS strapped to ground, so the board is permanently
selected and per-window banking comes entirely from EDIN. Nothing to configure
beyond
SOCKET_CS_CODE 0xFF. - DS4 — CS is CE0, which means "the 100000 window belongs to the on-board
ROM rather than to RAM or a cartridge". It says nothing about the other three
windows, which have no CE and are gated by EDIN alone. Set
SOCKET_CS_CODEto03so the check is scoped to that window. Applying CS globally would let one deasserted CE silence three windows it has no authority over — the whole system ROM disappearing whenever software banked something into 100000.
Whichever socket you use, every original it answers for must come out. Two devices driving the same RPLY and the same AD lines is a bus fight, and the 1801RE2 has no idea it has been replaced. If only DS4 is socketed and the other three are still soldered down, generate an image set containing only the 205 — that is a straight one-for-one replacement and needs no desoldering.
The PP has no spare address space: 000000-077777 is RAM, 120000-176777 is the three fixed ROM windows, 177000 up is I/O. The only place more ROM can go is the 100000 window, which port 177054 switches between PP RAM, the on-board 205, and cartridge banks — two slots of three 8 KB banks each.
The bus really is shared. XS1 carries the same 1AD lines, K1SYNC, K1RPLY and K1DIN as the ROM sockets, plus all four chip enables: CE0 on B12, CE1 on B14, CE2 on B13, CE3 on A12. A cartridge is not on a separate bus, it is on this one with its own enable.
What the DS4 socket does not carry is which bank is selected. Pin 23 is CE0 alone, meaning "the on-board ROM owns this window" — so when software switches to a cartridge, CE0 deasserts and the board correctly goes quiet, but it has no way to know that bank 2 of slot 1 is now wanted. Serving cartridge banks from this socket therefore needs CE1, CE2 and CE3 brought in by wire.
There are pins for it. Socket pins 21 and 22 are not connected in the machine and land on GPIO 12 and 14, and the X1/X2 jumper pads give GPIO 9 and 8 — four spare inputs for three enables. RAM is not the constraint either: eight windows already cost 128 KB of the RP2350's 520 KB, and six more banks would add 96 KB.
One thing to establish first: whether EDIN asserts for reads the machine has directed at a cartridge, or only for the on-board ROM. If it is qualified by CE0 as well, the read strobe never arrives for a cartridge access and K1DIN has to be wired in too. A scope on pin 1 while software selects a cartridge bank settles it.
The alternative is to stop fighting the socket and build into a cartridge, which carries the bus, the strobes and the enables by design. That trades three bodge wires for a mechanical adapter from a 24-pin DIP footprint.
An MS 0511 running this ROM set should show a boot menu — ЗАГРУЗКА, with disk, ROM cartridge, network, C2, tape, debug and тестирование as options. The emulator draws it from the identical 32 KB image, so the menu is unquestionably in the ROM the board is serving. A machine that reaches a cursor and answers УСТ but never shows the menu is therefore taking a different branch, not missing code.
Disassembled with tools/pdp11dis.py, the entry at 160300 reads:
160300 013704 172660 mov @#172660, r4 ; 172660 is in ROM: 000450
160304 005000 clr r0
160306 010406 mov r4, sp
160310 100465 bmi 160464 ; warm-start check
160312 032737 000020 177716 bit #20, @#177716
160320 001404 beq 160332 ; bit 4 clear -> cold start
160322 013700 000000 mov @#0, r0 ; bit 4 set: restart vector
160326 001401 beq 160332
160330 000110 jmp (r0)
160332 012737 000040 177716 mov #40, @#177716 ; hold the CPU in reset
160340 004767 012706 jsr pc, 173252 ; load the CPU's planes
160344 012737 070045 177010 mov #70045, @#177010
160352 016437 000042 177014 mov 42(r4), @#177014
160360 005037 177716 clr @#177716
160364 012700 000100 mov #100, r0
160370 077001 sob r0, 160370 ; settle
160372 012737 100000 177716 mov #100000, @#177716 ; release the CPU
160400 004767 000004 jsr pc, 160410 ; checksum all four ROMs
Two things in the previous version of this section were wrong and are corrected
here. 100465 is bmi, not bpl. And bit 4 of 177716 is not a strap: writing
177716 drives the central processor's control lines — bit 4 is HALT, bit 5
DCLO, bit 15 ACLO — which is exactly what the sequence above is doing when it
writes 40, then 0, then 100000. The bit-4 branch is a warm-restart hook that
jumps through location 0 if one is set; it has nothing to do with the menu, and
the guess that it explained the missing menu was wrong.
The first frame read back from hardware was:
| pulse | reading | |
|---|---|---|
| 1 | long | the monitor started from our vector — the instrument is valid |
| 2 | long | a ROM block failed its checksum |
| 3 | short | the PP RAM test found no fault |
| 4 | short | the error-printing routine was not entered |
| 5 | long | the startup test ran to completion |
| 6 | short | the boot menu header was never printed |
Pulse 2 is the one that matters, and it reverses the conclusion below. The images pass that same checksum offline — byte-identical to the reference, all four blocks verified — so the machine reading a bad block means it is not receiving what the board holds. That is a board-side fault, and it is not confined to the checksum: a processor fed a wrong word executes it. The missing menu may be downstream of this rather than a separate question.
Pulses 2 and 4 together are informative. A checksum failure leaves a bit set in
the error mask, and the monitor at 172732 branches to the printing routine when
that mask is non-zero — so pulse 4 should have been long too. The path that
skips it is the bhi at 172726, taken when the byte the central processor sent
over channel 0 (port 177060) is greater than 2. So the CPU is reporting
something as well. Pulses can only be lost, never invented, so this is a lead
rather than a proof.
There are exactly two ways our correct data becomes the processor's wrong sum:
reads we decline to answer, and replies we assemble that are never taken. The
watch build now scores both as pulses 7 and 8, and -DMPI_IGNORE_CS=ON builds
a twin that answers every window unconditionally — two firmwares differing in
one variable, which is what turns this from an argument into a measurement.
A scope on socket pin 2 in the running machine shows nRPLY doing something a reply line should not: sharp falls, then an RC ramp taking on the order of a microsecond to climb back, and at the cycle rate the machine actually runs at (the capture reads ~2.5 µs between cycles) several ramps are cut off by the next assertion before they arrive anywhere near a logic high. Peak is 4.24 V, so the level is fine — the edge is not.
That is the signature of a line released to hi-Z against a weak pull-up and a few hundred pF, which is exactly what the firmware does: nRPLY is open drain by construction, driven low to assert and released to let the bus pull it up. Open drain is the right model for an MPI reply line, but it assumes the bus can restore the line quickly, and here it plainly cannot.
The consequence fits the fault. A reply line that has not finished coming back up is indistinguishable from one still asserted. A host reading it that way takes a reply for a cycle nobody has driven yet and samples the AD lines early — data that is right in the board and wrong in the processor, intermittently, which is what the checksum reported.
-DMPI_RPLY_ASSIST=ON selects the mpi_respond_assist program, which drives
nRPLY high with the pad's own 8 mA for ~320 ns after the host takes the data
before going back to hi-Z. The window is bounded on purpose: nRPLY is shared,
every slave asserts it for its own cycles, and the drive has to be gone before
the next one could. From the observed ramp the line looks like a few hundred pF,
which 8 mA slews in 100–200 ns, so 320 ns does it with margin and still fits
inside the shortest turnaround the PP could produce at 6.25 MHz.
Pin 2 with the board removed sits at a solid 5 V. Three consequences, one of which cuts against the section above:
There is a real pull-up. The machine restores its own reply line, so open drain is the correct model and the assist is a workaround for a slow rise, not a correction of a wrong model. An earlier note here suggested that a line which drifted when unloaded would mean the machine expects its slaves to drive high. It does not drift, so that reading is off the table.
The board is loading the top of the swing. Unloaded the line reaches 5.0 V; with the board fitted the ramps peak at 4.24 V. That 0.76 V appears only when we are in the socket, and 4.24 V is about what a 3V3 rail plus a diode drop looks like — the signature of the pad's clamp conducting, which is what a non-5V- tolerant input on a 5 V bus does. The current involved is well under a milliamp against any sane pull-up, so it is a loading effect rather than a hazard, but it is why the ramp flattens as it approaches the top instead of arriving.
It weakens the timing argument. An RC curve crosses a TTL VIH of 2.0 V early in its rise — well before the flattening that the clamp causes — so the line probably does reach a valid high in time more often than the shape suggests. The "still looks asserted" mechanism is therefore a plausible lead, not an established cause.
What settles it is a two-channel capture: nRPLY on pin 2 against nDIN on pin 1. If nRPLY is reliably above threshold before nDIN next falls, the reply line is exonerated and the corruption is elsewhere. If it is not, the mechanism is confirmed and the assist is the fix. One capture decides it.
A two-channel capture with the plain build — nRPLY on pin 2 against nDIN on pin 1 — reads nRPLY 3.92 V peak, 39% high, against a clean 5.6 V nDIN. 3.92 V is 3V3 plus a diode drop, so the pad clamp is conducting exactly as expected, but it is a comfortable TTL high and the machine reads it fine.
The frame settles it more firmly than the waveform can. Pulse 8 was short: no reply we prepared ever went untaken. The handshake completes on every cycle. Whatever the reply line looks like, it is not breaking transfers, and the assist is not the fix. Ugly is not the same as broken.
That frame — long on 1, 2 and 5, short on everything else — also puts the board in an odd position: we answer every read, on time, every reply is taken, startup runs to completion, and a ROM block still fails its checksum.
Flashed on hardware it made the machine worse — the screen never reached the cursor, the video memory was not being initialised at all — and the frame read long on pulse 1, long on pulse 8, short on everything else.
The cause is worth recording because it is a PIO trap rather than a bus
subtlety. SET maps one pin. With the SET base at nRPLY, set pindirs, 0
releases nRPLY and nothing else, so the sixteen AD lines stayed driven until the
.wrap — which the assist had just pushed ~320 ns later. Any cycle the host
opened inside that window met our drivers on every address line. That is
contention on the address, not a timing effect, and it is exactly as destructive
as it sounds.
The IRQ was wrong too. It sat after the delay, so the CPU's served-versus-missed check at the next address strobe ran before the state machine had raised it, scored a completed cycle as missed, and re-armed the response machine underneath itself. Pulse 8 in that frame is mostly this, which makes that run uninformative about the untaken replies it appears to report.
Both are fixed: the IRQ is raised immediately after the host takes the data, at the same point in the cycle as the plain program, and a third direction mask — nRPLY driven, AD released — is preloaded into ISR so the AD lines are let go two instructions after the strobe while nRPLY alone is held high.
The lesson generalises: an instrumented build has to be timing-identical to the one it is measuring, and OUT and SET see different pins.
With a per-boot frame that could be trusted, the picture became a contradiction:
| pulse | ||
|---|---|---|
| 1 | long | the monitor started from our vector |
| 2 | long | a ROM block failed its checksum |
| 4 | short | no reply went untaken |
| 5 | short | the bus carried exactly what we drove, every cycle |
| 6 | long | we were listening from the machine's first cycle |
| 7 | long | all four windows fully covered |
Every word came from us, every reply was taken, the wire carried what we sent — and the checksum still failed. When every measurement says the data is right and the processor says it is wrong, the measurements are answering the wrong question.
They were. Look at what the response program did:
out pindirs, 24 ; drive the AD lines
mov osr, y
out pindirs, 24 ; assert nRPLY -- handshake complete
Two instructions. 13 ns at 150 MHz. nRPLY is the slave saying the data is on the lines, and this program said it 13 ns after starting to drive sixteen lines into a capacitive 5 V bus — the same bus whose reply line was measured needing the better part of a microsecond to cross a threshold. The AD lines are no faster. The host was being invited to latch data that was still on its way.
It fits every symptom exactly. A line already at the right level is correct immediately, so most words are fine and the machine runs; a word needing a long transition somewhere is not, so a sum over 16127 of them fails. The boot menu appears when the dice fall well and does not when they do not.
And it explains why the readback saw nothing. That sample sat at the end of the cycle, after the host had released the strobe — the most forgiving instant there is, by which time every line has long settled. Pulse 5 was dark because it was measuring the wrong moment, not because the data was good.
So both halves change. The program now waits ~430 ns after driving the lines before asserting the reply, and takes its readback at the instant it asserts — the earliest the host could latch, and therefore the only instant worth checking. Being late costs wait states rather than data, which is the whole reason a CPU-in-the-loop ROM is viable on this bus; 430 ns against a 2.4 µs cycle is margin bought at no real price.
It changed nothing. The frame came back identical. The setup time is kept because a slave asserting its reply before its data is valid is wrong however the machine behaves, and the cost is wait states — but it is not the fault, and that is the second mechanism proposed here that the hardware has refused.
Which is the point at which guessing at mechanisms should stop. The failure is
also intermittent: boots that print a CPU or CPU-RAM error with no
- ОШИБКА ПЗУ beside them are boots where all four blocks verified, and the
monitor prints the ROM line whenever the mask is non-zero. So the next thing to
establish is not why but which — see pulses 8–9.
Worth stating plainly, because it constrains everything above.
The peripheral processor is the only one with a real reset. It takes DCLO
and ACLO from the machine's power-on circuitry, and a 1801 starts on the
falling edge of ACLO with DCLO already low — an edge, not a level. In the
emulator that is the whole of CMotherboard::Reset(): assert both pins on the
PPU, clear the peripherals, release both. Nothing else in the machine is reset
by hardware.
The central processor has no reset of its own at all. Its DCLO, ACLO and
HALT pins are driven exclusively by the PP writing port 177716 — bit 5 is DCLO,
bit 15 is ACLO, bit 4 is HALT. Reset() never touches them. So the CPU starts
if and only if the PP executes this, out of our ROM:
160332 mov #40, @#177716 hold it: DCLO asserted
160340 jsr pc, 173252 load its memory through the plane ports
160360 clr @#177716 release DCLO
160364 mov #100, r0 / sob settle, a few hundred microseconds
160372 mov #100000, @#177716 release ACLO -- the edge that starts it
Two consequences.
The CPU's power-on is five instructions read from this board. A single word
misread anywhere in 160332–160376 and the central processor is never started, or
started with the wrong pin sequence, or started before its memory was loaded.
That is a direct route from "an occasional bad ROM read" to - ОШИБКА ЦП, and
it does not require the CPU or its RAM to be faulty at all. Pulses 3 and 4 watch
the two ends of it, which separates the CPU was never started from the CPU was
started and failed — otherwise pure guesswork.
The PP's own start is an edge from an ageing supervisor. If that circuit releases ACLO before the rails have settled, the PP starts erratically, and the symptom is intermittent trouble that clears on a manual reset — which is the pattern this machine has shown throughout, including the one boot that reached the menu. It is also indistinguishable, from the outside, from our own startup race: the board must be serving before that edge arrives, and nothing in the firmware can outrun the RP2350 bootrom. Pulse 8 tells the two apart, because it is short only when the machine asked before we were listening.
So yes — it could be a contributor, and it is worth a scope on the reset line at power-on to see whether ACLO comes up cleanly or chatters. A supervisor that retriggers would show as the frame's bits appearing and vanishing between passes, since each start clears them.
- ОШИБКА ОЗУ ЦП is the machine saying its central processor's memory is bad,
and that is worth testing properly — which the stock test structurally cannot
do. It runs on the central processor, using code that was itself copied
into the memory under test. A fault there corrupts the tester before it can
report on the testee, and every result it produces is suspect.
The peripheral processor can do the job instead, and this matters more than it sounds because of how the memory is wired:
The central processor's RAM is planes 1 and 2. Its word at address A is plane 1 byte A/2 in the low half and plane 2 byte A/2 in the high half — which is exactly the format of port 177014. So the PP can reach every location of it through the plane registers, with the central processor held in reset, and the two halves of what comes back name which plane is at fault.
That last part is the useful bit. If a machine has three RAM banks and one of
them is the PP's own (plane 0, which the monitor already tests and passes),
then a - ОШИБКА ОЗУ ЦП points at the other two — and a bank that tests fine
out of circuit may simply be the one that was never implicated.
$ ./tools/make_ramtest.py -o ramtest.bin
$ ./tools/gen_rom_images.py --logical ramtest.bin -o firmware/rom_images.c
$ cmake -S firmware -B firmware/build-ramtest -G Ninja -DMPI_BEACONS=ONThe test holds the CPU in reset for its whole run, checks PP RAM first so a fault there is not mistaken for a plane fault, then walks both planes twice — each location holding its address and then the complement, so every bit takes both values everywhere. It accumulates the XOR of what came back against what went in, so the result holds exactly the bits that were ever wrong.
Results come back as beacons, and -DMPI_BEACONS=ON blinks them. Two things
about that build are worth knowing, because both were got wrong first:
MPI_BEACONS needs MPI_WATCH, and used not to say so. Beacon counting and
the LED frame both live inside #if MPI_WATCH, so -DMPI_BEACONS=ON alone
produced a firmware that ignored every beacon and ran the ordinary status LED —
which on hardware looks like the board blinking away busily while reporting
nothing whatsoever. It now implies it, resolved before either reaches the
compiler rather than by defining MPI_WATCH twice and trusting the later flag
to win.
The beacons live in ROM, not RAM. A beacon works by being an address the board sees go past. At 077700 in PP RAM that depends on the capture machine latching cycles for addresses we do not serve — probably true, but never demonstrated: every address this project has confirmed seeing has been one of ours. At 176700 it needs no assumption at all, because we answer the read ourselves. Reading ROM is harmless; only the address matters.
| pulse | |
|---|---|
| 1 | alive — the test is running |
| 2 | PP RAM (plane 0) passed |
| 3 | PP RAM failed |
| 4 | planes 1 and 2 passed |
| 5 | plane 1 failed |
| 6 | plane 2 failed |
| 7 | finished |
| 8 | the bit was already wrong on an immediate reread — dead, not leaky |
| 9–16 | bit 0…7 of the failing plane's byte |
| 17 | every bit position failed within a single pass — not a chip, a subsystem |
The dead bit is fixed and the machine boots, runs its own diagnostic and takes keyboard input — but it still degrades as it warms, which a hard stuck bit cannot do. So there is a second fault, softer than the first, and it needs a different kind of measurement: the machine hot, and something still asking.
The plane test now soaks rather than parking. It loops for as long as the machine is left on, and because the firmware's beacon bits are set-only, a bit that fails on pass four hundred lights its pulse and stays lit. Leave it running an hour and read the frame afterwards.
Warm, the machine now freezes, blanks, or falls back to the uninitialised vertical lines after about ten minutes of sitting at the boot screen. All three are plane 0 symptoms rather than data corruption: the video tag list lives at 0000270 in plane 0, which is also the PP's own RAM, so a fault there makes the display stop making sense and the dispatcher stop dispatching. The soak tests plane 0 as its first phase, and now names its failing bits too.
That last part exposed a flaw the simulator caught and hardware would not have: the accumulator recording PP RAM faults was itself in PP RAM, so a bad bit erased the record of itself and the test reported a clean pass over a faulty bank. It lives in r6 now — the stack pointer, free because this program never uses a stack and takes no traps with interrupts masked. The witness to a memory fault cannot live in the memory under test.
One combination is worth reading deliberately: on an intermittent fault both
"planes passed" and "plane N failed" end up lit, because over hundreds of
passes both happened. A hard fault cannot produce that pairing, so the frame
distinguishes "always broken" from "sometimes broken" without any timing
information at all. test/test_ramtest.py demonstrates it with a fault present
on only one pass in three, rather than leaving it as a claim.
And a clean frame after a long hot soak is just as useful: it would put the thermal fault outside all three RAM planes, which is most of what is easy to suspect.
Two frames, minutes apart, from a cold start:
l l s l s s l s s s s s s s s s s healthy: alive, PP RAM ok, planes ok, done
l l l l l l l l l l l l l l l l l everything, including pulses 8 and 17
There is no third reading between them. That absence is the finding. A weak cell warming past its retention limit fails one bit, in one plane, and the frame grows a pulse at a time; this goes from a completely clean pass to every bit of every plane, plus pulse 8 (wrong on an immediate reread, so microseconds after the write, which no retention failure can reach) and pulse 17 (all eight bit positions lost inside a single pass).
Three banks on two different physical groups of chips do not degrade in unison. Whatever they share does — the RAS/CAS timing and refresh generator that drives all 24 chips, or a supply local to the array. Both of those fail as a step, and that is exactly the shape of the reading.
The latching frame answers "does this machine ever fail". It cannot answer "did
cooling this chip just fix it", because it is built not to: g_beacons is
set-only, cleared only when the PP restarts, and the ROM's accumulators sit
outside its loop. A fault lit on pass four hundred stays lit however cold the
guilty part gets. For chasing a thermal fault with a can of freeze spray that
makes it worse than useless — it will report a fault for as long as the machine
is powered, no matter what you do to the board.
So there is a second pairing that reports only the most recent pass:
$ ./tools/make_ramtest.py --live -o ramsoak-live.bin
$ ./tools/gen_rom_images.py --logical ramsoak-live.bin -o firmware/rom_images.c
$ cmake -S firmware -B firmware/build-live -G Ninja -DMPI_BEACON_LIVE=ON
$ cmake --build firmware/build-live--live moves the clears inside the loop, so each pass starts from nothing;
-DMPI_BEACON_LIVE=ON makes core 1 snapshot and clear the beacons on the DONE
beacon, which is the pass boundary. Both halves are needed — either alone still
latches, one because the ROM keeps re-asserting a mask the firmware clears, the
other because the firmware keeps displaying a fault the ROM has forgotten.
The LED stops being a frame and becomes a lamp:
| steady on | the last pass failed |
| dark, with a brief blip each pass | the last pass was clean |
| fast flicker | no pass has completed in fifteen seconds — the PP is stuck |
A pass takes a few seconds, so the lamp follows the machine closely enough to
spray one chip and watch. test/test_ramtest.py checks the property this whole
mode exists for: a fault present on pass 2 only must leave pass 3 reporting
clean, because otherwise cooling the guilty chip looks identical to cooling an
innocent one.
Six single-shot captures of RAS and CAS at the DRAM pins, in both states, and they cost three wrong conclusions before producing a right one. The waveform shapes changed between captures — CAS ramping in one, square in another — and each time the shape looked like the answer. It was not: those are Single captures of a signal whose content varies cycle to cycle, twelve cycles out of a continuous stream, and whichever cycle the trigger landed on is what you get. Two of the captures were relabelled mid-investigation, in opposite directions, and the "obvious" reading flipped with them both times.
What settled it was picking a statistic instead of a shape. With CAS on CH1 so its Vrms is reported:
| healthy | failed | |
|---|---|---|
| CAS Vrms | 3.84 V | 3.92 V |
| RAS | 2.38 MHz, 40.0%, 6.00 Vp-p | 2.38 MHz, 40.0%, 5.84 Vp-p |
Two percent, running the wrong way for any story. The strobes do not change
when the machine fails. Frequency and duty at that timebase are unusable —
Duty+: 100.0% means the scope never found a falling edge on the ramp — which
is its own lesson about which numbers on a cheap scope are load-bearing.
The rule that would have saved all of it: take several captures in the same state before comparing states, so you know what normal variation looks like.
Every bit of every plane failing at once is not a statement about chips. Nothing true of eight independent DRAMs in two separate banks is true of all of them in the same instant; what they share is the address bus and the strobes, and the strobes had just been ruled out.
tools/make_addrtest.py asks the other question. Three passes over planes 1 and
2, with the CPU held in reset:
$ ./tools/make_addrtest.py -o addrtest.bin
$ ./tools/gen_rom_images.py --logical addrtest.bin -o firmware/rom_images.c
$ cmake -S firmware -B firmware/build-addr -G Ninja -DMPI_BEACONS=ON -DMPI_BEACON_COUNT=19| pass | fill | sees |
|---|---|---|
| 1 | every cell 0 |
data lines stuck high |
| 2 | every cell 177777 |
data lines stuck low |
| 3 | every cell its own index | addressing and data |
The first two are structurally blind to addressing: every cell holds the same value, so a read that lands on the wrong cell still returns the right answer. That makes them a pure data-line test. The third is sensitive to both, and it carries more than pass/fail — with each cell holding its own address, whatever comes back is the address of the cell that actually got selected, so the XOR against what was asked for names the address bits that went wrong.
Bits the constant passes implicated are masked out of the address report. A data
line stuck at 0 differs from its index in that position too, and calling that an
address fault would point at the wrong half of the board. test/test_addrtest.py
holds it to that with a dead data bit, a dead address line, and both at once.
| pulse | |
|---|---|
| 1 | alive |
| 2 | pass finished |
| 3 | data lines bad |
| 4 | addressing bad |
| 5–19 | plane index bit 0…14 |
The machine's answer, warm:
l l s l s l s l s l s l s l s l s l s
^ ^ bits: 0 1 2 3 4 5 6 7 8 9 ...
| addressing bad
data lines clean
Address bits 1, 3, 5, 7, 9, 11, 13 wrong; 0, 2, 4, 6, 8, 10, 12, 14 clean. Every odd bit, no exceptions, with bit 15 outside the tested range.
Two things follow immediately. The 24 DRAMs are exonerated — both constant passes read back perfectly, so every data line in both planes carries what it is given. And a perfectly alternating pattern is never four independent faults; it is one part. Which part depends on how the schematic maps the plane index onto the DRAMs' eight address pins:
| if row/column is | then the failing bits are |
|---|---|
| index 0–7 / index 8–15 | pins A1, A3, A5, A7 dead in both phases — one package, if the two multiplexers are split odd/even to shorten the routing |
| even index / odd index | the entire column address, with the row address perfect |
The second is the simpler failure, and it fits the rest: the column address is what gets latched on CAS, so a wrong column behind a clean CAS waveform is exactly the combination that was measured. A wrong column selects the wrong cell in all 24 chips at once — which is why all eight bit positions in all three planes go wrong together and immediately, and why the display falls to the uninitialised vertical lines at the same instant.
One measurement separates them: probe a DRAM address pin with the fault present and watch whether it changes value between the RAS phase and the CAS phase. Pins that stop changing at column time, all of them, means the column half of the multiplexer; specific pins dead in both phases means those lines.
The answer is D22, a КР1801ВП1-055 gate array, which drives RPLY on the peripheral processor's bus. Cooling it keeps the machine running indefinitely; cooling anything else -- the PP, D8, D11, the DRAM banks -- changes nothing. That control is the whole result: a single component, repeated, with every other candidate tried the same way.
Not on our instruments. With D22 held cold the MS 0511 boots the stock monitor and runs its own ТЕСТИРОВАНИЕ suite -- ПРОХОД 4, ОШИБОК 0 -- which it had not managed since the thermal fault appeared. Stop cooling and it fails before that pass finishes: about a minute, against the ten minutes the whole board takes to reach the fault from cold.
That gap is the control. A chip that has been chilled returns to ambient in a minute or two while the board takes ten, so failing inside a minute says the part's own temperature is the variable and not the air around it. Cooling every other candidate the same way -- the PP, D8, D11, the DRAM banks -- changed nothing.
It also retroactively clears the memory. The CPU planes pass the machine's own
test now, so the - ОШИБКА ОЗУ ЦП that started this was the dead 4164 on DC7,
and everything after the repair was D22 making good memory look bad.
A 16-bit bidirectional bus transceiver with inversion — the transceiver between the machine's inverted AD bus and the peripheral processor's internal bus, so it sits in the path of every instruction fetched and every word read or written. Its logic diagram is published; the pin table, the reasoning and a substitute built from two 74x640s are in docs/D22-KR1801VP1-055.md.
That it is a pass-through rather than a storage element is the whole explanation for the shape of this investigation: nothing is ever stored wrong, so every memory test, register pattern and retention check came back clean while the machine plainly did not work.
The bypass capacitor. D22's decoupling was pulled and measured: 47 nF, which is a correct decoupling value and in tolerance. Replaced with a modern 100 nF anyway -- no change in behaviour. The board uses the same part at every IC, so there was no odd-one-out to find either. Cleared twice over: the value is right, and improving it does nothing.
(The marking reads 47н. Read from a blurry photograph it looked like 47п,
which would have been 47 pF -- a thousand times too small and a very promising
lead. It was worth chasing and it was worth measuring rather than believing.)
Latched state of any kind. The machine recovers without a power cycle: cool D22 and a soft reset on the front panel button brings it straight back. So nothing is latching -- not latch-up, not a corrupted internal register, nothing that needs power removed to clear. What crosses a threshold recrosses it as soon as the temperature drops, which is what a propagation delay against junction temperature does and what almost nothing else does.
That leaves the die. Worth trying before sourcing a КР1801ВП1-055, in this order, because both are free and reversible:
- supply at 5.2 V. 5 V ±5% puts 5.25 V in spec, and logic of this era gets faster with more supply -- more drive, shorter delay. If the part is marginal by a little, this may buy back more than cooling does.
- a fan, not a heatsink. D22 does not get warm to the touch, so it has almost nothing of its own to shed and a heatsink has nothing to remove; its junction still sits a few degrees above the surrounding air, and moving that air is the only remaining lever.
RPLY is the signal that completes a bus cycle. It carries no data, so nothing it does can be caught by a test that checks values -- and every test in this repository checks values. That is why the list of things measured clean grew so long and so useless:
| measured | result |
|---|---|
| both CPU planes, zeros and ones fills, with a third of a second between write and check | clean, every run |
| the plane address register, alternating and constant patterns | clean |
| RAS and CAS at the DRAM pins, healthy vs failed | CAS Vrms 3.84 V against 3.92 V |
| the data we drive, sampled at our own pads on every served cycle | clean |
| plane 0's eight chips, all replaced | no change |
| D8, D11, the CPU, the ROM socket, all reflowed | no change |
Nothing is ever stored wrong, so every storage test passes. What fails is when the processor is told a cycle is finished.
Every observation in this file's history is a completion failure seen from a different angle:
- hangs. RPLY never arrives, the cycle never ends, and a 1801 has nothing to time it out. The last-address frame froze on 0000246 with no trap, no restart, and no further bus activity -- a cycle simply left open.
- wrong data. RPLY arriving late means the processor latches the bus at the wrong instant, when some lines have settled and others have not. Which bits are wrong then depends on the pattern -- which is very likely the odd-bit "adjacent-line coupling" fingerprint that was reproducible for days and then refused to appear on a test that used the same register with no memory behind it.
- wild jumps. An instruction fetch completed at the wrong moment returns a word that is not the instruction, and the PC goes somewhere arbitrary.
- plane 0 first. DRAM cycles have the least timing margin and register
transfers the most, so as margin erodes the PP's own memory -- reached by
ordinary
mov r0, (r0), the slowest path in the test -- fails while 177010 and 177014 still work. The soak's digit readout showed exactly that ordering: 1, then 2, then 4. - everything else, afterwards. Plane 0 is the PP's own RAM, holding its vectors, its stack and the video tag list. Once it goes the PP corrupts itself, the test running on it reports the CPU planes bad, and the screen falls to the uninitialised vertical lines. The delay before that second stage varied from one pass to seven, which is what a consequence looks like rather than a second independent fault.
The instrument was wrong more often than the machine was surprising. Frames that latched when they should have cleared and cleared when they should have latched; 32 KB of ROM filled with HALT instructions, which turned every excursion the real machine survives into a death; a trap-vector capture that could not represent zero; a per-frame sample of an event one pass wide; and three iterations of an LED format before it could be read reliably at a bench.
The two habits that eventually worked: run the control -- cool the other chips too, or the effect belongs to the board and not the part -- and prefer a number to a shape, because six single-shot scope captures of RAS and CAS produced three confident and contradictory conclusions, and one Vrms reading settled it.
With the chip on DC7 replaced, the plane test passes and the MS 0511 reaches the ЗАГРУЗКА menu, cursor blinking, in 80-column mode.
The whole chain, in the order it was actually established rather than the order it was guessed: the board serves four windows of a Soviet mask ROM over a multiplexed bus; the machine's own monitor said its central processor's memory was bad; a test written for the peripheral processor, run from the ROM socket with the central processor held in reset, said plane 1; then bit 7; then that the bit was dead rather than leaky; and the schematic's data-line names put that bit on one 4164.
Two things worth keeping from it. The board was never at fault — coverage, reply handshake, chip select, startup race and bus data all measured clean, and every failure that looked like the board's turned out to be either the machine or a bug in the instrument. And the instrument was wrong often enough that the habit of testing it against a model, rather than against the machine it was pointed at, is what made its answers worth anything.
From the MS 0511 schematic, the RAM is 24 × K565RU5 (4164, 64K × 1) in two groups, and the arithmetic pins each plane to a group:
| chips | data lines | width | = |
|---|---|---|---|
| 8, on D10 | DG0–DG7 | 8 bits × 64 K | plane 0 — the PP's own RAM |
| 16, on D8 | DC0–DC15 | 16 bits × 64 K | planes 1 and 2 — the CPU's RAM |
24 chips × 64 Kbit is 192 KB, which is three 64 KB planes, which is what the emulator allocates. The 8-bit group can only be plane 0, because plane 0 is the one the PP reaches a byte at a time through 177012. The 16-bit group is the pair the CPU sees as words, reached together through 177014 — and since that port puts the low byte in plane 1 and the high byte in plane 2:
DC0 … DC7 = plane 1 = the low byte of every word the CPU executes
DC8 … DC15 = plane 2 = the high byte
So plane 1 bit 7 is the chip on DC7. Physically the 16-bit group is likely two rows of eight, and DC0–DC7 is the row that is plane 1.
One bound worth knowing before trusting a clean result later: the test walks 32768 addresses of each plane, which is half of a 64 KB plane. That covers everything the central processor can reach — its window below 160000 maps to about 28 KB of each plane — but not the upper half, which is video-only. A fault there would be in the same chip regardless, so it does not change which part is implicated; it does mean a pass is a pass over the CPU-visible half.
On hardware the frame read long long short short long short long short — alive, PP RAM good, planes not ok, plane 1 bad, plane 2 fine, done.
The final frame reads long long short short long short long long, then seven short and a long: plane 1 bad, stuck rather than leaky, bit 7 and no other. A single bit of the central processor's RAM that cannot hold a value at all — wrong on an immediate reread, not merely wrong later.
One bit is one column of the array, so on 1-bit-wide DRAM this is one chip. The
central processor's RAM is plane 1 and plane 2, and plane 1 is the low byte of
every word it executes. That is the machine's - ОШИБКА ОЗУ ЦП confirmed from
the outside, by a test running on the other processor with the faulty one held
in reset — and it explains the whole cluster of symptoms at once. The CPU's
program is copied into planes 1 and 2 before it is released, so corrupt low
bytes give a CPU that reports itself broken, halts with *** СТОП ***, or runs
just well enough to paint a menu, depending on where the damage lands.
It also explains why a bank tested out of circuit came back clean: plane 0 is the PP's own RAM, which the monitor tests and passes on every boot, and which this test passes too.
On the second run, pulses 9–16 read short for bits 0–6 and long for bit 7: a single bit, and therefore a single column of the array — one chip, not the bank, not the supply, not a shared strobe.
Pulse 8 separates the two ways one bit can be wrong. The test's first pass over the planes now writes a word and reads it straight back before anything else touches the array, in both polarities, and records what was already wrong at that moment. A bit wrong there cannot hold the value at all; a bit right there and wrong in the later passes held it and lost it. Dead chip against leaky one, which is what decides whether to suspect the part or its refresh — and, given the machine has been reported to worsen as it warms, worth reading cold and warm.
Two details in that pass are load-bearing. Writing 177014 updates the register as well as the array, so reading it straight back returns what was just written and proves nothing; the address register has to be rewritten to re-latch the data registers from memory. And the pass needs the complement half, because writing the address as the data leaves the high byte counting only 0…127 over a 32768-word plane, so bit 7 of plane 2 would never once be set — the test would have been structurally blind to exactly the kind of fault it just found, had it been in the other plane.
A probe on the suspect line during the test reported "much less activity than the other bits", which looked like evidence and was not. The test writes the address as the data, so the low byte counts 0…255 over and over: bit 0 toggles on every write, bit 7 once per 128. A perfectly healthy bit 7 looks sluggish next to bit 0 for that reason alone.
So the test no longer parks silently. Once it has reported, it sits in a loop writing all-zeros then all-ones to one address, which toggles all sixteen plane data lines at the same rate. Every bit becomes comparable with every other, and in particular bit 7 of plane 1 with bit 7 of plane 2 — the same position in the same kind of chip, and the only genuinely fair comparison available. A weak or dead line stands out against its own twin instead of against a faster neighbour. The beacon read stays in the loop so the LED keeps reporting while the probe is on.
Pulses 9–16 are the follow-up. The test already knew which bits — that is the mask it leaves at 077662 — it simply had no way to say so on hardware. Each pulse is one bit position of the failing plane's byte, and on a bank built from 1-bit-wide DRAM each bit is one column: a list of chips rather than a diagnosis to interpret.
ukncbtl needs no patching to run our ROMs. Emulator_LoadUkncRom() looks for a
file called uknc_rom.bin in its working directory and loads exactly
32256 bytes from it, falling back to the built-in resource only when the
file is absent. 32256 is our image size exactly, so:
$ cp ramtest.bin /path/to/ukncbtl/uknc_rom.binand the emulator boots the test ROM in place of the system ROM — the same substitution the board performs in hardware, for free.
Reading the result there is the only gap, because beacons are addresses on the bus: precisely what the board can see and what an emulator cannot show without being modified. So the test also leaves its verdict in memory, written after every test has finished:
| address | |
|---|---|
077660 |
status: 1 PP RAM ok, 2 plane 1 bad, 4 plane 2 bad, 10 done (octal) |
077662 |
every plane bit that was ever wrong |
A healthy machine parks with 077660 = 11 and 077662 = 0. Open the
memory view at 077660 and the answer is there, with no virtual LED to build.
The status bits are powers of two, and are written that way after an earlier
version set "done" to a Python 10 — decimal ten, binary 1010 — which
overlapped the plane 1 bit, so a healthy machine and one with a dead plane
reported the identical status word. The test did not catch it because both sides
of the comparison shared the error, which is how a diagnostic ends up lying with
confidence.
The video path is now fully specified, and it is simpler than expected. The
display is built from a tag list in plane 0 starting at 0000270 — and plane
0 is the PP's own RAM, so the PP writes the whole thing with ordinary MOV
instructions, no ports involved. Each 2-word tag is [addressBits, next], where
addressBits is where that line's pixels live and the low bits of next say
whether the following tag is the 4-word form that sets the palette or the scale.
307 lines are walked; drawing starts at line 19.
So a PP-side diagnostic can put a picture on screen using nothing but normal
memory writes, with the central processor held in reset — which works
identically in ukncbtl and on real hardware, and needs no LED and no beacons.
That is the natural home for test names, ПРОХОД/ОШИБОК counters and the
monitor test patterns. It is a real piece of work rather than a quick addition,
and the memory verdict above answers the immediate question without it.
test/test_ramtest.py runs the assembled image against a model of the PP with
three planes behind the registers and the ability to break one bit of one
plane, and checks that a healthy machine reports pass while a broken plane
reports itself and not its neighbour. It earned its keep twice: the first
version accumulated the OR of the observed and expected values rather than
their difference, so it reported both planes bad whenever either was; and it
caught the assembler bug below.
_num used int(tok, 0) — Python's rules — so a bare 177716 was decimal,
assembled as 133064, and the instruction meant to hold the CPU in reset wrote
into the middle of the ROM instead. Nothing complained; the program was simply
wrong. make_testrom.py had never noticed because it interpolates Python ints,
which round-trip through decimal by luck.
Bare numbers are now octal, as on any PDP-11, with 0x/0o/0b prefixes and a
trailing dot for decimal (10. is ten, 10 is eight). Both generators emit
octal into their templates, and test/test_testrom.py still passes, which is
what makes the change safe to have made.
With coverage instrumented, the frame came back long on 1, 2, 6 and 7 — and all four windows complete. Every word of every window we serve was asked for and answered by us. Nothing was banked away, nothing went unanswered, no reply went untaken, and we were listening before the machine asked its first question.
The screen is now doing the talking. Repeated resets produce, variously, the
boot menu; a *** СТОП *** halt display with a PC and PSW; and the startup test
screen:
СТАРТОВЫЙ ТЕСТ
- ошибка ЦП
ЦП is the central processor, and there is no - ошибка ПЗУ line beside it.
On that boot the monitor checksummed all four of our windows and was satisfied,
then found the central processor faulty.
That matters more than it looks, because of how the UKNC is built: the central processor has no ROM at all. Everything below 160000 in its address space is RAM and everything above is I/O, so it runs entirely on code the PP writes into it through the plane ports — the copy loop at 173252, sourcing from our image at 160000–173212 while the CPU is held in DCLO reset. The PP then releases it and reads its verdict back over channel 0 at port 177060.
So the chain is: our ROM → PP → CPU RAM → CPU self-test → a byte at 177060 → the message on screen. Our end of that chain now measures clean at every point we can instrument, and the checksum verifies the bytes.
The one link still unmeasured is the wire itself, which is what pulse 5 is for: the response machine samples the AD lines at the instant the host releases the read strobe, with our drivers still on, and compares against what it was asked to drive. If that stays short while the CPU error persists, the board has been exonerated at every point it is possible to exonerate it, and the fault is in the machine.
With the plain watch build flashed, a restart followed by a reset produced the boot menu on the real machine for the first time. That is the single most important result in this file, and it changes the shape of the problem: nothing is missing and nothing is fundamentally wrong. The machine is marginal.
The way it happened is the lead. A reset is not a power cycle — the board was already up, clocked and serving when the PP restarted. On a cold start it is not, and it cannot be:
| who is doing what | |
|---|---|
| power applied | the PP leaves reset and fetches its power-up vector at 160000 |
| meanwhile | the RP2350 runs its bootrom, sets up XIP, starts our code, sets the clock, loads two PIO programs, launches core 1 |
Nothing in the firmware can make the bootrom faster, so if the machine asks before we are listening, its first reads go unanswered. A ROM that is missing for the first instructions of a startup sequence is exactly the kind of fault that is intermittent, that clears on a reset, and that leaves the machine in a state no amount of reading the ROM contents will explain.
Pulse 9 tests it directly. The PP's first read after reset is its power-up vector, so if the first cycle we ever capture is 160000 we were in time; if it is anything else we came up mid-stream. Long on reset and short on cold power-on confirms the race.
The prediction is worth making before the measurement, because it is cheap and sharp: power on, wait a second, press reset should boot far more reliably than a cold start. If it does, the ROM board is not at fault in any interesting sense — it is simply late, and the fix is to keep it powered or to hold the machine off until it is ready.
The reasoning below is left because it is still sound as far as it goes; it is simply not evidence about what the machine reads, which is what pulse 2 settles. Both arguments are about the bytes we hold, and both remain true of a board whose data never arrives intact.
The menu text lives 40 bytes from text the machine already displays. The ЗАГРУЗКА block is at 103116, and the УСТ settings text the user can reach sits immediately before it at 103040:
103040 ый|3 - выключен |1 - включен |2 - выключен|..ЗАГРУЗКА.......
103140 ...(0.3): 0.........(1,2): 1|1 - диск |2 - кассета ПЗУ |
103240 3 - сеть |4 - стык С2 |5 - магнитофон |6 - отладка
103340 |7 - тестирование|...
Same chip, same window, same 128 bytes. A board that serves one and not the other is not a failure mode that exists.
The monitor checksums its own ROMs and is satisfied. The routine at 160410
sums each window as a ones'-complement sum and compares against four values
mask-programmed into the top of the last chip; a mismatch prints - ОШИБКА ПЗУ.
tools/rom_checksum.py runs that same algorithm on the images we serve:
$ ./tools/rom_checksum.py uknc_rom.bin
160000..176774 3839 words computed 103607 stored 103607 ok 208
140000..157776 4096 words computed 162125 stored 162125 ok 207
120000..137776 4096 words computed 133314 stored 133314 ok 206
100000..117776 4096 words computed 063160 stored 063160 ok 205 (DS4, ...)
all four blocks pass: the monitor's own ROM test is happy with these images
Note the block boundaries: the machine's own test partitions the ROM exactly by chip window, so a failure names a chip. Flip one bit anywhere in the 205's window and only that line goes MISMATCH.
Which decision is still open, and inference has gone about as far as it usefully
can. The board, though, sees every instruction the PP fetches — so it can be
asked directly. -DMPI_WATCH=ON builds a firmware that scores a handful of
monitor addresses and blinks the result; see Which build to run.
The decisive one is 101000, the emt 44 whose inline argument points at the
ЗАГРУЗКА string. If it hits, the menu was drawn and something happened to it
afterwards, which makes this a display problem. If it does not, the machine
branched away earlier and the search moves upstream. Either answer removes half
the remaining possibilities, and it costs one reflash and a look at the LED.
Selecting 7 (тестирование) from the menu in the emulator gives a screen headed
Т Е С Т И Р О В А Н И Е with ПРОХОД: and ОШИБОК: counters, and the pass
counter increments — it is a continuously looping test with an error tally, not
a one-shot report. That is a good model for our own suite: loop, count passes,
count errors, and stay readable while running.
Because the board is the ROM, it owns the machine from reset — which makes it possible to replace the system monitor with a diagnostic that runs on the peripheral processor itself.
The hook is the power-up vector. The PP fetches its starting PC and PSW from a HALT-mode vector at 160000/160002, which is offset 0 of the code 0 image. The stock monitor points it at 160300; put your own value there and your code runs before anything else in the machine does, with no monitor to work around.
Reporting results is the harder half — at reset the PP has no screen it can reach unaided. So the program signals by reading from reserved addresses. Every address is on the AD lines at the strobe and the board latches every strobe, so the board sees each beacon go past. It is the ROM under the program and the instrument watching it at once, which is not a thing a mask ROM could ever be. Beacons live in PP RAM near the top, so a read there disturbs nothing; only the address matters, never the data.
$ python3 tools/make_testrom.py -o testrom.bin
wrote testrom.bin: 32256 bytes, 48 words of code
power-up vector at 160000 -> 160300, PSW 000340
beacons at 077700: 0=alive 1=RAM ok 2=RAM bad 3=done
$ python3 test/test_testrom.py
healthy RAM: beacon sequence: [0, 1, 3]
RAM with a bit that will not set: beacon sequence: [0, 2, 3]
all checks passedtools/pdp11asm.py is a small assembler covering what test code needs, and
test/test_testrom.py runs the assembled image in a model of the PP before it
goes near hardware. That simulator immediately earned its place: the first RAM
test wrote each word's own address into it, which is a good address-decode test
but a weak stuck-bit test — a location only proves bit N works if its address
happens to have bit N set. A bit stuck low at an address that never sets it went
straight through. Hence the second pass with the complement, so every bit takes
both values at every location.
Next, in rough order of usefulness: run the image in ukncbtl, which takes the same 32 KB file and costs nothing to be wrong in; teach the firmware to watch the beacon range and report on the status LED; then extend the suite outward into the I/O page and the channel to the central processor.
Beacons are enough for a pass/fail, but not for a diagnostic anyone would want to read, and the obvious objection to a PP-side test suite is that the video memory belongs to the central processor. It turns out not to matter: the PP has its own port into it, and can draw the screen with the CPU held in reset.
| port | what it is |
|---|---|
177010 |
plane address register — a byte address into the 64 KB frame |
177012 |
plane 0 data — a byte written straight into the CPU's RAM |
177014 |
plane 1 & 2 data — low byte to plane 1, high byte to plane 2 |
177016 |
sprite colour |
177020/177022 |
background colour, planes 0–2, bits 0–3 and 4–7 |
177024 |
pixel byte: writes all three planes through the colour registers |
177026 |
plane mask |
Write an address to 177010, then a byte to 177012, and a byte of plane 0 has changed. Three planes give eight colours; 177024 with the background registers loaded does all three in one write, which is how you fill an area quickly.
The stock monitor uses exactly this, and its startup sequence is the worked
example: the routine at 173252 sets 177010 to 70000 and streams 3839 words into
177014, then walks a table of addresses writing 600 to each — all with the CPU
held in DCLO reset from the mov #40, @#177716 two instructions earlier. The
screen is up before the central processor has executed anything.
So the diagnostic can look like the stock one — a heading, a list of tests with
results beside them, ПРОХОД and ОШИБОК counters — and the test patterns for
setting up a monitor (greyscale ramp, colour bars, a border-to-border grid,
convergence crosshatch) are simply fills through 177024. None of it needs a
working CPU, a working keyboard, or a working disk, which is the point: it runs
on a machine that is too broken to run anything else.
The one thing to establish on hardware is the frame layout — where in the 64 KB
the visible lines actually are, which on this machine is set by a line table
rather than being a flat bitmap. The monitor's own drawing code is the reference
for that, and tools/pdp11dis.py reads it.
Before going further, know that this problem has been solved. The RE-mulator is a DIP-24 board built around an STM32F205 at 120 MHz that drops straight into a 1801RE2 or 1801RR1 socket, takes its power from the socket, can be reflashed in place, and — with additional external inputs — stands in for up to four separate chips at once. It has been verified in a BK-0010 running at 3 MHz replacing a 1801RE2-017. Four chips at once is exactly the UKNC shape, and those "external inputs" are almost certainly the per-window enables discussed above.
Two honest conclusions:
- If your goal is simply a working UKNC with reprogrammable ROMs, build or buy a RE-mulator. It is purpose-made, proven on real hardware, and a fraction of the work of this.
- If your goal is to do it on One ROM hardware, its existence is still good news: it settles that the part is DIP-24, that socket power is usable, and that a 120 MHz CPU-in-the-loop design is fast enough — which is the core assumption this design rests on. Its manual also documents the pinout, which is the one thing blocking this repo.
This is the part worth internalising before worrying about speed. A 2364 gives you a hard deadline: address changes, and valid data must be on the pins inside the access time, with nothing to say otherwise. MPI is a fully asynchronous handshake — the host waits for nRPLY. Responding late costs wait states, not corruption, and the only real limit is the host's bus timeout, which is orders of magnitude longer than anything the RP2350 will take.
That is what makes a CPU-in-the-loop design safe here, and it is why the address decode below is allowed to be a dozen instructions of table lookup rather than something exotic. The PIO handles the edges; the CPU has from nSYNC to nDIN to think, and even overrunning that window is survivable.
From One ROM's own board description (rust/config/json/fire-24-e.json), the
Fire 24 rev E maps its 22 signal socket pins onto GPIO 0–7 and GPIO 10–23.
GPIO 8 and 9 go to the X1/X2 jumper pads, not to the socket.
That gap is why this cannot be a straightforward port of One ROM's approach.
One ROM leans on being able to read a contiguous run of address pins in a single
PIO in pins, n and use the raw value — however scrambled the bit order — as a
table index, with the pre-processor pre-scrambling the image to match. Order
does not matter; contiguity does. On Fire 24 rev E the longest run of
socket-connected GPIOs is 14, so sixteen AD lines can never be one window, no
matter how the chip's pins fall.
The way out is to stop trying. Read GPIO 0–23 as a single 24-bit field — including the two jumper bits, which are static and harmless — and let the CPU un-scramble it with three 256-entry lookup tables. Sixteen output bits are handled the same way, in reverse, but entirely at boot: the drive pattern for every one of the 4096 words is precomputed, so the serving path never scatters a bit at runtime.
With the real pinout in, the sixteen AD lines land on GPIO
0 1 2 3 4 5 6 7 10 11 13 19 20 21 22 23
with nSEL on 15, nDIN on 16, nRPLY on 17 and nSYNC on 18. Nothing contiguous anywhere, which settles the question — but every one of them is inside the 24-bit field, so one read and one write still cover the whole bus. Socket pins 21 and 22 are not connected; they land on GPIO 12 and 14, feed no address bit, and are pulled down so they do not float.
GPIO 8 and 9 are deliberately never muxed to the PIO and never appear in a direction mask, so a fitted X jumper cannot be shorted by an output driver.
Two PIO state machines and one CPU core.
mpi_capture waits for nSYNC high (so a machine that starts mid-cycle
resynchronises rather than latching garbage), then for the falling edge, then
snapshots all 24 GPIOs into the RX FIFO. The MPI address hold time after nSYNC
is far longer than the ~20 ns of synchroniser plus one instruction, so sampling
straight off the edge lands well inside the valid window.
Core 1 pops the snapshot, gathers the logical address through
g_gather[3][256] (pin inversion baked in), picks the window table by the top
three bits, and pushes the precomputed 24-bit drive pattern. If the address is
not ours it pushes nothing, and the response machine simply stays blocked — a
non-matching cycle produces no bus activity at all, which is the behaviour you
want when you are one of several devices on a shared bus.
mpi_respond waits for nDIN, presents the data, enables the AD drivers, and
one instruction later enables the nRPLY driver — giving the host data setup time
before the handshake completes. nRPLY is open-drain by construction: its output
value is always 0, so enabling the direction pulls it low and clearing the
direction returns it to the bus pull-up. The line is never driven high.
Roughly six PIO cycles from the nDIN edge to data and nRPLY, about 40 ns at 150 MHz — faster than the chip being emulated.
One correctness wrinkle drives a design choice worth spelling out. We latch on the address strobe, which is asserted for every cycle on the bus, but the read strobe only arrives for a read the host has decided belongs to us — on the UKNC, EDIN, which the CGM withholds when the window is banked to RAM. So a prepared cycle routinely ends without ever being served: writes, cycles for other devices, banked-out windows.
Once the response machine has executed its pull, the pattern is in the OSR and
the TX FIFO reads empty, so neither a FIFO check nor watching the address strobe
can tell it is sitting on a stale value. It would then apply that value to
whatever read came next — wrong data, driven confidently. So rearm_respond()
puts the machine back to idle unconditionally at the start of every cycle: a few
register writes inside the strobe-to-strobe gap, in exchange for a machine that
cannot carry state across cycles.
| File | |
|---|---|
firmware/mpi_rom.pio |
the two state machines |
firmware/main.c |
hardware setup and the core 1 serving loop |
firmware/decode.c |
the per-cycle arithmetic, free of SDK dependencies |
test/test_decode.c |
host test of that arithmetic |
firmware/board_fire24e.h |
socket-to-GPIO map and chip pinout |
firmware/rom_images.h |
image table interface |
tools/re2_convert.py |
dump format conversion, tested |
tools/gen_rom_images.py |
emits firmware/rom_images.c from dumps |
tools/diagnose_dump.py |
tells a bad dump from a bad chip |
tools/selftest.py |
the self-test pattern, shared by generator and checker |
tools/check_selftest.py |
diagnoses a dump of the self-test build |
firmware/watch.h |
the watchpoint table for the MPI_WATCH build |
test/test_watchpoints.py |
checks those watchpoints against the ROM |
tools/pdp11asm.py |
small PDP-11 assembler, for test ROMs |
tools/pdp11dis.py |
small PDP-11 disassembler, for reading the stock ROM |
tools/make_testrom.py |
builds a test ROM that replaces the system monitor |
test/test_testrom.py |
runs that ROM in a model of the PP |
tools/rom_checksum.py |
the machine's own ROM test, run offline |
A reader that ignores RPLY cannot tell "the chip did not answer" from "the chip answered with these bits". When the chip stays silent the AD lines float to whatever the rig's pull resistors give, so a bit position reads as one constant value across the whole dump — indistinguishable, by eye, from a stuck bit.
The chip only answers when nAD13–nAD15 match its mask-programmed code, so addressing the wrong 8 KB window silences it entirely. That makes "wrong window" and "dead chip" look alike, and the four UKNC chips have four different codes.
diagnose_dump.py separates them by shape. Whole bus constant means no reply;
one or two bit positions constant means a stuck bit or an open line; nothing
constant but still disagreeing with a reference means addressing or timing.
Pass --reference with a known-good image — the k1801 archive has all four
UKNC chips — and it will say which pattern the differences fit.
The analysis is invariant under the programmer-order transform, so it does not matter which orientation either file is in.
The 1801RE2 dumps in the 1801BM1/k1801
archive are in "Sterkh programmer" order: every byte inverted, and the word
addresses inverted too. Reduced to bytes, and matching the original rev16
utility exactly:
out[addr ^ 0x1FFE] = ~in[addr]
0x1FFE inverts byte-address bits 1–12 — the twelve word-address bits — leaving
bit 0, which selects the byte within the word. The transform is its own inverse.
Files carry two trailing bytes: the chip code, then a constant 0x03.
This part is done and checked. Round-tripping is an identity across all 50 dumps
in the archive, the codes recovered from the trailers match the archive's own
table, and converted images are unmistakably PDP-11 code where the raw ones are
not — the BK-0010 monitor goes from 0 to 80 occurrences of RTS PC and 0 to 159
of JSR PC.
The four UKNC images are 205_mc0511.rom through 208_mc0511.rom in that
archive, and they generate as a complete set:
$ ./tools/re2_convert.py 205_mc0511.rom
205_mc0511.rom: code=3
as read RTS PC= 0 JSR PC= 0
converted RTS PC= 110 JSR PC= 76
$ ./tools/gen_rom_images.py -o firmware/rom_images.c \
205_mc0511.rom 206_mc0511.rom 207_mc0511.rom 208_mc0511.rom
wrote firmware/rom_images.c: 4 image(s), windows ['0', '1', '2', '3']which yields, with 208 correctly stopping short of the I/O page:
const mpi_image_t mpi_images[] = {
{ "205_mc0511", 3, 4096, img_re2_205_mc0511 },
{ "206_mc0511", 2, 4096, img_re2_206_mc0511 },
{ "207_mc0511", 1, 4096, img_re2_207_mc0511 },
{ "208_mc0511", 0, 3840, img_re2_208_mc0511 },
};Power is settled: the 1801RE2 follows JEDEC, so pin 24 is Ucc and pin 12 GND, exactly what the Fire 24 hard-wires to its regulator and ground plane. The board drops in unmodified. CS polarity is settled too: the schematic straps it to ground on three of the four sockets, so grounded means selected, so active low. The whole pin configuration is now read off documentation rather than guessed. What is left is behavioural, not electrical.
Write behaviour. Worth answering from the datasheet rather than guessing:
does a real 1801RE2 assert nRPLY on a write into its window? If it does not,
writes to ROM produce a bus timeout trap, and software may depend on that. The
BK-ROM-Disk GAL (RPLY = CHIPSEL & (DIN # !_DOUT)) does reply to writes, but
that is a RAM-disk controller, not a ROM. If the answer is "no", nothing needs
adding: this firmware only ever responds to nDIN.
After the RAM repair the plain build would not initialise the machine and the status LED blinked fast — which is precisely what the LED is defined to mean: cycles being served, replies prepared, none of them taken.
The response program samples the bus on every served cycle and autopushes the
result. Something has to empty that FIFO whether or not anyone reads it, and the
drain was written inside #if MPI_WATCH. The watch builds emptied it; the plain
build never did. Four served cycles filled it, autopush stalled the state
machine at the in, and the board stopped answering — while core 1 carried on
capturing addresses and queueing replies nobody would ever take, which is the
5 Hz blink exactly.
The comment above the drain said "it also has to happen every iteration regardless: autopush stalls the state machine on a full FIFO", and it was behind a conditional. Knowing the rule and encoding it are different things.
It drains unconditionally now, with only the comparison under MPI_WATCH, so
the plain and diagnostic builds share the one path and cannot diverge again.
Both 150 MHz and 200 MHz boot an MS 0511, so the deadline is not tight and the default is 150 MHz — the RP2350's rated speed. Overclocking to buy margin against a deadline you are comfortably inside is a real cost for an imagined benefit: more power, more heat, more radiated noise, inside a machine whose other problems are usually analogue.
MPI_SYS_CLK_KHZ in main.c raises it if that ever changes. The status LED is
what should decide it rather than caution — it blinks when replies are prepared
and not taken. Solid in normal use means the clock is not the constraint.
Keep SOCKET_CS_CODE at 03 for the DS4 socket. Ignoring CS made no difference
to booting, but that only shows CE0 is irrelevant while the window holds the
on-board ROM. It exists to say when the window belongs to RAM or a cartridge
instead, and honouring it is what keeps the board off the bus then.
$ cmake -S firmware -B firmware/build-watch -G Ninja -DMPI_WATCH=ON
$ ninja -C firmware/build-watchThis serves the bus exactly as the normal build does — the scoring happens after the reply has been queued, so a watch build cannot introduce the timing fault it might be used to rule out — and repurposes the status LED to report which of a handful of monitor addresses the machine has executed since power-on.
Each pass is one frame: 1.5 s dark, then one pulse per watchpoint in table order, long (1 s) for hit, short (0.1 s) for miss, 0.4 s apart. Every watchpoint gets a pulse whether or not it hit, so a position can never be miscounted — the failure mode of any scheme that blinks only the hits. Nothing is ever cleared, so a frame that changes between passes is itself a fact.
| pulse | a long pulse means | |
|---|---|---|
| 1 | 160302 after 160300 | the monitor started from our vector — if this is short, stop, nothing else means anything |
| 2 | 160450 after 160446 | a ROM block failed its checksum |
| 3 | 160342 after 160340 | the PP called the routine that loads the central processor's memory |
| 4 | 160374 after 160372 | the PP released the central processor — the ACLO edge that starts it |
| 5 | 101006 after 101004 | the boot menu header was printed |
| 6 | 174170 after 174172 | the PP is scanning its idle task queue — alive and dispatching |
| 7 | (bus) | a reply was prepared and the host never took it |
| 8 | (bus) | the bus carried something other than what we drove |
| 9 | (bus) | the first cycle of this boot was the power-up vector fetch — we won the startup race |
| 10 | (bus) | the bus went quiet — a whole second with no cycle served |
| 11 | (bus) | all four windows fully covered: every word we serve was asked for |
| 12–13 | (bus) | a 2-bit number, most significant first, naming the block whose checksum failed — meaningless unless pulse 2 is lit |
Pulses 6 and 10 are a pair, and they exist for the frozen screen: a boot menu drawn with no blinking cursor and no response to the keyboard. The cursor blink and the keyboard scan are both tasks on the PP's dispatcher, so a static menu means the PP is not dispatching — and these say which kind:
| 6 | 10 | |
|---|---|---|
| long | short | the PP is alive and scanning, but no task ever becomes ready — interrupts or the timer |
| short | long | the PP has stopped fetching altogether — a halt, a trap, or a reply that never came |
| short | short | it left the dispatcher and is running somewhere else — a wild jump |
00 = 205, 100000-117777 10 = 207, 140000-157777
01 = 206, 120000-137777 11 = 208, 160000-176777
The loop body is the same code for all four blocks, so no fetch address
distinguishes them. The data does: cmp 176766(r5), r3 at 160442 reads its
stored sum from 176776, 176774, 176772 and 176770 as r5 walks 8, 6, 4, 2, and
that read lands directly behind the fetch of the instruction's second word at
160444 — a pair, so the summing loop cannot forge it while reading those same
words as data on its way down.
A frame describes one boot. Hits were originally never cleared, which made a frame the union of every boot since the board was powered — and that ambiguity bit immediately, when a boot printed a CPU error with no ROM error while the checksum pulse was still lit from an earlier attempt. The frame now clears when the PP takes PC and PSW from its power-up vector, detected as the pair 160002 directly behind 160000 so the checksum cannot forge it by reading those same two words on its way down.
That clearing was got wrong once, in a way worth recording. It was first done on core 0, at the top of the next frame — which is up to ten seconds after the restart, by which time the machine has finished booting. The wipe therefore landed squarely on the startup evidence it existed to isolate, and the frame came back with the sanity pulse short while the "we saw this boot begin" pulse was long: a combination that cannot happen on a machine that is booting at all, since it says the monitor never ran while we watched it start. The reset now happens on core 1, two cycles into the new boot, and coverage is a generation tag rather than a bitmap precisely so that clearing it costs one increment instead of a 4 KB memset between two bus cycles.
The frame has been trimmed as questions closed. The PP RAM test, the error
printer and the end of the startup test answered consistently — RAM clean, no
error line, startup completes — and every pulse spent re-confirming a settled
fact is one the reader has to count past. The CS pulse went with them: it never
lit, so no read was ever declined, and -DMPI_IGNORE_CS=ON is moot.
Pulses 6–9 exist because the instrument had a blind spot the size of the remaining question. With "we declined a read" and "a reply went untaken" both negative, we answer every read we are asked and the host takes every answer, and the machine still computes a bad checksum. Either the data is corrupted electrically between our pins and the processor, or some reads never reached us — and the second is invisible to any counter of cycles we saw, because the read strobe is EDIN and the CGM withholds it for a window banked elsewhere.
The startup checksum reads every word of every window exactly once, so one bit per word settles it: all four windows covered means every word came from us and the corruption is electrical; a window short of its count means its reads went somewhere else, and that window is the failing block.
The first version of this scored a single address, and on hardware it reported every watchpoint as hit. That was not a finding, it was a bug, and the reason is worth keeping: the bus does not distinguish an instruction fetch from a data read, and the startup test checksums all four ROMs — it reads 16127 of the 16128 words, every watchpoint among them, within moments of power-on. "This address was read" is true of the entire ROM.
What separates a fetch from a data read is the company it keeps. Straight-line execution reads consecutive words back to back; the checksum walks downwards and puts three fetches of its own loop body between every pair of data reads, so a data read of A is followed by 160434, never by A+2. The plane-copy loop at 173270 ascends but does the same. So a watchpoint is two addresses — the one to score and the one that must have arrived immediately before it — and picking the second word of a multi-word instruction makes the predecessor its own first word, which holds whether the instruction was reached by fall-through or branch.
test/test_watchpoints.py checks both halves against the ROM image: that each
pair is really adjacent in the code, that the disassembly at the predecessor is
what we think it is, and that neither ROM-scanning loop can forge the adjacency.
It also reports what the old rule would have done, which is how the bug is kept
fixed:
$ python3 test/test_watchpoints.py uknc_rom.bin
160300 -> 160302 mov @#172660, r4 2 words ok
160446 -> 160450 beq 160452 1 word ok
(single word, falls through with no bus cycle)
...
no ROM-reading loop forges these adjacencies:
checksum (descending) clear
plane copy (ascending) clear
under the old address-only rule the checksum alone would light 6 of 6 watchpointsThe pairing can only produce false negatives: another master interleaving a cycle, or a write landing between the two fetches, breaks the pair and loses a hit. A short pulse means "not seen", not "did not happen".
Edit firmware/watch.h to watch something else; the table is the interface.
Needs arm-none-eabi-gcc, CMake, Ninja and the Pico SDK. The board carries an
RP2354A — an RP2350A with 2 MB of stacked flash — so it builds as a plain
rp2350 target.
$ ./tools/gen_rom_images.py -o firmware/rom_images.c \
205_mc0511.rom 206_mc0511.rom 207_mc0511.rom 208_mc0511.rom
$ cd firmware
$ PICO_SDK_PATH=/path/to/pico-sdk cmake -B build -G Ninja -DCMAKE_BUILD_TYPE=Release .
$ cmake --build buildFlash build/mpi_rom.bin with One ROM Web or the CLI's
--firmware option — the same raw-binary route One ROM uses for its own
onerom-rp235x.bin, loaded at 0x10000000. A .uf2 is produced too, but the
board has no BOOTSEL button, so the USB route is the practical one.
One ROM's USB is a TinyUSB device stack presenting its own vendor interface on VID 0x1209 / PID 0xF542 — "picobootx", an extended picoboot. The web flasher therefore talks to One ROM's running firmware, not to a bootloader. Flashing this firmware replaces that stack, so the board stops appearing in One ROM Web entirely.
One ROM Web goes further and validates what you hand it — a binary that is not recognisable One ROM firmware is refused outright, so it will not flash this project at all. Two routes do work:
- pico⚡flash, by the same author. It speaks picoboot rather than checking for One ROM metadata, and reaches both One ROM's running firmware and the bare bootrom.
- The bootrom's mass-storage volume, which takes a
.uf2by drag and drop.
For the second, header J2 on the Fire 24 rev E carries everything needed:
| J2 pin | J2 pin | ||
|---|---|---|---|
| 1 | SEL_A | 2 | GND |
| 3 | SEL_B | 4 | GND |
| 5 | BOOT | 6 | SWCLK |
| 7 | RUN | 8 | SWDIO |
Short pin 5 to pin 4 and power up, and the volume mounts. BOOT reaches QSPI_SS through a 1K against its 10K pull-up, so grounding it wins. If the board is already powered, briefly ground pin 7 (RUN) instead of power-cycling.
J2 is a footprint, though, not necessarily a fitted header — on a board that shipped without it there is nothing to short. In that case SWD is the route, and it is the better one anyway since it works regardless of what is on the flash:
$ openocd -f interface/cmsis-dap.cfg -f target/rp2350.cfg -c "adapter speed 5000" -c "program firmware/build/mpi_rom.elf verify reset exit"SWCLK and SWDIO are J2 pins 6 and 8, ground on 2 or 4. Those pins double as image-select jumpers C and D, so leave those jumpers off while programming.
That route needs no working firmware at all, which makes it the real safety net: a corrupt image sends the bootrom to USB by itself.
On top of that, image-select jumper 0 doubles as a recovery jumper, so you do
not have to go looking for the BOOT pad. Fit it and power on:
before a single socket pin is touched, the board hands straight back to the
bootrom's USB mode. It is checked first thing in main() so that it still works
when the rest of this firmware does not — though note it does depend on this
firmware booting at all, which is why BOOT-to-GND remains the fallback beneath
it.
The check does not assume which rail the jumper ties to. A floating pin follows whichever internal pull is applied and a driven one does not, so comparing a read under pull-down with a read under pull-up detects a fitted jumper either way — worth the few microseconds when the cost of getting it backwards is a board that never runs, or one that cannot be recovered.
rom_images.c is generated and gitignored — it holds actual ROM contents, which
have no business in the repository.
decode.c holds the whole per-cycle arithmetic and deliberately has no SDK
dependency, so the interesting half of the firmware can be tested with nothing
but a host compiler:
$ make -C test check
pin map
window index
round trip over all four windows
16128 addresses served correctly
silence outside the served windows
16640 addresses correctly left alone
AD0 does not change the word selected
all checks passedThe round-trip test presents every address in every window the way the host
would — inverted and scattered across the real pin map from board_fire24e.h —
runs it through the same gather tables the firmware builds, and reads the answer
back off the resulting drive pattern with an independently written model of the
bus. It also checks the board stays off the bus everywhere it should: the four
windows it does not serve, and the I/O page the code 0 chip must stop short of.
The tests have been mutation-checked. Dropping the complement in
mpi_window_index(), forgetting that data is inverted on the wire, and serving
the full window over the I/O page are each caught.
What this does not test is the PIO. That needs hardware — see below.
Read back through an Arduino-based 1801RE2 reader, all eight windows answered with the right data in the right places: address capture, window decode, data drive and bus turnaround all work. The code 0 window returned exactly 3840 words, so the I/O page truncation holds too.
About 3% of reads came back displaced by one bit — the correct word, correctly
selected, misaligned. A parallel bus cannot displace bits, so that belongs to
whatever samples the lines, not to the board; check_selftest.py now names it
rather than blaming an address line.
The four converted images, concatenated in window order (205, 206, 207, 208), are byte-identical to the reference ROM that ukncbtl ships — all 32256 bytes:
ours = b"".join(convert(split_dump(Path(f"{n}_mc0511.rom").read_bytes())[0])
for n in (205, 206, 207, 208))
assert ours[:32256] == Path("uknc_rom.bin").read_bytes()Worth running before suspecting the board of anything. It checks the image choice, the programmer-order conversion, the window ordering and the I/O page boundary in one line — the reference is 32256 bytes rather than 32768 precisely because 177000-177777 is the I/O page.
An MS 0511 with all four mask ROMs removed and one board in the DS4 socket initialises its video RAM and reaches a cursor. That settles the two things a bench reader cannot exercise: nRPLY, which a reader latching on a fixed delay never looks at, and latency against a live bus, where being slow costs wait states right up until the read strobe has come and gone before the reply is ready, and then the cycle gets nothing at all.
What is still unexercised is CE0. It only speaks for the 100000 window and only matters when software banks that window to RAM or to a cartridge, so ordinary booting never touches it. The board treats an open CS as selected, so a surprise there fails toward answering rather than falling silent.
The status LED reports the remaining unknown directly: lit means answering, dark means nothing is asking, and blinking means replies are being prepared and not taken — some of which is normal, but a reply assembled too late to be sampled lands in the same count.
If you have a reader that can dump a real 1801RE2, you already have an MPI bus master, and it is a far better first test than the machine: it exercises address capture, window decode, the data drive and the bus turnaround, with nothing expensive attached and no video timing to corrupt.
Flash the self-test build rather than a real ROM image:
$ ./tools/gen_rom_images.py --selftest -o firmware/rom_images.c
$ cmake --build firmware/buildThat fills all eight windows with a pattern whose every word encodes both its own index and the window it belongs to. Dump it back through the reader, then:
$ ./tools/check_selftest.py dumped.bin --window 3The point is that the three failures which look identical in a dump of real code are separable here:
| symptom | reported as |
|---|---|
| board never drove the bus | every bit stuck at one level |
| wrong window answered | a clean pattern, but from another window |
| a data line not driving | that bit position never changes |
| an address line swapped or stuck | valid words at permuted indices, and the XOR names the bits |
A stuck data line below bit 12 also perturbs the index a word claims, so the checker masks those bits out before blaming the address lines — otherwise one bad data line reports as an address fault too.
Two things to get right on the bench. Power the board and the reader from the same supply so they rise together: that is the case One ROM validated for the brief window where 5 V is present before the regulator has brought VDD up. And make sure the reader either grounds pin 23 or leaves it open — the firmware pulls CS down internally so an open pin reads as selected, matching the three UKNC sockets that strap it to ground.
If the reader does not wait for RPLY but latches after a fixed delay, a passing dump proves the data path but says nothing about the reply. Worth knowing which you have before reading too much into a green result.
Do not go straight into a host.
- Build, and check the pin map against your own reading of the datasheet.
board_fire24e.hcarries the pinout in a comment for exactly that. - Drive the board from a second RP2350 running a synthetic MPI master, or from a logic analyser plus a hand-clocked cycle. Check the address snapshot decodes to what you presented, and that the AD lines and nRPLY stay hi-Z for an address outside the configured window — that is the test that protects the host's bus drivers.
- Scope the real machine before committing — on the peripheral processor's bus. Capture nSYNC, nDIN, nRPLY and a couple of AD lines during a ROM read with the original chip in place. That gives you the actual SYNC-to-DIN gap, which tells you how much slack core 1 has, and shows whether the AD lines are driven push-pull or open-drain. This firmware drives them push-pull during its response window; if the host's pull-ups are weak and something else is contending, that assumption needs revisiting.
- On the same capture, watch EDIN across a port 177054 write that banks RAM into that window, and confirm the strobe stops arriving. That is the gate the whole design leans on, so it is worth seeing rather than assuming.
- Only then, one image, one window, in the host. Start with 206 or 207 — the plain system ROM windows, no banking games, no I/O page adjacency, and CS grounded so there is one less variable.
On levels: the RP2350 GPIOs are directly connected to the socket, with no
buffers, and are 5 V tolerant to 5.5 V once VDD is up. One ROM uses 8 mA drive
strength for 5 V hosts and has validated the 20 µs power-on window where VDD is
still rising; the same analysis applies here and is worth reading in the
project's docs/VOLTAGE-LEVELS.md.
- One ROM — board description
(
rust/config/json/fire-24-e.json), serving algorithms (docs/firmware-rewrite.md), plugin API (firmware/ora/api.h), levels (docs/VOLTAGE-LEVELS.md) - KR1801RE2 datasheet, table 11.26 and figure 11.30 — the pinout
- Elektronika MS 0511 schematic, revision 5 (sheet 1) — DS1–DS4 wiring, the CGM at D10, and the EDIN / CE0–CE3 gating
- 1801BM1/k1801 — 1801RE2 description, code
table, dump archive, and the
rev16conversion utility - vldmrrr/BK-ROM-Disk — GAL equations for a device on the same bus; useful confirmation of how nSYNC latching and nRPLY generation are done in practice
- nzeemin/ukncbtl — the UKNC PP memory map
and the port 177054 window gating, in
emulator/emubase/Memory.cpp,CSecondMemoryController::UpdateMemoryMap(); the same file'sGetPortWord/SetPortWordare the reference for the plane ports 177010–177026 and for 177716 driving the CPU's HALT, DCLO and ACLO pins, andemulator/res/uknc_rom.binis the reference image - RE-mulator — the existing DIP-24 1801RE2/RR1 in-circuit emulator
- Elektronika MS 0511 — processors, clocks, and the 1801RE2-205..208 ROM set
- One ROM Fire 24, onerom.org