Repository navigation
Conversation
Changes over the sepsf-2 copy (e14d8ea). Every EL restores benchmarkoor's jochemnet/24402727 datadir instead of a canonical snapshot. jochemnet started from the mainnet 24350000 snapshot and ran 1 s blocks of bloat state up to 24402727 (head ts 2026-01-31). Its head is not on mainnet. Snapshots (snapshots.ethpandaops.io/jochemnet/<client>/24402727/), layouts checked from the first 64 MB of each tarball: - geth/besu/nethermind snapshot.tar.zst, reth snapshot-v2.tar.zst (v2 storage), erigon snapshot-pruned.tar.zst. Sizes 0.7-1.26 TB compressed, so the 1.8 TB droplets (DO's largest local disk) stay. - ethrex snapshot.tar.zst (0.3 TB) is new compared to sepsf. It extracts into /data/ethrex/chain-1, which is where ethrex 29 keeps a chain-id-1 custom-genesis db. It runs --syncmode=full; the May db is migrated on start. - nimbus-el still has no snapshot and snap-syncs from our peers. - shadowfork.yaml: snapshot_fetcher_filename moves from play vars (which beat group_vars) to a template default, so it can be set per EL. Genesis (egg 6.2.0, CHAIN_ID=1, mainnet deposit contract): - jochemnet publishes no _snapshot_eth_getBlockByNumber.json. The one inside the erigon tarball is the stale mainnet 24350000 block. SHADOW_FORK_FILE is therefore files/shadowfork_block.json, the full-tx eth_getBlockByNumber of a restored EL; TODO before playbook.yaml. Tested with a stand-in head: the real genesis role keeps osaka/bpo1/bpo2 at mainnet times and puts amsterdamTime at genesis + 40 epochs. genesis.ssz builds; a hash-only block makes eth2-testnet-genesis fail. - Mainnet has no EIP-8282 builder contracts, so the devnet-7/8 deploy-eip8282-contracts assertoor startup test is back. It sends the same 0x4e59 factory calldata as the sepolia deployment txs; checked with cast create2. Clients: - nethermind: --config=mainnet, nethermind_db/mainnet, FlatDb (the tarball's mainnet/flat layout). - erigon: keeps --externalcl and --keep.stored.chain.config; db.size.limit is 2TB. The sepsf-2 --snap.stop host_vars are dropped: segments are frozen up to 24401000, so ~1.7k blocks are left to retire (sepsf-2 had 220k). - Static EL enodes reset (they named sepsf-2's hosts), plus a note to firewall EL p2p fleet-wide: mainnet peers have none of our blocks. Other: - Images: nimbus v26.10.0, prysm v7.2.1, ethrex 29.0.1 (released). - terraform: VPC 10.32.0.0/16. - Secrets: the metrics/logs usernames and the xatu ingest address are now glamsterdam-msf-1. Launch TODOs (marked TODO(msf-1)): genesis time, gloas_fork_epoch, files/shadowfork_block.json, shadowfork_el_static_enodes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
This comment has been minimized.
This comment has been minimized.
…ates) From a five-way review against the client release tags, plus local runs of every EL image (2026-10-07). Genesis: - egg's mainnet genesis.json does not load as is. nethermind 2.1.0 fails on alloc keys without 0x, and besu 26.9.0 needs the EIP-7002/7251 request contract addresses. scripts/normalize-shadowfork-genesis.py fixes both; a play right after the genesis role runs it, for chain id 1 only. All seven ELs load the result with genesis hash 0xd4e567..cb8fa3 and our amsterdamTime. Keep real mainnet out (chain id 1 has mainnet's fork id until amsterdamTime): - erigon: --nodiscover (the only way to stop its mainnet DNS tree) and --no-downloader (jochemnet segments carry mainnet names). Tested on v3.7.1. - reth: --disable-dns-discovery. - nethermind: --Network.EnableEnrDiscovery=false and --Discovery.UseDefaultDiscv5Bootnodes=false. - bootnode geth: explicit --bootnodes, so it no longer falls back to mainnet's. - The restore deletes each tarball's p2p identity and saved peers (nodekey, nodes/, discovery-secret, known-peers.json, key, node.key, node_config.json, <net>/peers, <net>/discoveryNodes), so hosts restored from the same tarball don't share a node id. - The DO firewall admits 30303 only from our droplets (el_p2p_extra_source_addresses for a local join host). - The iptables script also matches --sport 30303 and reports failed or unreachable hosts. It runs right after init-server. Wallets / tooling: - fund-shadowfork.sh skips every EIP-7702-delegated wallet. mainnet's mnemonic-4 (faucet_private_key) is delegated to a sweeper. - k8s: faucet disabled (that key) and homepage disabled (its MetaMask button offers chain id 1); faucet-agents stays. - nimbus-el: --debug-snap-sync-resume=true; start only after Gloas finalizes, because v0.4.2 needs a post-Amsterdam pivot. Gate: - scripts/shadowfork-pregenesis-check.sh. genesis.json checks: 0x keys, request addresses, fork and blob keys, CL head hash. Per EL host: head and hash, eth_config next activation, unique node id, fleet-only peers, disk, erigon preverified.toml, ethrex metadata.json. Tested against a mock RPC with injected failures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
76 droplets in total. NUMBER_OF_VALIDATORS stays 36000: terraform splits each entry's 1000-key range over its 2 nodes (checked with terraform console: contiguous 0..36000, all supernodes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
Replaces the 2-per-pair matrix (76 droplets). 3000 keys per node, 36000 total; every CL and EL client holds 1/6 of stake and no CL/EL pair repeats. Names keep the hosts group_vars reference as EL bootnodes. Plus bootnode, 2 buildoor and 1 nimbus-el. Checked with terraform console. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
There was a problem hiding this comment.
Summary
msf-1 shadows mainnet from the jochemnet benchmark snapshots: per-EL snapshot filenames, chain-id-1 shadowfork genesis with a committed block file, mainnet-specific EL peer isolation, a pre-genesis gate, and a regenerated 41-node inventory. The last commit aligns inventory.ini with the restored 36-validator nodes.tf layout and I verified all node names, validator ranges and terraform-computed regions match. The one open item is the still-absent shadowfork_block.json the genesis role requires.
Issues
- 🟡
ansible/inventories/msf-1/group_vars/all/all.yaml:107— shadowfork_block.json lookup targets a file not in the repo — see the thread on that line
Reviewed @ 6cbfefd5
"Weeks of coding can save you hours of planning."
| export DEPOSIT_CONTRACT_ADDRESS=0x00000000219ab540356cBB839Cbe05303d7705Fa | ||
| export SHADOW_FORK_FILE=/config/shadowfork_block.json | ||
| # TODO(msf-1): commit files/shadowfork_block.json. | ||
| shadowfork_block.json: "{{ lookup('ansible.builtin.file', (ansible_inventory_sources[0] | dirname) ~ '/files/shadowfork_block.json') }}" |
There was a problem hiding this comment.
🟡 shadowfork_block.json lookup targets a file not in the repo
This resolves to ansible/inventories/msf-1/files/shadowfork_block.json, but no files/ directory is tracked. The ethereum_genesis role consumes the dict via ethereum_genesis_generator_config_files | dict2items (roles/ethereum_genesis/tasks/generate_genesis.yaml), so the missing file lookup raises and aborts the genesis play before any host runs. It is marked TODO, but until the file is committed playbook.yaml cannot get past genesis.
| fmt="${RPC_URL_FMT:-}" | ||
| [ -n "$fmt" ] || fmt="https://$user:$pass@rpc-{host}.srv.glamsterdam-$net.ethpandaops.io" | ||
| fleet="$(grep -oE 'ansible_host=[0-9.]+' "$inv" | cut -d= -f2 | sort -u)" | ||
| hosts="$(cd "$root/ansible" && ansible -i "$inv" 'ethereum_node:bootnode' --list-hosts 2>/dev/null | tail -n +2 | tr -d ' ')" |
There was a problem hiding this comment.
🟡 pre-genesis gate always fails on the nimbusel node
The gate queries ethereum_node:bootnode, which includes lighthouse-nimbusel-1. nimbusel has no snapshot, gets no static peers (shadowfork_el_static_enodes is empty) and its snap-sync needs a post-Amsterdam pivot (nimbusel.yaml), so it cannot be at shadowfork_height before genesis; the loop reports "no RPC answer" (line 59) or a head mismatch and exits 1. Exclude nimbusel from the check (or special-case it) or the gate can never pass as documented.
It has no snapshot and only starts after Gloas, so it cannot be at the shadowfork head before genesis (redpandabot review). The disk check still covers it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
41 droplets: full 6x6 CL/EL matrix (36 x 1000 validators), bootnode, 2 buildoor and 2 nimbus-el. Applied 2026-10-07 10:46/10:57 UTC (49+52 resources; the 12 earlier droplets got tag updates only). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo
Changes over the sepsf-2 copy (e14d8ea). Every EL restores benchmarkoor's
jochemnet/24402727 datadir instead of a canonical snapshot. jochemnet started
from the mainnet 24350000 snapshot and ran 1 s blocks of bloat state up to
24402727 (head ts 2026-01-31). Its head is not on mainnet.
Snapshots (snapshots.ethpandaops.io/jochemnet//24402727/), layouts
checked from the first 64 MB of each tarball:
storage), erigon snapshot-pruned.tar.zst. Sizes 0.7-1.26 TB compressed, so
the 1.8 TB droplets (DO's largest local disk) stay.
/data/ethrex/chain-1, which is where ethrex 29 keeps a chain-id-1
custom-genesis db. It runs --syncmode=full; the May db is migrated on start.
beat group_vars) to a template default, so it can be set per EL.
Genesis (egg 6.2.0, CHAIN_ID=1, mainnet deposit contract):
the erigon tarball is the stale mainnet 24350000 block. SHADOW_FORK_FILE is
therefore files/shadowfork_block.json, the full-tx eth_getBlockByNumber of
a restored EL; TODO before playbook.yaml. Tested with a stand-in head: the
real genesis role keeps osaka/bpo1/bpo2 at mainnet times and puts
amsterdamTime at genesis + 40 epochs. genesis.ssz builds; a hash-only block
makes eth2-testnet-genesis fail.
deploy-eip8282-contracts assertoor startup test is back. It sends the same
0x4e59 factory calldata as the sepolia deployment txs; checked with cast
create2.
Clients:
tarball's mainnet/flat layout).
is 2TB. The sepsf-2 --snap.stop host_vars are dropped: segments are frozen
up to 24401000, so ~1.7k blocks are left to retire (sepsf-2 had 220k).
firewall EL p2p fleet-wide: mainnet peers have none of our blocks.
Other:
glamsterdam-msf-1.
Launch TODOs (marked TODO(msf-1)): genesis time, gloas_fork_epoch,
files/shadowfork_block.json, shadowfork_el_static_enodes.
🤖 Generated with Claude Code
https://claude.ai/code/session_01NnamtoTL7VJzozbP63tWdo