Skip to content

Latest commit

 

History

History
42 lines (36 loc) · 53.5 KB

File metadata and controls

42 lines (36 loc) · 53.5 KB

This is an upgrade project to bring vutuv, a legacy Phoenix application, up to the latest Elixir and Phoenix Framework.

vutuv is mostly a classic Phoenix controller + HTML module + .html.heex template app (Phoenix 1.8 *_html.ex modules with embed_templates, living in lib/vutuv_web/views/* — there is no phoenix_view dependency), and LiveView is being adopted incrementally on top of it. Most pages are still controller + template; do not rewrite a controller page as a LiveView unless a task explicitly asks. There is still no core_components.ex.

What is already LiveView (the real-time shell, see docs/architecture/realtime.md): the app shell VutuvWeb.ShellLive (top bar + mobile bottom tab bar with live unread badges) is embedded in the shared app layout via live_render and shows on every page; the Messages (/messages) and Notifications (/notifications) pages are LiveViews under a live_session; the profile (/:slug, VutuvWeb.UserProfileLive) is a LiveView embedded by its controller via live_render (the controller keeps owning agent-format negotiation, so the .md/.txt/.json/.xml/.vcf siblings are untouched) — every state-changing control on it is reload-free (follow pill, the header card's bookmark/like glyph toggles, ⋯-menu mute/block, list follow buttons, tag endorsements) and its counts/tags update live over PubSub; the post permalink's conversation (VutuvWeb.PostLive.Thread, embedded the same way) renders a long thread as a window around the permalinked post whose "Show earlier / more" expanders load the rest over the socket (Vutuv.Posts.thread_window/3; agent formats keep the whole capped thread); real-time updates flow over Vutuv.Activity (PubSub on "user:<id>") and VutuvWeb.Presence. The layout is split into root.html.heex (document shell) + app.html.heex (chrome), shared by both dead and live pages. When you touch the shell, layouts, or real-time features, LiveView is expected. The email chokepoint and CSRF/PIN rules below still apply unchanged.

Framework conventions (Elixir, Ecto, Phoenix, HEEx, LiveView, assets) live in .claude/rules/ and load automatically only when you edit a matching file, so they stay out of context the rest of the time. Use the VutuvWeb.UI components and the assets/css/components.css reskin rather than inventing new styles.

Project guidelines

  • Use the mix test alias when you are done with all changes and fix any pending issues (it runs ecto.create + ecto.migrate first).

  • New site pages live under /system/, never at a new root path word. Profiles own the URL root (/:slug), so every new root segment (a listing, a directory, a tool, a stats page) permanently burns a word members could otherwise claim as a handle and must be listed in Vutuv.Accounts.ReservedSlugs. system is already reserved: route new site pages as /system/<name> (the member directory at /system/members set the pattern) — no new ReservedSlugs entry needed. Before ever claiming a new root word instead (only for a genuinely member-facing top-level feature, agreed with Stefan first), check the production data for members already holding it as a username (query the dev DB, a prod copy).

  • Every id is a UUID v7 — nothing else. All primary keys and foreign keys use Vutuv.UUIDv7 (set once in use VutuvWeb, :model; never override @primary_key/@foreign_key_type per schema). Migrations default to :binary_id via the repo config in config/config.exs. Mint ids in code with Vutuv.UUIDv7.generate/0 — never integer ids, never UUID v4, never Ecto.UUID.generate/0. Inside fragment/1 wrap pinned id params as type(^id, Vutuv.UUIDv7) (fragments can't infer the type). Id order matches creation order (the timestamp is in the id), so keyset tiebreakers on id keep working. A regression test (test/vutuv/schema_uuid_chokepoint_test.exs) fails the build on drift.

  • Never store a bare (unsalted, unkeyed) hash of low-entropy or guessable user data — key it with a server secret. SHA-256 (or any fast hash) of a value an attacker can enumerate — an email, phone number, username, postal code, PIN, small integer id — is not privacy: whoever reads the table (or a backup) recomputes hash(guess) and confirms a match, so the "we only store a hash, a leak reveals nothing" claim is false. This is exactly what shipped in the invitations table (email_hash = SHA256(email)) until issue #942. Instead key the hash with a pepper derived from secret_key_base so a DB/backup leak alone can't recompute it: for a deterministic dedup/lookup key use :crypto.mac(:hmac, :sha256, pepper, normalized) |> Base.encode16(case: :lower) where pepper = :crypto.hash(:sha256, "vutuv/<purpose>/pepper/v1" <> secret_key_base) (see Vutuv.Invitations.hash_email/1 and the login-PIN pepper in Vutuv.Accounts); for anything password-like also add a per-row random salt. Real random tokens (Vutuv.Token.random_token/1, ~165 bits) stay fine with a bare SHA-256 — the entropy, not a key, is what protects them. Second, hashing plaintext out of the database does not remove it from other surfaces: the mail server still logs recipient addresses in cleartext, web/access logs carry query params, etc. When you reason about "who could learn X", enumerate every channel (DB, backups, mail logs, request logs, error trackers), not just the column — and treat those logs as equally sensitive (retention/access).

  • Use the already included :req (Req) library for HTTP requests; avoid :httpoison, :tesla, and :httpc.

  • Making an existing NOT NULL column nullable is never just a migration: walk every IN / NOT IN, every inner join and every Repo.get on it first, because each one fails SILENTLY or fails LOUDLY and you need to know which. This is the standing cost of the nullable-pair model the organization milestone chose (posts.user_id beside organization_id, follows.followee_id beside followee_organization_id, issues #1334/#1336), and it bit five times in one milestone. Three shapes, each with its own tell. (1) x NOT IN (…, NULL) is never true in SQL, so one NULL in a subquery's result set makes the whole predicate false for every row — the feed's discovery rail returned [] for anybody who followed an organization, and mention notifications from a page vanished from both the list and its unread count. No error, no log. (2) An inner join to the old owner table drops the new rows — the same mention list joined users on the author and simply never saw an organization post. (3) Repo.get(Schema, nil) and where: x == ^nil RAISE rather than answering nothing, so those fail loudly, and the crash can hide behind a gate: Webhooks.emit/3's nil-comparison only ran once an installation had its first subscriber for that event, so liking an organization post would have 500ed in production while passing every test. Practical rules: give the widened column an explicit is_nil(x) or … branch wherever it feeds a NOT IN; build such id lists in a named private function (followees_of/1) rather than an inline subquery, so there is one place to fix; make nil a no-op at the notification/delivery chokepoints (Activity.notify/2 and broadcast/2 already do, Webhooks.emit/3 now does); and calibrate the regression test against the un-fixed code — a NULL-trap test that passes either way is worth nothing, so revert the fix once and watch it go red; revert via a patch you re-apply, never git checkout -- — with the fix still uncommitted that discards it too.

  • A visibility or permission clause must pattern-match on a COLUMN, never on a preloaded association. visible_to?(%Post{organization: %Organization{}}, …) looks right and silently does not match the bare %Post{} that half the callers hand over straight from a query — which dropped organization posts into the member branch and Repo.get(User, nil), so liking one crashed. Match %Post{organization_id: id} when is_binary(id), and when the answer genuinely needs the associated record, take the preload if it is there and fall back to a lookup if it is not (organization_author?/2). A permission answer that depends on whether somebody remembered a preload is not a permission answer. The same rule covers a URL: one function owns a thing's path (Vutuv.Posts.path/1), and every hand-built ~p"/#{post.user}/posts/#{post.id}" at a call site is a place the next author kind has to be remembered again. The same goes for its text (Vutuv.Posts.text/1), which the three post kinds hold in two differently named columns — that choice sat in three modules and had already drifted apart on nil. That is not hypothetical — it broke a member's own notifications page with cannot convert nil to param the moment a page mentioned them, two deploys after the column that made it possible. Grep for hand-built copies whenever you widen what a thing can be, and route them through the owning function; and remember the two halves travel together, because the owning function needs the association preloaded (visible_posts_by_ids/2 was loading :user alone).

  • Every user-writable :string field needs a validate_length matching its column (varchar(255) unless the migration says otherwise). Ecto does not enforce column limits, so without the validation an oversized value raises Postgres 22001 (string_data_right_truncation) — a 500 from a plain form submit, and inside a multi-insert transaction (the LinkedIn import) it aborts everything after it. This exact gap 500ed real imports until v7.38.4: description columns were varchar(255) while LinkedIn allows 2,000-char descriptions (those two are text now). Also validate generated columns that derive from user input (work-experience/education slug builds from title + organization, so it can overrun its own varchar(255) even when each source fits). Prose fields that legitimately run long belong in text columns with a generous cap (the descriptions use max: 10_000), not in varchar(255). And when a new column stores a value copied from an existing one, give it that column's type — look the sibling up, never default to :string. The new fediverse_post_deliveries.inbox_uri (issue #1102) copies fediverse_followers.inbox_uri, which is :text because a remote server's inbox address is not ours to bound; declaring the copy varchar(255) would have raised 22001 on the post-publish path, i.e. a 500 on posting, for any follower with a long inbox URL — and no changeset validation would have caught it, because nothing user-facing writes that column. If the new column also carries a btree unique index, check the combined entry against Postgres' ~2704-byte limit and note in the migration which validation bounds it (both URI sources cap at 2048 bytes).

  • Member uploads land in the repo working tree in dev/test, so every upload storage-dir root must be gitignored — never committed. config :vutuv, :uploads_dir_prefix is empty outside production, so every uploader in lib/vutuv/uploaders/* writes into the checkout (avatars/, qualification_documents/, organization_images/, …). A tree missing from .gitignore gets swept into a commit by a routine git add -A after a manual upload smoke test — which is how two dev-side credential-proof scans (a PDF + a JPEG) reached the public repo until #1037. When you add an uploader, add its served storage-dir root to the upload-tree block in .gitignore and to @upload_trees in test/vutuv/uploads_gitignore_test.exs (the private originals/ copies are covered by the single /originals rule, so they need no per-type entry); that test fails the build if any known tree is tracked or un-ignored, and would have caught F21 the moment the files were committed. A file that reaches git history is not un-leaked by a later git rm: it stays in every old commit, clone and fork (and in GitHub's blob store), so on a public repo the only real fix is to never commit it — a history rewrite is disruptive hygiene, not a recall. Treat any real member file that does slip in as already-disclosed.

  • Never render a bare integer count in the UI — always format it. Route every user-facing number through a formatter from VutuvWeb.UI (imported into every view/component): compact_count/1 for a glanceable badge (exact up to 999, then 1K / 60K / 5M) or delimited_count/1 for an exact, locale-grouped figure (German 60.023, English 60,023). A run-together integer like 60023 is a bug. compact_count/1 is the default for counts; reach for delimited_count/1 when the precise number is the point (membership totals, money, audit figures). Gotcha: ngettext/3 auto-binds %{count} to the raw integer and a count: binding does not override it, so to show a formatted number inside a pluralised string use a separate placeholder (e.g. %{formatted}) bound to the formatter, or render the count outside the gettext call (a plain delimited_count/1 in the markup, as the admin dashboard's member tile does).

  • vutuv is installable by third parties — design every feature for other installations, not just vutuv.de. Before shipping any feature or change, answer explicitly: what does this mean on someone else's installation (internet or intranet), and how do they configure it? Concretely: (1) No vutuv.de assumptions — derive host/URLs from PHX_HOST/Endpoint.url(), never literal vutuv.de/www.vutuv.de; anything naming the operator (email From, operator-notice recipients, footer credit, postal address, security.txt contact) reads the "Operator identity" block in config/config.exs (env-overridable in config/runtime.exs; defaults are the vutuv.de values, so our production needs no .env entries). (2) New knobs get the right seam: per-installation values → env var read in runtime.exs with the vutuv.de value as default; on/off product switches → a config flag in config.exs (like :ads_enabled); per-installation content → data with an admin UI (like the legal pages: Vutuv.Legal, trusted Markdown at /admin/legal, rendered with DevDocMarkdown.to_html(body, breaks: true) — Earmark's ast pipeline escapes inline HTML, so trusted Markdown needs real line breaks, never <br/>). Never require a source edit to run vutuv elsewhere. (3) Outbound network calls need a flag (intranet installs run air-gapped — see :fetch_gravatar, :fetch_mastodon_posts, :generate_screenshots) and external services must degrade gracefully when unreachable. (4) Document it: a new env var or flag is not done until it is in the docs/ADMINS.md reference (env-var table, and the intranet section if it calls out), and a new page under /admin is not done until the dashboard or its parent page links it — /admin/ads/discounts shipped with neither, so only somebody who knew the URL could make a discount code; operator-relevant behavior changes update that manual, developer-facing ones the matching subsystem file under docs/architecture/ (one document per subsystem, index in docs/architecture/README.md; docs/DEVELOPERS.md stays the slim entry point for setup, tests and deployment). The README stays a short entry point and makes no capacity claims (no "N profiles on one server" numbers — we don't know the limit); the scaling message is: one modest server goes a long way, and Elixir scales out to multiple nodes for very large installations.

  • Recognising one of our own URLs: match on host + path, never on a prefix of Endpoint.url() — and www. is us. A member pastes whatever their browser, their mail client or a share button handed them, and that is the same page in half a dozen spellings: the www. alias, a trailing slash, an appended ?utm_source=, a fragment, a shouted host, plain http, a dev port. Every one of them names the page and every one of them misses a whole-string prefix match — and the failure is never a polite "not found", it is a fall-through to the foreign path. That is exactly how the fediverse post lookup shipped: a pasted https://www.vutuv.de/<slug>/posts/<id> was treated as another server, so this installation sent itself a signed ActivityPub GET, got nginx's 301 (which ap_get deliberately does not follow), and told the member their own site was unreachable — after spending a slot of their hourly budget on it (v7.197.0, fixed in v7.198.3). So parse with URI.parse/1, ask Vutuv.Fediverse.local_host?/1 whether the host is ours (it strips a leading www. on both sides: serving a site at both the apex and its www. alias is the oldest convention on the web, and nothing about that is vutuv.de-specific, so third-party installations get it too), and read the path with String.split(path, "/", trim: true), which drops a leading and a trailing empty segment in one go while the query and the fragment are not in path at all. URI.parse/1 is lenient: https://a.example\@b.example/ yields host b.example while a browser opens a.example. Before trusting a host, refuse a backslash, userinfo, or anything URI.new/1 rejects (ChangesetHelpers.web_url?/1). Then resolve the record by the id in the path, not by the handle beside it: a handle goes stale the moment its owner renames, and landing on the post beats a dead end. The one place that legitimately demands handle and id agree is local_note_post/1 on the boost path, because it decides whether a member's post may be redistributed on a remote actor's say-so rather than where to send a reader who is already here. The same care applies to every other "is this us" test — the follow gate, the search page's follow offer and own_object?/3 all read local_host?/1, and all three had the wrong answer for a www. address until it learned this.

  • Email is sent through one chokepoint. Build every message from Vutuv.Notifications.Emailer.base_email/0 and send it with Emailer.deliver/1. Never call Vutuv.Mailer.deliver/1 or use Swoosh directly outside Vutuv.Notifications.Emailer. base_email/0 stamps the From and the auto-generated robot headers (Auto-Submitted, X-Auto-Response-Suppress) that keep out-of-office responders silent. Builders return a %Swoosh.Email{}; the single deliver/1 sends it. A regression test (test/vutuv/notifications/mailer_chokepoint_test.exs) fails the build if anything bypasses this. For bulk mail, add Emailer.bulk_headers/1. A new email needs BOTH bodies — lib/vutuv_web/templates/email/<name>_<locale>.text.eex and lib/vutuv_web/templates/email_body/<name>_<locale>.html.heex, per locale — and email_html_drift_test.exs fails the build on a missing half. Beware the "compiled from a wildcard" trap in general: a module that builds functions from Path.wildcard/1 at compile time tracks each matched file as an @external_resource, so editing one recompiles it but adding one does not — and since CI caches _build, a brand-new template can pass a full local mix precommit and then fail CI with function …/1 is undefined (exactly what happened to the username-change PIN mail, issue #1086). Such a module must define __mix_recompile__?/0 comparing the current wildcard to the compile-time one (see VutuvWeb.EmailText); when you add a file to any wildcard-compiled directory, verify with mix compile on an already-built tree, never only after a --force.

  • Green tests are necessary but not sufficient; smoke-test the critical paths in a real browser before deploy. The test harness does not behave like production in three ways that have already shipped bugs: Phoenix.ConnTest sets plug_skip_csrf_protection on every conn (so a plain post/3 never checks CSRF — this hid a login 403, issue #759), the Swoosh test adapter never renders or sends a real email, and a ConnTest that PUTs a hand-built path never exercises the form's rendered action= — this hid that every Save button on /settings/privacy and /settings/notifications posted to retired /:slug/settings/* URLs and 404ed in production from v7.34 until v7.42. So after any change to auth, sessions, forms, or email, drive the affected flow end to end in a browser (the verify / run skills help) and read the actual email — and when a test covers a form submit, assert the rendered action= (or submit through it) instead of hardcoding the route you know exists. Smoke-testing from a fresh git worktree needs three setup steps first, or you are testing a different app than production: a worktree carries neither deps/ nor a database of its own (each one gets its own since #1876), so mix deps.get and mix ecto.create && mix ecto.migrate come before anything else, and without them mix phx.server dies inside Vutuv.Prefs.Cache with a DBConnection.ConnectionError about a pool that "cannot serve requests fast enough", which reads as a busy database and is an absent one; priv/static/assets/* and assets/node_modules are gitignored, so until mix assets.setup && mix assets.build runs, /assets/app.js 404s and every page renders with no JavaScript at all (a form falls back to a full submit, LiveView never connects) — which reads exactly like a broken enhancement; and the upload trees are gitignored too, so every avatar 404s until the worktree is linked to the shared image store — which scripts/restore-snapshot.sh --link-only does, and its nightly run does anyway. Leave those symlinks alone. They are infrastructure, not a smoke-test convenience: :uploads_dir_prefix is empty in dev, so Plug.Static from: "avatars" and Vutuv.Uploads.disk_dir/1 resolve every image path against the running server's own directory, and a worktree without the links serves 404s for every picture while the files sit in the main checkout. Deleting them does not protect anything either — test/vutuv/uploads_gitignore_test.exs lstats the link instead of reaching through it, so it runs green with all thirteen in place; "beyond a symbolic link" never fires when the target is inside the same repository, and the run is silent. The real hazard is an upload as an existing member, which writes through the link and replaces that member's served sizes and their private original, which no derivation brings back (that destroyed a real avatar on 2026-09-07, and a snapshot does not carry originals/, so a restore cannot undo it). So drive every picture smoke test as a freshly registered throwaway account, never an existing member, page or review. frozen/ is the one tree that is not linked, so a takedown tried in a worktree moves the store's files into it — reversible with Vutuv.Images.unfreeze/1, and gone for good if the worktree is torn down first (#2066). Confirm the bundle loaded (typeof window.liveSocket) before concluding that any JS behaviour is broken. A browser smoke test is not a substitute for reading the request log: the browser tooling swallows the odd first click and does not always submit on Return, so a step that "did nothing" is usually the tool, not the app — check POST/GET lines and the Processing with … action in the server log before you start debugging code that is fine. And when a wizard-style flow fails only in the browser, suspect a GET that mutates state: an action that clears session state (a pending change, a draft) is destroyed by any accidental navigation — Back button, sidebar link, breadcrumb, link prefetch — none of which a ConnTest ever performs. Keep GETs safe and age the state out instead (this exact shape produced a bogus "this confirmation expired" in the username rename, issue #1086). An HTTP-stub (plug:) test must answer with the real API's content-type (put_resp_content_type("application/json"), never a bare send_resp): Req's decode_body step branches on that header, so a content-type-less stub hands the client a binary while the real server's answer arrives as a decoded map. That gap let v7.95.4 drop decode_body: false when it added the into: streaming collector — into: does not disable the decode step, so a client that decodes for itself must keep decode_body: false beside the collector — and every Mastodon/Bluesky feed and code-forge stats fetch failed in production for 18 days while the whole suite stayed green, each account silently walking the backoff ladder to permanent fetch_disabled_at (fixed 2026-07-30). Never stop a dev server with a pattern that can match another one — several worktree sessions run their own mix phx.server on their own port against the shared vutuv1_dev DB, so pkill -f "mix phx.server" kills a colleague session's server mid-smoke-test (done 2026-07-26, port 4009 went down with mine). Start yours on a port of its own, remember the background task id, and stop it by that id or by kill <pid> resolved from lsof -nP -iTCP:<your port> -sTCP:LISTEN — never by a bare process-name pattern. The same "shared dev DB" caution applies to the data: if you mutate another member's row for a check (pinning their post to look at a page), restore it in the same step. And what a browser sees is not always what you deployed: markup a LiveView newly streams meets the previous release's CSS and JS in every tab open across a deploy — a deploy reloads nothing, the socket simply reconnects to the new release and patches into an hours-old document — so anything that brings its own stylesheet or hook belongs behind static_changed?/1 (phx-track-static on both assets in root.html.heex is what lets it answer). The feed's tab ticker shipped without that gate and drew as an unstyled 200-character paragraph across the tab bar that no clock ever took away (v7.347.0, fixed in v7.347.1).

  • Before you loosen a heuristic that classifies stored content from elsewhere, run the old and the new rule over the dev database and count the changes in BOTH directions. That database is always yesterday's copy of production, so the real corpus is already on this machine and the measurement costs a psql dump plus one script — reach for it rather than reasoning about what the rule should match. A calibrated regression test only proves the case you thought of: widening the closing-hashtag-line rule so tagesschau's #FlughafenLeipzig/Halle would fold into chips was green, calibrated and still a net regression — over 5,398 cached fediverse posts it fixed 8 posts and made 12 worse, because authors glue their closing link straight onto the last tag (#linuxhttps://flathub.org/apps/…) and lifting that line deleted the URL from the card. Read the losers, not just the total, and name the invariant the rule must not break ("lifting a line is a move, never a delete") — that is what turns a count into a rule.

  • Locale is a test and reproduction dimension; a plain English curl hides German-only crashes. vutuv is a German site: real visitors send Accept-Language: de, but Phoenix.ConnTest and a bare curl default to en, so a bug that only fires in the German render passes every English check. This shipped a real production outage: the logged-out landing page 500ed for every German visitor (a .po merge duplicated the German consent sentence, and the template's hard [a, b] = String.split(gettext(...), "{placeholder}") raised on the doubled placeholder) while curl / and ?lang=de-without-header both rendered English and looked fine. Two standing rules from it: (1) when a page 500s in production but you cannot reproduce it, vary the failing browser's real request dimensions before concluding it is transient — first curl -H "Accept-Language: de-DE,de" <url>, then Accept header and cookies; do not write it off as a deploy blip until the German render is green. (2) Never hard-pattern-match on a translated (or otherwise user/data-influenced) string — [a, b] = String.split(gettext(...)) turns any .po slip into a 500; split with parts: 2 and a fallback (see VutuvWeb.UI.split_marker/2). Cover public pages across every config[:locales] via an Accept-Language header, not just the default (see page_locale_render_test.exs). This compounds the gettext extract/merge hazard: a merge can duplicate a msgstr (concatenate two variants), not only prune translations. And worst of the three, mix gettext.extract --merge FUZZY-FILLS a brand-new msgid with the translation of some unrelated string it looks similar to — it does not leave it empty. Adding this page's labels produced "Now" → "Nein" (No), "Nothing has changed yet." → "Noch nichts Neues.", "Not yours yet." → "Noch keine Reposts." Each is flagged fuzzy in the .po and nothing fails the build, so a German page ships confident nonsense while every English test stays green. After any gettext.extract --merge, git diff priv/gettext and grep the new entries for fuzzy — treat every one as untranslated and write it yourself, then delete the flag. Grep for ", fuzzy", never for "#, fuzzy": the flag line this tool writes is #, elixir-autogen, elixir-format, fuzzy, so the obvious pattern matches nothing and reports a clean file. That false negative nearly shipped "This page does not follow anyone yet." as "Diese Seite gibt es nicht" (this page does not exist) and "Members and organizations" as "Organisation hinzufügen" (add organization) on 2026-08-11; grep -c ", fuzzy" on that same file said 12. And msgstr "" followed by indented continuation lines is the ordinary multi-line form, not a missing translation: writing into it appends a second copy that gettext concatenates into one doubled sentence, so read the line after any empty-looking msgstr before filling it. A second trap sits beside it, and grep cannot see it at all: reusing an existing msgid whose German was written for a different speaker. gettext("Following") is translated "Folge ich" (I follow) for a member's own list, so putting that msgid under an organization's nav made the page claim the reader follows these accounts when it is the page that does. A msgid is a key, not a phrase: when the same English word is said by a different voice or about a different subject, give it its own msgid (here "Follows") rather than sharing one and hoping the translation fits both. Assert the new German strings by name in a test (Accept-Language: de-DE,de), including the short labels: a one-word msgid like "Now" is the likeliest to be fuzzy-matched and the least likely to be noticed. And a change of wording or address is not finished in the catalogs: lib/vutuv_web/templates/email/*_<locale>.text.eex, lib/vutuv_web/templates/email_body/*_<locale>.html.heex, priv/help/*_<locale>.md and the per-locale branches in user_helpers.ex (email_greeting/1, image_kind_label/2) and email_components.ex carry translated prose gettext never sees, so a sweep over priv/gettext alone leaves them behind — 58 of the 61 Italian mail templates still said Lei after both catalogs had moved to tu (#2240), which is precisely the seam a member crosses walking from a post to the mail about it. Grep the whole tree for the wording, not just the .po files. And an edited .po changes nothing the running app reads until the catalog is recompiled — gettext compiles the translations into a module, so the dev server keeps serving the old string and you debug a translation that is already correct (this cost a whole round on the ad mails). Force it with touch lib/vutuv_web/gettext.ex && mix compile; deleting the .beam is not enough and leaves the module undefined.

  • Public pages have agent-format siblings — keep them in sync. Every public page (profile, post permalink/archive, follower/following/connections lists, every profile section page — work_experiences, links, social_media_accounts, addresses, phone_numbers, emails, tags — and their single-entry show pages, tag pages, the most-followed listing) is also served as Markdown / plain text / JSON / XML (and the profile as vCard) under the same URL plus .md/.txt/.json/.xml/.vcf, or via Accept negotiation — see VutuvWeb.AgentDocs. All formats render the anonymous public view from one doc map per page (VutuvWeb.AgentDocs.*Doc). If you change what one of these HTML pages shows, change its doc builder too; test/vutuv_web/agent_docs/agent_docs_drift_test.exs fails on drift. New public pages should join the system (controller calls AgentDocs.respond/2 with an :html and a :doc fun; only intricate flows branch on negotiate/2 themselves) and be listed in /llms.txt (PageController.llms). Never let a .md URL serve HTML — the endpoint plug VutuvWeb.Plug.AgentFormat 404s unhandled extension requests by design.

  • Never give a Phoenix <.form> an id just so a button elsewhere can reach it with form=. The id becomes the prefix of every input id the form builds (user_avatar turns into profile-form_avatar), which breaks every label for=, data-crop-target and test that names a field. Put the button inside the form instead; a <dialog> may sit there too.

  • Two-step PIN flows must survive CSRF. Login, email change, and account deletion each render a form whose CSRF token lives in the session, so step 1 must not drop the session (renew it instead — see Vutuv.Accounts.login_by_email/2). Cover such flows with the CSRF-enforced submit_with_csrf/3 helper in VutuvWeb.ConnCase, not a bare post/3; see test/vutuv_web/controllers/csrf_pin_flows_test.exs.

  • Authenticate every LiveView socket from the session token, never a bare session["user_id"]. Any LiveView that decides who the viewer is — especially the off-router live_render children a controller or the layout embeds (ShellLive, PostLive.Actions, SectionReorderLive, and the profile/CV LiveViews) — must resolve identity through VutuvWeb.Live.InitAssigns.session_user/1 (Sessions.active_session/1 on the cookie's session_token + the Moderation.login_block/1 gate), the same source of truth VutuvWeb.Plug.ConfigureSession uses on every HTTP request. Reading session["user_id"] instead trusts a value that (a) is never re-checked against server-side revocation, so a remotely logged-out device (#794) or a suspended member keeps socket access, and (b) also rides the controller-curated live_render session map, which is signed but not encrypted and valid for days, so a captured payload can be replayed to authenticate as the member it names. Nested live_render children do receive the cookie session merged into their mount session (Phoenix.LiveView.Static + Channel), so session["session_token"] is always available — never thread a bearer credential or a trusted user_id through a curated map. The one allowed use of the curated user_id/display fields is the throwaway dead (disconnected) render, which its own HTTP request already authenticated and which is replaced the instant the socket connects and re-checks the token (see ShellLive.mount_static/3 vs mount_authenticated/3). This bit twice — #1034 (the two router mounts) and #1036 (these three embedded children); resolve on connected?/1 for a per-post-card bar so a many-card page pays one lookup per card only on connect, not on the discarded dead render. (Ecto gotcha the fix surfaced: where: x.user_id == ^nil raises, it is not a silent no-op — guard a write handler for an unresolved (nil) owner before scoping an Ecto query to it.)

  • A background sweeper that picks work "least recently done first" must advance that clock on EVERY outcome, including the ones where it does nothing. vutuv runs ~18 such sweepers (CountsRefresher, NoteSweeper, FollowerPruner, Deliverability.Sweeper, the recheck sweepers …), and they all share one shape: select a bounded batch ordered by a *_checked_at/*_at column, do the work, stamp the column. The trap is the branch that decides the item cannot be worked on at all — no signer, no inbox, nothing to do — and returns a :skip that writes nothing. That item is due again on the very next run, forever (nothing about it changes in two minutes), and because the ordering is oldest-first it now holds the front of every batch, so the batch cap is spent on work that can never complete. That is not a slowdown, it is a deadlock: Vutuv.Fediverse.refresh_counts/1 made zero real requests for hours while 307 objects were due, and every fediverse post from the current day showed 0 likes (v7.228.1, #1316). Three things to carry over: (1) stamp the clock on a skip too — it is the scheduler's clock, not a claim that the work happened — and take a strike only when the remote side failed, so a transient impossibility can still be retried; (2) a nil clock that means "always due whatever the age" (a one-off backfill rule) is permanent poison when combined with a skip, since the item can never leave that state; (3) any per-host/per-tenant fairness cap applied to an already-sorted list is an amplifier — the starved items at the front spend their host's whole quota, so one blocked host silently blocks every healthy item behind it. Cover it with a test that asserts an unworkable item is no longer returned by the due query after one pass.

  • Work that outlives one request must be started again after it dies, and here that means a row saying it is unfinished plus a sweeper that acts on it. Every long job here is fire-and-forget on Vutuv.TaskSupervisor (newsletter sends, fediverse deliveries and media fetches, screenshots, image scans), and a blue/green deploy stops the old slot mid-loop while the member sees nothing and no error is logged anywhere. Copy the shape that already survives it, Vutuv.Newsletters.BroadcastResumer: the row carries the state (status: "sending"), each recipient gets its own newsletter_deliveries row as it completes, a sweeper (once a minute) resumes every send with no delivery activity for five minutes, and resume_broadcast/1 CAS-locks on updated_at so the two slots of a deploy overlap can never work the same send. Three consequences. A resumed run must skip what already has its per-item row — never re-send, never re-notify — and report the whole protocol rather than its own pass (broadcast_sent_count/1, not the recipients this run handled). The staleness window is what makes a resume safe during the deploy overlap, so it must be longer than one item takes, not shorter. And the test has to interrupt: kill the task mid-loop, then assert the sweeper finishes exactly the rest — a test that never interrupts anything proves nothing about a deploy. The same question decides where periodic work belongs: a Process.send_after loop holds nothing across a restart, so the schedule may live in memory but the due list must be a query (see the sweeper-clock rule above).

  • Migrations must be backward-compatible for exactly one step (N-1). Production deploys are blue/green (scripts/deploy.sh): migrations run while the previous release is still serving traffic, and it keeps serving the migrated schema until the new release passes /health and nginx switches. So every migration must keep the currently deployed release working — one release back, no further; the release after next may assume the change. Removing things therefore takes two deploys (expand/contract): deploy 1 ships code that no longer touches the table or column, deploy 2 ships the migration that drops it. Never drop, rename, or repurpose a table or column in the same deploy whose code stops using it. Plain additions (new tables, new nullable columns, new indexes) are fine in one deploy. A migration that cannot be made N-1 compatible is a deliberate planned-downtime deploy, agreed with Stefan beforehand and never pushed casually to main (as was done for the one-time UUID v7 re-key, which shipped to production on 2026-06-18). Column-type changes (even safe widens like varchar → text) invalidate the old release's cached prepared statements: Postgres answers them with 0A000 "cached plan must not change result type" until the connection re-prepares. The repo config sets disconnect_on_error_codes: [:feature_not_supported] so the pool heals itself, but still expect a one-request blip per affected statement on the old slot during the switch window (observed: exactly one request during the v7.38.4 description-widen deploy) — fine for a widen, but don't stack type changes on hot tables into one casual deploy.

  • A PR is opened only on Stefan's word — and once it is open, the merge is not a second decision. Most session work is never meant to reach GitHub at all: without an explicit "ship it" / /deploy / "mach einen PR" / "merge das" in that request, finished work is committed on its branch and pushed nowhere, however green it is. But the moment a PR exists, opening it was the decision to deploy, so never ask "soll ich mergen?" and never end a turn with a green PR left open. Across 30 merged PRs the normal open→merge time is 8 to 10 minutes — exactly one CI run, i.e. merged the second it went green; the eight that took 27 minutes to 8 h 36 are all this failure, and two of them were finally merged by a different session that found them lying around. So watch CI as a background task (gh pr checks --watch --fail-fast) and keep the turn alive until it answers: the nine-minute wait is never a reason to hand the PR back. Wait for that one notification — never poll it with repeated sleep+check calls, which cost tokens without advancing the wall clock. Green → merge at once, whatever the hour, hot and cold deploy alike. Red → at most two repair rounds (fix, push, watch again); if it is still red after that, leave the PR open and report the failing job by name. Tell peer sessions you are merging, never wait for their reply. And if you genuinely cannot finish — context exhausted, a hard block — the last thing you say names the orphan: "PR #N is open, CI green, merge it with gh pr merge N --squash --delete-branch". A PR nobody is told about is the failure this rule exists to prevent.

  • /deploy in this repo means: open a PR, wait for CI to go green, then merge it — never a direct push to main. A push to main fires .github/workflows/deploy.yml and ships to production straight away, so the only path to main is a merged, CI-green pull request; the merge commit is the deploy trigger. The full flow lives in the project skill .claude/skills/deploy/SKILL.md, which overrides the personal deploy skill in ~/.claude/skills/deploy/ (that one commits and pushes to the current branch, which on main would deploy unreviewed) — when both are offered, take the project one. Short version: branch → mix precommit green → commit → git push in a Bash call of its own (the PreToolUse hook .claude/hooks/precommit-before-push.sh aborts the whole command on a red precommit, so a chained git push && gh pr create dies with it) → gh pr create → gh pr checks --watch --fail-fast → gh pr merge <nr> --squash --delete-branch → delete the local branch by force and prune. Red CI means fix and re-watch, never merge. Only an explicit, unambiguous "push straight to main" for that specific push skips the PR; a plain /deploy never does. That last step is not optional, and --delete-branch does not do it for you: the flag reliably removes the remote branch but routinely leaves the local one, because the squash replays the work as one new commit, so the branch tip never becomes an ancestor of main and git's safe delete refuses it with "the branch is not fully merged". Every deploy then leaks a branch, which is how 16 stale ones piled up by 2026-07-26. After the merge run git checkout main && git pull --ff-only, then git branch -D <branch>, then git fetch --prune origin — force is correct, not risky, since the PR is merged and the content is on main. In a worktree session the branch is checked out there and git refuses to delete a branch checked out anywhere, so tear the worktree down instead (ExitWorktree / git worktree remove <path>) and then git branch -D <branch> in the main checkout: the teardown only removes the checkout, which makes the branch deletable — it does not delete it. Neither teardown knows about Postgres, so the worktree's own dev and test databases (vutuv1_dev_<name>, vutuv1_test_<name>) outlive it: run cwgc afterwards, which lists every worktree database with no worktree behind it and drops them with cwgc -y. That is the path by which 33 orphaned databases and 1.7 GB piled up by 2026-09-02, because a session tears its worktree down from the inside and cwrm never runs. To keep the worktree, git checkout --detach origin/main first and the local branch deletes normally. In that session gh pr merge --squash --delete-branch also reports a failure it did not have — it merges and deletes the remote branch, then dies on its own local checkout step with "fatal: 'main' is already used by worktree at …", because main lives in another worktree. Never re-merge on that message: check gh pr view <nr> --json state,mergeCommit (it will say MERGED) and finish the cleanup by hand. Before you report the deploy done, git branch -vv must show no branch marked [origin/<name>: gone] — that marker is the signature of this leak (remote deleted, local survived). The same squash is why a branch must never be reused for a second PR: rebase it onto origin/main first. Because the squash replays the work as a new commit, a branch that keeps its old commits has both sides claiming the same changes, so the next PR opens CONFLICTING — and a conflicting PR gets no CI run at all, because GitHub cannot build the refs/pull/N/merge ref that on: pull_request workflows check out. The symptom is silence, not an error: gh pr checks says "no checks reported on the '' branch", gh run list shows nothing new, the workflow is active and the repo's Actions are enabled, and closing/reopening the PR changes nothing — a picture that reads exactly like an account-level Actions block or a spent billing limit, which is how it was misdiagnosed on 2026-08-11. So when a PR reports no checks, run gh pr view <nr> --json mergeable,mergeStateStatus before suspecting anything about GitHub: CONFLICTING/DIRTY is the answer, and git rebase origin/main + git push --force-with-lease starts CI within seconds. In a long worktree session that ships several PRs from one branch, rebase onto the freshly merged origin/main immediately after each merge, before starting the next piece of work.

  • Run mix precommit and get it green before every push to main or to any branch headed for a PR — and do not push if it fails. CI runs the whole mix precommit alias (compile --warnings-as-errors, credo --strict, mix format --check-formatted, mix test), so run the whole alias locally first, not a hand-rolled subset. A subset is not enough: a mix test + grepped-credo check once let a Credo.Check.Design.AliasUsage failure reach main and turn CI red right after a green deploy. Note credo --strict exits non-zero on suggestions too — fully-qualified Foo.Bar.baz() that could be aliased — so alias nested modules at the top of the file. (Pushing to main auto-deploys to production via the blue/green pipeline, so red-after-deploy is the worst case: the deploy ships but main is broken for the next contributor.) CI runs on the pull request, and on main only when mix.lock or config/*.exs change — with the squash-merge flow the post-merge run re-tested the tree the PR had just proved green, so it cost a second full run per change and caught next to nothing; what it still earns is the shared build cache, whose key is exactly those two files (a cache written on a PR branch is invisible to other PRs, so only a default-branch run refreshes it). Two consequences to hold in mind: do not wait for a CI run on main after merging — for most changes there will not be one, and the merge itself is the deploy trigger; and the one thing the old run could have caught, two PRs each green alone but broken once merged, now falls to the deploy's own build and /health gate. So the local mix precommit is carrying more weight than it used to.

  • There is no version number to bump — mix.exs derives its version from the commit, and nobody writes a number there. version: is version(), the date of the commit being built (2026.8.29; 0.0.0 without git), and what identifies a release is the commit itself: the footer names its short sha and its time (Vutuv.BuildInfo), NodeInfo, the Mastodon API and the user agent show the date. test/vutuv/version_test.exs fails the build on a hand-written version: "…" line. The number it replaced was one line every open PR had to change, so every merge set every other PR to CONFLICTING (six rebases and six CI runs on one afternoon, 2026-08-26, none for code), two PRs picking the same number collided with no conflict and no warning, and nothing ever compared it. Three consequences: a PR has no version step and no pre-merge re-check of mix.exs; a branch opened before 2026-08-29 still carries its bump, so when it rebases resolve that one hunk to main's version: version(), — never git checkout --theirs on the whole file, which drops any dep change the branch made; and a shipped change is named by its PR number and merge date (git log -1 --format='%h %cs' <sha>), not by a version.

  • An issue is three sentences of prose and one Where: line — the labelled-field shape is retired. The **TL;DR** / **Where:** / **Now:** / **Want:** skeleton looks organised and reads as six things to hold at once, which is how 37 agent-written issues piled up that Stefan could not tell apart: the median body was already short (148 words, none over 303) and still opaque, because every one of them opened on a module name. Write instead what happens today in words somebody who has never opened the code understands, then what should happen instead, then one sentence saying who notices — a member, an operator of another installation, or nobody outside the code — and what it waits on or what waits on it. Then a single line, Where: <a URL, a module, a function>, naming where to start and nothing more: whoever implements it reads the code anyway, so an inventory of 58 call sites with their frequencies spends the reader's minute doing the implementer's first ten. The title names a surface, never a code construct. It is the only thing the list shows, so "the caches behind the feed's rails each repeat the same timer" beats "share one refresh loop across the three snapshot caches"; an internal change still has a surface Stefan can picture, and naming the construct instead is what makes forty open issues unreadable at a glance. Three sentences land far inside the 120-word ceiling in the global CLAUDE.md, so that ceiling stops being the binding constraint — shape is. docs/ISSUES.md holds the shape with worked before/after examples and is the file to update when this changes; the two .github/ISSUE_TEMPLATE/*.yml forms teach the same shape to the rare human-filed issue (2 of 39 open) and must not drift from it.

  • A PR body is read in under a minute, by somebody who does not hold this codebase in his head. The word ceilings are in Stefan's global CLAUDE.md (PR body 150, comment 80) and they are hard; what this repo adds is what to spend them on. Symptom first, then what you changed, then what he has to decide — the reasoning that got you there belongs in the code comments and the commit message, beside the thing it explains. Name the entry point (URL, module, function) instead of treating internals as common ground. No screenshots. A picture of the changed surface would read better, but gh cannot put an image into a PR body (only the web UI can), and every way around that leans on some other part of GitHub to host what the PR should carry itself: an orphan shots/pr-<nr> branch leaves a branch per PR standing forever, a gist parks repo content outside the repo, and whichever one gets tidied up later turns the merged body into broken placeholders. So name the URL and the module instead, precisely enough that the reviewer can open the real thing, and say in words what changed about it. Still drive the surface in a real browser before you open the PR (the smoke-test rule above) — that part was never about the picture.

  • The language of a GitHub text is set by what you are writing into, never by what you just wrote. Pull requests are English — title and body alike, every time, whatever language the branch's commits are in. Commit messages here are German; an issue body and its comments follow the issue, which is almost always English. The rule is in my global CLAUDE.md and still loses, because a closing note is written moments after a German commit and the momentum picks the language before the rule gets asked — three English issues have been closed in German that way (#1447, #1448, #1502). So re-read the issue (gh issue view <n>) and take its language before the first sentence, every time. One consequence to know rather than trip over: gh pr merge --squash takes the PR title as the subject for a multi-commit PR and the lone commit's subject for a single-commit one, so main's log carries the English title in the first case and the German summary in the second. That is expected, not drift.

  • Never name the production hosts in anything public. bremen1, bremen2, bremen3 and their *.wintermeyer.de addresses stay out of PR titles and bodies, issue bodies and comments, commit messages, and the docs under docs/. Naming the machine that serves the site points at it, and it tells a reader who does not run our infrastructure nothing anyway — vutuv is installable by third parties, so a doc that names our box is describing the wrong thing. Write "the production server", "the GPU host", "the mail host" instead, and where a command genuinely needs the address, leave a placeholder (root@<production-host>). Mentions that predate this rule are grandfathered (docs/production-email-and-bounces.md, docs/architecture/images.md, the Postfix log fixtures in test/); do not sweep them, just stop adding new ones.