Skip to content

refactor(worker): report provider provisioning-failure facts, keep recovery policy in core - #2552

Merged
steipete merged 2 commits into
mainfrom
codex/triage-20260925-provider-failure-facts
Sep 26, 2026
Merged

steipete merged 2 commits into
mainfrom
codex/triage-20260925-provider-failure-facts

Conversation

@steipete

Copy link
Copy Markdown
Contributor

Builds on @steipete's provider-private interpretation and synthetic HTTP tests in #2478. This replacement keeps AWS RunInstances uncertainty and Hetzner structured server/key evidence behind a pure, synchronous provider hook, while the coordinator owns recovery policy. The original draft is unchanged for the coordinator to close with credit.

Removed and why:

  • Removed retryPendingKeyImmediately. AWS reports ownedKeyCleanupPending; core selects immediate cleanup only when the current record has pending owned-key debt and no allocation uncertainty or recorded resource. Other cleanup retains the five-minute retry delay. A regression test supplies identical provider facts with different current allocation evidence and asserts different delays.
  • Removed canceledBeforeAllocation and the adapter's cancellation-context flag. Cancellation plus a missing cloud ID is not native proof of absence. Core retains dispatched-create uncertainty and request markers until existing recovery resolves it; the key-import cancellation test now verifies the inventory absence-confirmation window before cleanup completes.
  • Replaced settlesResourceUncertainty with the factual allocationRejected classification for independently established Hetzner key-only rejection evidence.

The hook remains after the current-lease reread, generation fence, and existing cleanup-custody/unresolved-resource guards. All state writes, retention, custody invalidation, completion, and scheduling remain in core. Foreign-provider errors, existing custody, definitive versus uncertain outcomes, replay conflicts, and real AWS/Hetzner clients against synthetic loopback HTTP endpoints retain coverage. Coordinator documentation and the Unreleased changelog are updated.

Validation:

  • npm ci --prefix worker — passed, zero audit vulnerabilities.
  • npm test --prefix worker -- --maxWorkers=2 — 3,431 passed, 15 skipped, including Node runtime and production-bundle tests.
  • Both Worker and Node typechecks, lint, and formatting checks passed.
  • Scheduling and cancellation regressions failed against the imported draft and passed after the reshape. Four existing cancellation assertions were updated to require unresolved cleanup rather than false completion; the full suite then passed.
  • node scripts/check-docs-links.mjs — passed, 253 Markdown files.
  • Go formatting passed. Local go vet ./..., CLI build, and the command-docs check encountered missing shared Go-cache artifacts following host disk exhaustion. No Go source changed; the coordinator's remote Go gate and CI remain required.
  • Independent Codex autoreview of the final committed branch through P2: scoped-clean, no accepted/actionable findings.

No live native-provider proof was performed: this host's permitted AWS configuration excludes lease creation, and Hetzner credentials are unavailable. Loopback client tests are synthetic transport proof, not native lifecycle qualification. Native proof still requires an approved AWS execution configuration or Hetzner environment.

No wire-schema, dependency, credential, configuration, or migration changes. Canceled creates can now retain unresolved cleanup status until recovery establishes allocation absence. Worker merges follow the existing coordinator deployment workflow; this PR is not merged or deployed. Rollback is a code revert and redeployment, with no data migration.

@clawsweeper

clawsweeper Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. labels Sep 25, 2026
@clawsweeper

clawsweeper Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Codex review: blocked before merge. Reviewed September 26, 2026, 1:19 PM ET / 17:19 UTC (Revision 3).

ClawSweeper review

What this changes

The branch moves AWS and Hetzner provisioning-failure classification into provider hooks, keeps recovery decisions in the coordinator, and adds failure-transport and cancellation tests.

Merge readiness

⛔ Blocked before merge - 5 items remain

Keep open. Current main still handles these provider failures in the coordinator, so this PR remains useful. No discrete patch defect is established, but the longer visible cleanup state and native-cloud rollout need maintainer judgment.

Priority: P2
Reviewed head: 5e3b133e5e4d03c505ebf9ec393b610c4ddf6f3b
Owner decision: Required. See Decision needed.

Review scores

Measure Result What it means
Overall readiness 🦐 gold shrimp (3/6) The patch and production-client transport proof are useful, while existing-state upgrade and native rollout evidence limit merge confidence.
Proof confidence 🐚 platinum hermit (4/6) Sufficient (terminal): The PR reports passing tests that exercise the changed FleetCoordinator path through real AWS and Hetzner clients against injected loopback HTTP failures, observing requests and after-fix lease states, including delayed cancellation cleanup. Native-cloud behavior remains a rollout risk; a pre-upgrade pending AWS record is not directly exercised, so stored-state compatibility evidence is insufficient.
Patch quality 🦐 gold shrimp (3/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Verified Sufficient (terminal): The PR reports passing tests that exercise the changed FleetCoordinator path through real AWS and Hetzner clients against injected loopback HTTP failures, observing requests and after-fix lease states, including delayed cancellation cleanup. Native-cloud behavior remains a rollout risk; a pre-upgrade pending AWS record is not directly exercised, so stored-state compatibility evidence is insufficient.
Evidence reviewed 13 items Applicable repository policy: The full root policy places provider-specific behavior in adapters and generic recovery decisions in core; no applicable nested AGENTS.md or maintainer-notes directory was found.
Introduced recovery boundary: The introduced hook receives failure context after the current lease reread and generation fence; the coordinator writes the lease and schedules recovery.
Changed cancellation state: A canceled dispatched create retains allocation uncertainty and request markers; core delays cleanup when allocation remains uncertain.
Findings None None.
Security None None.

How this fits together

Crabbox’s coordinator receives lease requests, asks cloud providers to create resources, and stores the resulting lease and cleanup state. The shared coordinator runs on Cloudflare Workers or Node.js with PostgreSQL.

flowchart LR
A[Lease request] --> B[Coordinator]
B --> C[Cloud provider]
C --> D[Failure facts]
D --> E[Recovery decision]
E --> F[Stored lease]
E --> G[Cleanup retry]
Loading

Decision needed

Question Recommendation
Should canceled creates remain unresolved until allocation absence is confirmed, and what native AWS or Hetzner qualification is required before deploying that behavior? Qualify before deployment: Demonstrate recovery of a pre-upgrade pending lease, document the pending state, and exercise an approved native cancellation and cleanup case before deploying.

Why: The delayed completion is intentional and safer against uncertain allocation, but accepting its operator-visible timing and the remaining native-cloud gap requires rollout ownership.

Before merge

  • Resolve merge risk (P1) - Canceled creates can remain visibly unresolved through the allocation-absence confirmation window; maintainers have not explicitly accepted that compatibility change for existing clients and operators.
  • Resolve merge risk (P1) - The changed persisted cleanup-state transitions lack a direct recovery check using a pre-upgrade pending AWS cancellation record.
  • Resolve merge risk (P1) - Native AWS and Hetzner cancellation and cleanup remain unqualified; loopback HTTP proves coordinator and client behavior but cannot establish native-cloud lifecycle timing.
  • Complete next step (P2) - Decide the cancellation compatibility and native qualification threshold, and demonstrate recovery of a pre-upgrade pending AWS lease before merge.
  • Resolve maintainer decision - Resolve the maintainer decision shown above before merge.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Production and test delta production +39 net lines, tests +351 net lines The provider-boundary change has substantial transport and cancellation coverage.

Root-cause cluster

Relationship: canonical
Canonical: #2552
Summary: This PR is the stated replacement for the still-open draft covering the same provider-failure refactor.

Members:

Proposal only: this assessment does not dispatch repair, suppress jobs, mutate sibling items, close, or merge anything.

Merge-risk options

Maintainer options:

  1. Prove upgrade and native recovery (recommended)
    Exercise a pre-upgrade pending cancellation record and an approved native failure-and-cleanup case, then document the observed pending state.
  2. Own a monitored rollout
    Accept the longer pending state and native qualification gap explicitly, with a rollout owner watching cleanup debt.

Technical review

Best possible solution:

Keep provider failure facts in adapters and recovery policy in core, then qualify existing-record recovery and agree on the client-visible pending state and native rollout threshold.

Do we have a high-confidence way to reproduce the issue?

Not applicable: this is a refactor with a deliberate safety-behavior change, rather than a reported current-main bug. Focused source tests exercise the changed failure and cancellation paths.

Is this the best way to solve the issue?

Yes, subject to rollout qualification: provider classification and coordinator-owned recovery match the documented boundary. A direct existing-record recovery check would strengthen upgrade confidence.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against a90a781ec5e5.

Labels

Label changes:

No label changes.

Label justifications:

  • P2: This is a bounded coordinator reliability and architecture change with limited direct user impact.
  • merge-risk: 🚨 compatibility: Canceled creates now remain visibly pending longer, and recovery of a pre-upgrade pending record is not directly shown.
  • merge-risk: 🚨 availability: Native cleanup timing remains unqualified, so cloud resources or owned keys could remain pending after cancellation.
  • rating: 🦐 gold shrimp: Overall readiness is 🦐 gold shrimp; proof is 🐚 platinum hermit and patch quality is 🦐 gold shrimp.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Sufficient (terminal): The PR reports passing tests that exercise the changed FleetCoordinator path through real AWS and Hetzner clients against injected loopback HTTP failures, observing requests and after-fix lease states, including delayed cancellation cleanup. Native-cloud behavior remains a rollout risk; a pre-upgrade pending AWS record is not directly exercised, so stored-state compatibility evidence is insufficient.
  • proof: sufficient: Contributor real behavior proof is sufficient. The PR reports passing tests that exercise the changed FleetCoordinator path through real AWS and Hetzner clients against injected loopback HTTP failures, observing requests and after-fix lease states, including delayed cancellation cleanup. Native-cloud behavior remains a rollout risk; a pre-upgrade pending AWS record is not directly exercised, so stored-state compatibility evidence is insufficient.

Evidence

What I checked:

  • Applicable repository policy: The full root policy places provider-specific behavior in adapters and generic recovery decisions in core; no applicable nested AGENTS.md or maintainer-notes directory was found. (AGENTS.md:27, 5e3b133e5e4d)
  • Introduced recovery boundary: The introduced hook receives failure context after the current lease reread and generation fence; the coordinator writes the lease and schedules recovery. (worker/src/fleet.ts:4660, 5e3b133e5e4d)
  • Changed cancellation state: A canceled dispatched create retains allocation uncertainty and request markers; core delays cleanup when allocation remains uncertain. (worker/src/fleet.ts:25934, 5e3b133e5e4d)
  • Recovery timing: An empty provider inventory starts an absence-confirmation window before unresolved allocation cleanup completes. (worker/src/fleet.ts:17085, 5e3b133e5e4d)
  • Current main remains distinct: Pinned current main still contains coordinator-side AWS and Hetzner failure classification and the earlier AWS cancellation-completion branch. (worker/src/fleet.ts:25930, a90a781ec5e5)
  • Release check: The latest release is v0.66.0; the PR is unmerged and its new hook is absent from the pinned current-main source.

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • vincentkoc: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rank-up moves

Optional improvements that raise the rating; they are not merge blockers.

  • Show a pre-upgrade pending AWS cancellation record recovering under the new coordinator behavior.
  • Provide approved native AWS or Hetzner cancellation evidence, or obtain explicit acceptance of a monitored rollout.

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (2 earlier review cycles)
  • reviewed 2026-09-25T05:23:23.858Z sha 90bcee2 :: blocked before merge. :: none
  • reviewed 2026-09-26T16:07:45.618Z sha dd5f2f3 :: blocked before merge. :: none

@steipete
steipete force-pushed the codex/triage-20260925-provider-failure-facts branch from 90bcee2 to dd5f2f3 Compare September 26, 2026 16:03
@clawsweeper clawsweeper Bot added the proof: sufficient Contributor real behavior proof is sufficient. label Sep 26, 2026
…covery policy in core

Build on the provider-private interpretation and synthetic HTTP coverage from #2478 by @steipete. Remove adapter scheduling and cancellation-derived absence outputs; core retains unresolved allocations and owns retry timing.
@steipete
steipete force-pushed the codex/triage-20260925-provider-failure-facts branch from dd5f2f3 to 5e3b133 Compare September 26, 2026 17:14
@steipete
steipete merged commit 7073c2e into main Sep 26, 2026
48 checks passed
@steipete
steipete deleted the codex/triage-20260925-provider-failure-facts branch September 26, 2026 17:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. proof: sufficient Contributor real behavior proof is sufficient. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant