fix(tencentcloud): preserve acquisition causes and bound readiness - #2464
Conversation
|
🦞👀 Pull request received. I will update this pull request when review starts. ClawSweeper review completeClawSweeper finished reviewing this revision. The review result is being finalized. |
|
Codex review: blocked before merge. Reviewed September 25, 2026, 1:12 PM ET / 17:12 UTC (Revision 4). ClawSweeper reviewWhat this changesThe Tencent Cloud provider preserves acquisition and rollback error causes, stops a fresh-instance retry after reported termination failure, and bounds public-IP readiness, with tests and operator documentation. Merge readiness⛔ Blocked before merge - 4 items remain Keep open: current main and v0.66.0 still have the older Tencent Cloud behavior, and this PR provides a distinct fix. The patch has no identified correctness finding; merge readiness depends on an explicit decision about the retry policy and the pending native qualification. Priority: P2 Review scores
Verification
How this fits togetherCrabbox's Tencent Cloud provider turns a CLI lease request into a cloud instance, waits for its public IP and SSH access, then returns a claimed lease. If acquisition fails, it attempts termination and decides whether to retry. flowchart LR
A[CLI lease request] --> B[Tencent Cloud provider]
B --> C[Create cloud instance]
C --> D[Poll public IP]
D --> E[Check SSH and claim]
E -->|Ready| F[Usable lease]
E -->|Failure| G[Terminate and assess retry]
G -->|Cleanup succeeded| B
G -->|Cleanup failed| H[Error for operator]
Decision needed
Why: The retry veto deliberately changes a conditional recovery path, while the current proof exercises a loopback transport rather than a Tencent account; accepting that upgrade behavior and qualification threshold requires provider-owner intent. Before merge
Agent review detailsSecurityNone. Review metrics
Merge-risk optionsMaintainer options:
Technical reviewBest possible solution: Ship the bounded wait and cause-preserving rollback with clear operator guidance for a termination whose outcome is unconfirmed. Do we have a high-confidence way to reproduce the issue? Yes. Current-main source exposes the lost primary cause and unbounded readiness read, and the captured baseline tests and source-matched loopback CLI trace give a focused reproduction path; I did not execute the path in this read-only review. Is this the best way to solve the issue? Yes for reusing the existing shared helpers to preserve causes and bound polling. The conditional no-retry policy still requires explicit owner acceptance. AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against 5fa10bdaee68. LabelsLabel changes: No label changes. Label justifications:
EvidenceWhat I checked:
Likely related people:
Rank-up movesOptional improvements that raise the rating; they are not merge blockers.
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
History |
08aef04 to
aea8543
Compare
aea8543 to
a8211c4
Compare
Brings in 5fa10bd (telemetry workspace command cleanup, openclaw#2563) and e636e39 (Tencent Cloud acquisition causes and readiness, openclaw#2464). The merged tree passed vet, race tests of the Proxmox, Tencent Cloud and all-provider packages and cmd, the internal/cli race suite (except the nc-dependent test that fails identically upstream on devbox) and the docs checks before this merge.
Summary
Consolidate two Tencent acquisition policies behind existing shared helpers, keeping the changes in separate commits:
JoinAcquireCleanupErrorpreserves primary and termination error causes and primary typed exit-code precedence. Reported cleanup failure explicitly vetoes a fresh-instance retry; successful rollback retains the existing bootstrap retry.PollReadinessowns the five-minute elapsed IP-readiness budget and three-second wait. Cooperative reads now receive the bounded observation context, waits cannot overshoot the budget, and caller cancellation retains its custom cause without changing the canonical public cancellation diagnostic.Tencent still owns public-IPv4-only readiness (regardless of state), immediate API-error handling, completed-response precedence, and the exact timeout message/exit code 5. A completed ready response or typed API error can win at coincident cancellation; this is not a promise to reject every response from a client that ignores context. Lifecycle timestamp clocks remain separate from the readiness budget.
Termination order, its independent 45-second cleanup context, tags, claims, and local-key removal are unchanged. This does not qualify recovery-state retention. Docs and the maintainer-added Unreleased entry describe both fixes.
Verification
testing/synctestexercises the actual five-minute/three-second constants without production timing knobs. The successful-rollback control still retries; the new combined cancellation-plus-cleanup test preserves the private cause without exposing its text or allocating twice.Field-selected before/after CLI results:
[ { "phase": "before", "source": "e6c63a5970e43402fea33768d465da6f52fddb64", "sha256": "8a7be9fa074e63771d3b2b446df4af4bcb432fd4c84d096ef1394740f25db9e2", "exit": 1, "actions": [ "DescribeInstances", "RunInstances", "DescribeInstances", "TerminateInstances" ], "primaryDiagnostic": true, "cleanupDiagnostic": true, "retryVetoWarning": false }, { "phase": "after", "source": "08aef04d28c2d7f2bacd6878b23469ddac59b0b5", "sha256": "e5dda9de653928f7b7d29275b6a4a64c6791279f52ed2c91947dc66bd6352d16", "exit": 1, "actions": [ "DescribeInstances", "RunInstances", "DescribeInstances", "TerminateInstances" ], "primaryDiagnostic": true, "cleanupDiagnostic": true, "retryVetoWarning": true } ]Qualification status
This remains a draft pending native Tencent lifecycle proof. Known approved credential locations were unavailable in the previous check; no credentials or cloud account permissions were changed. The prior proof remains attributed to its original source; the results above are from the integrated head. No workflow, dependencies, release machinery, or provider mutation authority changes.