Skip to content

Node self-fences after S3 keeps rejecting lease renewals whose If-Match matches the current ETag (no other writer) #239

Description

@ridhwanhassan

Summary

On AWS S3, a node's lease renewal is refused once, and then every retry is refused too, even though each retry uses the ETag the node just read back. No new object version lands on the key. The TTL runs out and the node self-fences (node_lease_watchdog_fence code=3, HaltReason::NodeLeaseExpired). ECS replaces the node in about 60 s, and our app saw no failed requests, but it happened 3 times in the first 7.5 h after we moved to v0.6.0 (about 10 a day) on a 3-node fleet.

The log line does not say whether S3 answered 412 or 409, because put_cas folds a clean 412/409 into Ok(None) → rejected (per the comment in ownership_store.rs). That status is what we need to find the cause.

Environment

  • celld v0.6.0 (it also happened on v0.4.1, at about 2 a day)
  • AWS S3, ap-southeast-5, versioned bucket, no lifecycle rules
  • CELLD_DURABILITY=fleet, 3 nodes on ECS Fargate
  • The lease key is written every ~3.3 s

What the logs show (one fence, 2026-09-29, UTC)

Renewals apply normally every 3.33 s:

01:22:51.235  node lease attempt started   attempt="renew" prior_authority_headroom_ms=6665
01:22:51.261  node_lease_write outcome="applied" duration_us=25631

The next renewal is rejected, and so is every retry after a readback:

01:22:54.571  node lease attempt started   attempt="renew" prior_authority_headroom_ms=6664
01:22:54.629  node_lease_write outcome="rejected" duration_us=58384
01:22:54.656  node_lease_read  outcome="found" generation="3df88238…"
01:22:56.320  node_lease_write outcome="rejected" duration_us=16767
01:22:56.341  node_lease_read  outcome="found" generation="3df88238…"
01:22:57.583  node_lease_write outcome="rejected" duration_us=16876
   … 10 more rejected writes (13 in total), each ~17 ms, each followed by a readback that returns the same generation …
01:23:01.237  SELF-FENCE: node lease not renewed within TTL — halting event="node_lease_watchdog_fence" code=3

The object did not change

S3 version history of <prefix>/nodes/node_abd4….json around the fence (list-object-versions):

01:22:42  (applied renewal)
01:22:45  (applied renewal)
01:22:48  (applied renewal)
01:22:52  (applied renewal, the last one that succeeded)
— no version between 01:22:52 and 01:23:05 —

So no write landed on the key between the last applied renewal and the fence. (S3 data events were not enabled, so we cannot list rejected requests from other callers.)

Why a stale ETag does not explain it

crates/logic/lib.rs handles a Rejected renewal by reading the lease back. In read_self_node_lease, when the readback equals the prior record (same_node_lease), it adopts the fresh ETag (prior.record.etag = record.etag) and retries. The node kept retrying and finally halted with NodeLeaseExpired, not NodeLeaseMismatch or NodeLeaseMissing, which fits that branch. So, as far as we can tell from the code, each retry sent If-Match with the ETag S3 had just returned from a read, and S3 still refused it every time, for 6.6 s.

put_cas uses the cas_store client with max_retries: 0, so these are single requests, not SDK retries landing on their own earlier write.

Questions / suggestions

  1. Log the HTTP status and the S3 error code on a rejected CAS (412 PreconditionFailed vs 409 ConditionalRequestConflict), and the S3 request ID. Without it we cannot tell a real lost race from an S3-side conflict.
  2. The S3 docs say an If-Match write "can also receive a 409 Conflict response in the case of concurrent requests", and that a PutObject "may be retried after receiving a 409 Conflict error" (conditional writes). Should a 409 be kept apart from a 412 and retried with the same guard, without spending the readback and the renewal slot?
  3. Is there any known interaction between S3 conditional writes and a key with many noncurrent versions? Our lease keys gain about 26,000 versions a day because the bucket is versioned. (This may be unrelated: the first prod fence on the new prefix came 3.5 h after the move, when the key had about 3,800 versions.)

We can turn on CloudTrail S3 data events for the lease prefix and attach the exact status code from the next fence if that helps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions