Skip to content

Safe node retirement through the control plane #223

Description

@ewhauser

I'm working on a Kubernetes operator for celld (https://github.com/ewhauser/celld-operator) and need a way to know when a node's disk can safely be removed during scale-down or replacement.

The existing shutdown endpoint acknowledges the request, but neither that response nor process exit tells us whether the disk is still needed for recovery. In particular, the node may hold follower data for other nodes. I'd like celld to own that decision rather than have the operator inspect and interpret its recovery state.

I implemented a workaround in my fork that extends the shutdown endpoint with a disk-removal mode and exposes progress through the state endpoint. It stops accepting new work, drains outstanding work, and checks that both the node's own data and its follower obligations are recoverable without the disk. It then stays alive in a control-only state so the supervisor can observe completion before stopping the process and removing the disk. Requests are tied to a specific runtime incarnation and can be retried safely. If completion can't be confirmed, we retain the disk.

Would you be open to a patch for this? I'm happy to adapt the approach to how you'd like it to fit into celld's existing shutdown mechanism and control plane.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions