I'm working on a Kubernetes operator for celld (https://github.com/ewhauser/celld-operator) and need a way to know when a node's disk can safely be removed during scale-down or replacement.
The existing shutdown endpoint acknowledges the request, but neither that response nor process exit tells us whether the disk is still needed for recovery. In particular, the node may hold follower data for other nodes. I'd like celld to own that decision rather than have the operator inspect and interpret its recovery state.
I implemented a workaround in my fork that extends the shutdown endpoint with a disk-removal mode and exposes progress through the state endpoint. It stops accepting new work, drains outstanding work, and checks that both the node's own data and its follower obligations are recoverable without the disk. It then stays alive in a control-only state so the supervisor can observe completion before stopping the process and removing the disk. Requests are tied to a specific runtime incarnation and can be retried safely. If completion can't be confirmed, we retain the disk.
Would you be open to a patch for this? I'm happy to adapt the approach to how you'd like it to fit into celld's existing shutdown mechanism and control plane.
I'm working on a Kubernetes operator for celld (https://github.com/ewhauser/celld-operator) and need a way to know when a node's disk can safely be removed during scale-down or replacement.
The existing shutdown endpoint acknowledges the request, but neither that response nor process exit tells us whether the disk is still needed for recovery. In particular, the node may hold follower data for other nodes. I'd like celld to own that decision rather than have the operator inspect and interpret its recovery state.
I implemented a workaround in my fork that extends the shutdown endpoint with a disk-removal mode and exposes progress through the state endpoint. It stops accepting new work, drains outstanding work, and checks that both the node's own data and its follower obligations are recoverable without the disk. It then stays alive in a control-only state so the supervisor can observe completion before stopping the process and removing the disk. Requests are tied to a specific runtime incarnation and can be retried safely. If completion can't be confirmed, we retain the disk.
Would you be open to a patch for this? I'm happy to adapt the approach to how you'd like it to fit into celld's existing shutdown mechanism and control plane.