fix: doctor answers whether a drain would work - #55
Merged
Merged
Conversation
"Would a drain refuse here?" is the question an operator has after every upgrade, and until now the only way to answer it was to attempt the drain — on a busy pool, with real jobs at stake. Checking the thing by doing the thing is exactly wrong when the thing is destructive if it misjudges. doctor now reports, per pool, which loaded runners predate the graceful-shutdown setting. A note rather than a failure: such a pool runs and takes work perfectly normally, and only --drain is affected, which refuses safely on its own. Nothing is wrong; this exists so nobody has to find that out the hard way. Found immediately after the 0.10.0 release, when neither machine could be cycled and there was no read-only way to ask what state its runners were actually in.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
"Would a drain refuse here?" is the question an operator has after every upgrade, and until now the only way to answer it was to attempt the drain — on a busy pool, with real jobs at stake. Checking the thing by doing the thing is exactly wrong when the thing is destructive if it misjudges.
doctornow reports, per pool, which loaded runners predate the graceful-shutdown setting:A note, not a failure. Such a pool is running and taking work perfectly normally. Nothing is wrong until someone tries to drain it, and the drain refuses safely by itself. This exists so nobody has to find that out the hard way.
It reads the loaded agent's environment, not the plist on disk — the same distinction
_rp_drain_poolrelies on, and the one that mattered in practice at the 0.10.0 upgrade when both machines had correct plists on disk while every loaded runner still had the old behaviour.Found immediately after releasing 0.10.0, when neither machine could be cycled and there was no read-only way to ask what state its runners were in.
Verified live against a real pool. Five new cases in
tests/drain-guards.sh(28 total);bash -nandshellcheck --severity=warningclean, all four suites pass.