Repository navigation
feat: run the macOS e2e suite on CodeBuild - #1093
Open
andychoquette wants to merge 27 commits into
Open
andychoquette wants to merge 27 commits into
andychoquette wants to merge 27 commits into
Conversation
Platform gating, in both directions: - The two Linux job-bundle tests were gated to skip only on Windows, so they ran against a macOS fleet and Deadline auto-cancelled them as NOT_COMPATIBLE. The other tests labelled "Linux specific" are left running, since nothing shows they are actually incompatible and skipping them would hide coverage. - test_worker_shuts_down_host_machine_if_configured is skipped on macOS: it reads instance_id, which LocalMacWorker does not have, and expects the host to power off, which on a runner that hosts the worker ends the suite. - test_session_root_dir shares the POSIX branch, under /private/tmp because the macOS system volume is sealed and /tmp is a symlink whose resolved form the agent reports. Portability and correctness: - The session_runtime switch used a bare `sed -i`, which is GNU-only; BSD sed takes the script as the backup suffix. The suffixed form works under both. - The Metal probe piped system_profiler into grep under `set -e`, so a headless CI VM with no displays exited 1 and aborted the script before the Metal check it exists to perform. The inventory is informational and no longer fatal, and swift's stderr is redirected so a real failure is visible rather than looking like nothing ran. - The runtime recovery test's stop confirmation is removed. Each branch named its own service manager, so on macOS the POSIX branch ran `systemctl is-active`, exited 127 and satisfied both assertions immediately -- no settle time on the one platform where the restart raced. systemctl stop and Stop-Service are synchronous, so Linux and Windows lose nothing, and confirming the stop now sits behind stop_worker_service. - The path mapping test asserts the job succeeded before checking outputs, since wait_until_complete treats FAILED as complete and the failure otherwise surfaced as an empty mapping with no indication of the cause. - Worker annotations in the session runtime tests accept either concrete worker that can drive the agent's service, rather than only the EC2 one. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Replaces the capability probe in pipeline/e2e-macos.sh with the suite. The probe answered its question on the reserved MAC_ARM fleet: passwordless sudo, a hidden sub-500 account via dscl, dseditgroup GID allocation, an /etc/sudoers.d rule and a LaunchDaemon bootstrapped and running as that account are all permitted there. Reworked from aws-deadline#1080 rather than merged from it. That PR targeted a GitHub-hosted runner with no provisioned Deadline resources, so it carried a 366-line _scaffold_deadline_resources that created a farm, queues and fleets per run. On CodeBuild those resources already exist and arrive as the SSM-backed environment variables the Linux and Windows projects use, so that whole path is dropped along with aws-deadline#1080's own workflow and its six diagnostic commits. The bootstrap mirrors e2e.sh's TEST_TYPE block rather than shortening to `hatch run e2e:test`. The reusable workflow sets TEST_TYPE=WHEEL, and with WORKER_AGENT_WHL_PATH unset conftest installs the PUBLISHED agent instead of this commit -- a green run that exercised nothing. The built wheel's existence is asserted, because that fallback is silent. conftest's create_worker gains the macOS branch, ahead of the EC2 one. This is the half that is easy to miss: yielding LocalMacWorker is not enough, because the EC2 path asserts on SUBNET_ID and SECURITY_GROUP_ID and pulls a worker instance profile out of bootstrap_resources, none of which exist for a worker that installs the agent onto the host running the tests. It also drops the ec2_worker_type parameter, which every call site passed the same fixture value for. The probe's cleanup, stale gate and leftovers survey are replaced by a narrow pre-build reset. LocalMacWorker's _reset_host_state clears only worker.toml and worker.json within a run, so on reserved capacity a daemon left loaded by a killed build and the worker id it registered would poison the next one. The account, group and /opt/deadline venv are deliberately left: the installer is idempotent over them and reusing them saves minutes. The bootout is polled rather than read once, since it is asynchronous. Requires deadline-cloud-test-fixtures 0.18.23, which wires MACOS to LocalMacWorker (aws-deadline/deadline-cloud-test-fixtures#333). Until that releases, the e2e requirement does not resolve and this cannot go green. Verified locally as far as is possible without installing an agent on a developer machine: 79 tests collect under OPERATING_SYSTEM=macos, e2e mypy and ruff are clean, the macOS branch constructs with only configuration and deadline_client while AL2023 still requires SUBNET_ID, the CODEBUILD_BUILD_ID guard exits 1, and the wheel block builds a real wheel and fails when the path does not resolve. Not verified: a suite run, which needs the fixtures release and a dispatch from mainline. The workflow stays dispatch-only. Adding the push trigger before one green run would put a red macOS check on mainline. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…hat need two Review found that LocalMacWorker installs onto the host running the tests, so two workers cannot coexist: the second start() overwrites the first's worker.toml, worker.json and LaunchDaemon, and whichever configuration landed last is the one under test. Every worker fixture in test_session_runtime.py creates its own worker, and test_worker_config.py creates one per test, so on macOS all of them collide with the session-scoped worker that 44 call sites depend on. Skipping them was the obvious fix and it costs too much. macOS already runs 45 of the 79 collected tests against 56 on Linux and 62 on Windows, mostly for legitimate reasons -- cap_kill is Linux capabilities, domain_user is Windows AD, and both platforms' job-user and job-attachment variants are OS-specific. Skipping the two-worker tests would have taken macOS to roughly 37 and removed session-runtime routing entirely, which for a platform whose agent support is new is the more interesting half. So macOS installs them one at a time. Creating a worker stops whatever holds the host first, and the session worker is reinstalled on next use if it was the one displaced -- lazily, so a run that never returns to it pays nothing. stop() leaves the venv and the account, so reinstating is the installer plus a launchd bootstrap rather than a full provision. Counts are unchanged at 56/62/45. session_worker splits into a cached session-scoped impl and a function-scoped wrapper. Without a per-use hook the displaced worker would stay stale: pytest hands back the same object all session, so every later assertion would run against a host with no agent. The wrapper is a no-op on Linux and Windows. stop_worker gained a guard that turned out to be load-bearing. A displaced worker has already been stopped, and stopping it again raises: stop() deletes the worker record and the second DeleteWorker on the same id fails, so both the session worker's teardown and any displaced per-test worker's would have thrown at the end of a run. Only the worker still holding the host has anything to tear down. Nothing changes for Linux or Windows. Both new call sites are behind `if _MACOS`, where workers are separate instances and genuinely concurrent. Six tests in test/e2e/test_worker_exclusivity.py cover the state machine directly, including the double-stop that would have raised and a claim whose incumbent fails to stop. They skip off macOS, where the registry is inert. Not verified end to end: that needs the fixtures release and a dispatch from mainline, and the reinstall cost per transition is unmeasured against the 180-minute timeout. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…w gaps The rust fault injection moves openjd/model/_v1 aside instead of deleting it, and moves it back in rust_unavailable_worker's teardown. On EC2 either would do, because the instance is discarded with the class. LocalMacWorker installs into a venv at /opt/deadline/worker on a reserved-capacity host, and e2e-macos.sh keeps that venv deliberately, so a delete is permanent: build N would leave build N+1 failing TestExplicitModeRouting[rust], which runs before the class that broke it, reporting a missing adapter as a routing regression. Moved into the fixture rather than left in the test body, because that is the scope that owns the host mutation, and the path is resolved once before the move: openjd.model imports from _v1, so the import used to locate it stops working as soon as it is gone. e2e-macos.sh repairs a tree left aside by a build that died between the move and the restore, which the fixture teardown cannot cover. It finds the directory rather than asking python where the package is, for the same reason. ec2_worker_type is resolved inside the EC2 branch. It raises ValueError for an operating system it does not recognise, so resolving it up front would fail a macOS run at fixture setup before the branch that wants LocalMacWorker could run. That branch now names LocalMacWorker directly: ec2_worker_type is the override point for swapping in an EC2 subclass, so a cast there asserted a relationship the fixture does not guarantee and would have failed on the missing ec2_client and subnet_id keywords. _grab_bootstrap_log handles LocalMacWorker. macOS bootstrap -- account creation, the sudoers rule, the LaunchDaemon -- is the part most likely to fail, and the reusable workflow hides the CodeBuild log from GitHub, so without this the only record of a start failure was the exception string. One command per file, because LocalMacWorker runs under set -euo pipefail and a missing path would abort the rest. The session root moves from /private/tmp/mysessionroot to /opt/mysessionroot. /private/tmp is mode 1777, so an unprivileged process could pre-create the directory and have the agent adopt one it does not own. /Users/Shared is 1777 as well; /opt is root:wheel 0755, on the writable data volume, and already holds the agent's venv. The stop confirmation before the service restart is restored with a macOS branch. It was dropped because the POSIX branch names systemctl, which macOS has not got: there `systemctl is-active` exits 127 with empty stdout and satisfies both assertions immediately, contributing no settle time on the one platform where the restart actually races. macOS asks launchd instead. Collection is unchanged at 85 tests; runs/skips are 56/29 linux, 62/23 windows, 51/34 macOS. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…e registry on all platforms _claim_host now stops the incumbent through stop_worker instead of calling worker.stop() directly. A displaced worker was running jobs moments ago, which is exactly when DeleteWorker returns ConflictException, and stop_worker is where the retry for that lives. The bare call swallowed a transient conflict permanently: _installed_worker moved to the newcomer regardless, so the incumbent's own teardown took the already-displaced early return and nothing retried the delete. reinstate_session_worker goes through _claim_host too, so it releases the host before installing. A class-scoped per-test worker can still be alive when a session_worker test reinstates; without this the session worker overwrote its worker.toml in place and its teardown found nothing to stop, leaking the record. _displaced_session_worker is now cleared after a successful start rather than before it, so a failed reinstall stays displaced and the next use retries instead of being handed a host with no agent. stop_worker clears _installed_worker below the KEEP_WORKER_AFTER_FAILURE return, not above it. A kept worker is installed and running, so releasing the slot let the next _claim_host install straight over the state the flag exists to preserve -- destroying the post-mortem it promises while starting the new worker against a LaunchDaemon and worker id it did not create. _claim_host now refuses that case with a message naming the conflict, because only one agent fits on the host and neither silent option is defensible. The exclusivity tests use monkeypatch instead of assigning conftest's globals, and patch _MACOS on rather than skipping off macOS. Assigning directly left a MagicMock in _installed_worker for the rest of the session, which would have stopped the mock in place of a real incumbent and sent real teardowns down the already-displaced path. Patching _MACOS means the state machine is now covered by the Linux and Windows runs that gate merges, not only by the platform it changes: 11 tests, up from 6, passing on all three. Four fixtures still declared ec2_worker_type without using it, so pytest resolved it during setup and the unsupported-OS ValueError fired before create_worker could branch. That hit test_session_root_dir, which is one of the tests this PR enables on macOS, so the change would have gone untested behind a fixtures-compatibility error. Runs/skips: 67/23 linux, 73/17 windows, 56/34 macOS. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Dropped when the module was rewritten to use monkeypatch, which failed test_copyright_headers on every platform. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Mirrors the mutual-exclusion the fixtures package's own worker fixture applies. Without it the Docker branch wins: the run reports macOS results from a Linux container, and the macOS exclusivity registry claims and stops Docker workers it was never written for. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
The third ordering bug in stop_worker, same shape as the KEEP_WORKER_AFTER_FAILURE one: _installed_worker was cleared before _stop_with_retry(), so a stop that failed left a live agent on a host the registry recorded as free, and the next _claim_host installed straight over it -- the silent overwrite the registry exists to prevent. Reachable on the first attempt, because the backoff gives up on anything that is not a ConflictException: a throttle, an expired credential, or a bootout failure all land there immediately. The slot now clears in the try's else, so the registry says the host is free exactly when it is. A failed stop keeps the worker recorded; the next claim sees it and retries the stop -- covered by a test that fails the first stop, asserts the worker stays recorded, then claims with a second worker and asserts the incumbent was stopped again rather than overwritten. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
The header still described the capability probe this PR replaces: probe-specific exit codes 1/2/3 that no longer exist, and a rationale saying nothing had confirmed a CodeBuild macOS build can create users, write sudoers rules and bootstrap a LaunchDaemon. A probe build did confirm all of that. The workflow stays dispatch-only, but for the remaining reason rather than the settled one: the suite has never completed a run on this fleet. The trigger note now keys on a dispatched run passing instead of on the installer being proven. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
The launchctl stop gate was still vacuous. Keying on launchd's not-found message OR the absence of a running state is satisfied by every error launchctl can emit, because an error message has no state line either -- so any failure passed the gate on the first attempt with no settle time, which is what the previous revision claimed to fix. There are now three outcomes, not two: not-found is stopped, a loaded job is stopped only if not running, and anything else is evidence of neither and raises so the backoff keeps trying. Both call sites share one helper, since both wanted the same predicate and both got it wrong the same way. Verified the substring against a real launchctl: a missing label prints 'Could not find service "..." in domain for system', so the match itself was right. A bad domain prints 'Unrecognized target specifier', which is the case the old disjunction waved through. KEEP_WORKER_AFTER_FAILURE no longer cascades. The refusal keyed on 'a test failed and the flag is set', and testsfailed never decreases, so one failure anywhere converted every later macOS worker-creating test into a setup error -- a wall of noise on exactly the run the flag exists to diagnose. It now keys on _kept_worker, set only where stop_worker actually leaves a worker installed. Displacement also passes honor_keep_after_failure=False, because the first attempt at this moved the cascade rather than removing it: stop_worker declined to stop during displacement, and _claim_host then recorded the newcomer and installed it over a live agent. Two mutations moved inside their try blocks, the same defect already fixed on the rust fixture. _set_plist_env is a read-merge-write, so a failure after the plist landed left the override in place with the finally unreached -- and bootout does not rewrite the plist, so it survives into the next build and every job there runs as the override user. test_config_file_user_override had the same gap around its sed-then-grep, which poisons the rest of the current run. BUILD_RESIDUE's comment claimed it covered logs, which it does not. It now states why: _grab_bootstrap_log reads from /var/log/amazon/deadline after a start failure, so wiping it first would destroy the only diagnostic for the failure being hit. The plist and sudoers rule are absent for a different reason, that every start() rewrites both wholesale. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Two facts the suite's correctness depends on and nothing has established. test_worker_requires_no_instance_profile is skipped on macOS because a Mac has no IMDS, so _enforce_no_instance_profile raises IMDSUnreachableError rather than the agent refusing work under an instance profile. But CodeBuild reserved capacity is EC2 Mac underneath, so IMDS may answer here and that skip may be hiding a test that would run. The output distinguishes the three cases: a token plus a 200 on iam/info means the skip is wrong, a 404 means IMDS answers with no profile so the test would fail rather than be inapplicable, and no token confirms the skip's stated reason. /usr/bin/python3 --version is logged because it is the interpreter LocalMacWorker builds the agent venv from, and the xfail on TestServiceSelectedFollowsServiceHint is justified by it being 3.9 -- which caps botocore below the release that models AssignedSession.metadata. If a future macOS ships something newer the xfail should be removed, not left to pass as an xpass. Diagnostic only: both print and neither can fail the build. --max-time 3 keeps a silently dropped IMDS request from costing two minutes, and 169.254.169.254 is link-local so nothing leaves the host. Verified off-EC2 that the whole block takes 3 seconds and reports unreachable rather than hanging. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
crowecawcaw
previously approved these changes
Oct 1, 2026
andychoquette
had a problem deploying
to
mainline
October 2, 2026 16:23 — with
GitHub Actions
Failure
andychoquette
had a problem deploying
to
mainline
October 2, 2026 16:59 — with
GitHub Actions
Failure
…test on macOS The first dispatch failed after 16 seconds on PEP 668: the host's pip3 is Homebrew's and refuses system-wide installs, so `pip3 install --upgrade hatch` exits with externally-managed-environment. The Linux script gets away with the same lines because it runs in a container. --break-system-packages would work and is the wrong fix: this is reserved capacity, so mutating the host's Homebrew python is a change every later build inherits, which is the class of problem the rest of this script exists to avoid. The tooling goes in a venv under the build directory, which CodeBuild already cleans between builds, and PATH is prepended so hatch resolves to it. test_worker_requires_no_instance_profile runs on macOS again. I had skipped it on the assumption that a Mac host has no IMDS, so _enforce_no_instance_profile would raise IMDSUnreachableError rather than the agent finding a profile and refusing work. The diagnostic added for exactly this question disproved it on the real host: CodeBuild reserved capacity is EC2 Mac underneath, IMDSv2 answers, and iam/info returns 200, so a profile is attached and the test's premise holds the same way it does on the Linux and Windows instances. The comment records that, so the next person does not re-skip it on the same reasoning. The other diagnostic confirmed /usr/bin/python3 is 3.9.6, which is what the TestServiceSelectedFollowsServiceHint xfail depends on. Runs are now 66 linux / 72 windows / 63 macOS of 95 collected. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
andychoquette
had a problem deploying
to
mainline
October 2, 2026 17:07 — with
GitHub Actions
Failure
…odeArtifact The second dispatch got as far as LocalMacWorker.start() and failed in _install_agent with AccessDenied on codeartifact:GetAuthorizationToken, as arn:aws:sts::539109460720:assumed-role/customer-.../i-0c9b3233d0ad98a4b. That identity is an instance profile, not the CodeBuild build role. send_command runs every fixture command through sudo, whose env_reset strips AWS_* from the environment, so the CLI falls back to IMDS -- which this host has, being EC2 Mac underneath -- and picks up the CodeBuild host's own instance profile: an identity in an AWS-owned account with no access to our CodeArtifact domain. install_command_for_linux chains the login with && before pip install, so the denial aborts the install. On EC2 Linux the same lookup succeeds because the instance profile is one the fixtures bootstrapped and it carries the access. The host's profile here is not ours to configure, so the credentials go where a root shell will find them. aws configure export-credentials rather than copying AWS_* directly: on CodeBuild the build role usually arrives through the container credential endpoint rather than as static keys, so there may be nothing in the environment to copy. The transformation to a credentials file was verified locally end to end. Written with mode 600, piped through stdin so no secret reaches a command line or this log, removed by an EXIT trap, and listed in BUILD_RESIDUE so the next build clears it too -- live credentials in root's home must not outlive the build that needed them on reserved capacity. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
andychoquette
had a problem deploying
to
mainline
October 2, 2026 17:14 — with
GitHub Actions
Failure
…d them The previous attempt wrote /var/root/.aws/credentials and root still resolved to the host's instance profile. macOS sudoers carries env_keep+="HOME MAIL", so sudo preserves the invoking user's HOME and the AWS CLI running as root looks in the build user's home, not /var/root. Root gets /var/root first, and the build user's home only as a fallback when root still cannot see the build role. Writing [default] into the build user's home shadows its own credentials, which on CodeBuild come from the container endpoint and refresh, where a static snapshot does not -- and the suite is budgeted at up to three hours. Paying that only when the first location fails keeps the refresh intact on any host whose sudoers resets HOME. Both locations are removed by the EXIT trap and listed in BUILD_RESIDUE. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
andychoquette
had a problem deploying
to
mainline
October 2, 2026 17:22 — with
GitHub Actions
Failure
First real run: 33 passed, 3 failed, 11 errored. Both causes are fixed here. The 11 errors were 'Worker agent install failed: exit_code: 254' -- the CodeArtifact AccessDenied again, but only from 17:41 after succeeding from 17:24. Copying the credentials to a file was the wrong shape: CodeBuild's build credentials arrive through the container endpoint and rotate, so a file is a snapshot that expires about 25 minutes in. A sudoers drop-in adding the AWS credential variables to env_keep means root resolves credentials exactly as the build user does, refresh included, and nothing secret is written to a disk that outlives the build. Validated with visudo -cf before installing, because a malformed file in /etc/sudoers.d breaks sudo for every user on reserved capacity. The 3 failures were the tests this change enabled, all from one GNU-ism in the shared bundles: BSD wc -l right-pads its count to 8 columns, so [ "`ls ... | wc -l`" != "2" ] compares " 2" against "2" and fires on every check. 24 of those across the two bundles are now numeric -ne, which ignores the padding and behaves identically on GNU. Three other wc -l uses were already safe, being unquoted in [ $COUNT -lt 3 ] where word-splitting strips it. That is the third GNU-ism these bundles have yielded after /bin/env and the os.family pin, all found only by running on the host. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
andychoquette
had a problem deploying
to
mainline
October 2, 2026 17:48 — with
GitHub Actions
Failure
The run showed why it cannot work on macOS, and it is not the script. complex_bundle/linux hardcodes /tmp/storageprofiletest in its attachment manifests, while MacOSJobStorageProfile declares the resolved /private/tmp/storageprofiletest -- correctly, since /tmp is a symlink there and the agent reports resolved paths. Job attachment sync then fails with 'No path mapping rule found for the source path /tmp/storageprofiletest' before any task runs: step 1 FAILED, every dependent step CANCELED, the task NEVER_ATTEMPTED. Enabling it on macOS needs a complex_bundle/macos variant carrying the resolved paths, the same way the Windows suite has its own. That is an unwritten variant rather than something this change can fix, so the test keeps its Linux-only skip and the reason now says which difference blocks it. The two dependency-data-flow tests stay enabled on macOS: they pass now that the wc -l comparisons are numeric, and their bundle declares no storage profile locations. macOS runs 62 of 95. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
andychoquette
had a problem deploying
to
mainline
October 2, 2026 18:45 — with
GitHub Actions
Failure
…estart test The run after the previous fixes came back 57 passed, 1 failed, and the failure was not one of the new tests: test_worker_lifecycle_status_is_expected, which predates this change and runs on all three platforms. is_worker_stopped backed off to exhaustion, so the worker never reported STOPPED after stop_worker_service. test_macos_worker_restarts_process runs immediately before it and shares the same class-scoped worker. It killed the agent, then returned as soon as launchd showed a new pid and pgrep found a process -- both true within seconds of the respawn and well before the agent has registered. The next test then booted the service out mid-registration, so the agent died without reporting STOPPED and that test failed on a worker this one had left half-started. It now waits for the worker to report started again, which is the state the next test is entitled to assume from a class-scoped fixture. The worker id is bound to a local first, mirroring how the other tests in this package narrow it. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…doers drop-in Two review findings. _claim_host records the newcomer as holding the host before start() runs, which is right while an install is in flight and wrong once it has failed. On a failed reinstate the displaced incumbent was already stopped and this start never replaced it, so no agent held the host, yet the slot still named the session worker. The next claim routed that through stop_worker into a second DeleteWorker on a record start() never created, which raises, and _claim_host deliberately does not catch it: one failed reinstall refused every later claim. Routed through stop_worker rather than clearing the slot directly, matching what create_worker does on a failed start, so a start that registered before failing still has its record deleted. The existing test asserted _displaced_session_worker but not _installed_worker, so it did not catch this. Both are asserted now, plus a second test stating the consequence as behaviour. The sudoers drop-in granting env_keep for the AWS credential variables could outlive the build. A bare EXIT trap does not fire when the shell is killed rather than exiting, which is the case the rest of this script is built around, so a CodeBuild timeout or host reclaim left a host-wide Defaults env_keep on reserved capacity that no later build removed. That applies to every sudo on the host, not only the install it exists for, so a job launched through the agent's sudo -u <job-user> would see whatever credential environment the invoking process carried. The trap now also catches INT and TERM, and reset_agent_state removes a leftover by name for the kill -9 no trap can catch. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
create_worker's context manager wraps only start(); the yield is bare. So an AssertionError raised from this fixture's finally propagates straight out of the with, and the stop_worker that sat after it never ran. That made a failed restore -- the one condition this fixture is most defensive about -- also the condition that leaks the worker: on EC2 the instance and its record survive the rest of the build, which is what stop_worker's ConflictException retry exists to handle, and on macOS _installed_worker still names it so it is cleaned up only if some later test happens to claim the host. stop_worker now runs inside the finally, before the assert, and the trailing call after the with is gone so it is not attempted twice on the normal path. Also finishes the BUILD_RESIDUE half of the sudoers fix. The list still carried /var/root/.aws, left over from the credential-copying approach that was replaced, and did not carry the drop-in itself -- so a build killed by a signal no trap can catch would leave a host-wide env_keep for AWS_SECRET_ACCESS_KEY and the container authorization token on reserved capacity. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…eted Both failed-start paths called stop_worker with the default honor_keep_after_failure=True. Once any test had failed with the flag set, that takes stop_worker's early return: it does not stop the worker, does not clear the slot, and records it as kept. So after a failed reinstall _installed_worker named a worker with no agent behind it, _kept_worker named it too, and every later claim was refused with a message saying an agent was being kept for inspection -- which was not true, because the install never completed. That is the wall of setup errors _kept_worker was introduced to prevent, reached by a different route. The flag preserves a failed *test's* worker. A start that raised left nothing worth preserving, which is the same reasoning _claim_host already applies to its displacement stop, so both paths now pass honor_keep_after_failure=False and all three stop paths agree. The existing failed-reinstall tests could not catch this: isolated_registry clears KEEP_WORKER_AFTER_FAILURE, so they all ran the branch where honouring it is a no-op. The new test sets the flag and a prior failure, then asserts the slot is released, the worker is not recorded as kept, and the next claim proceeds. Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
andychoquette
enabled auto-merge (squash)
October 5, 2026 19:23
crowecawcaw
approved these changes
Oct 5, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces the capability probe in
pipeline/e2e-macos.shwith the suite. The probe answered its question on the reserved MAC_ARM fleet: passwordless sudo, a hidden sub-500 account viadscl,dseditgroupGID allocation, an/etc/sudoers.drule and a LaunchDaemon bootstrapped and running as that account are all permitted there.Reworked from #1080 rather than merged from it. That PR targeted a GitHub-hosted runner with no provisioned Deadline resources, so it carried a 366-line
_scaffold_deadline_resources. On CodeBuild those resources already exist and arrive as the SSM-backed environment variables the Linux and Windows projects use, so that whole path is dropped. #1080 can be closed once this lands.One host, one agent
LocalMacWorkerinstalls onto the machine running the tests, so two workers cannot coexist: the secondstart()overwrites the first'sworker.toml,worker.jsonand LaunchDaemon. Rather than skip every test that wants a differently-configured worker, macOS installs them one at a time. A new worker displaces the incumbent, and a displaced session worker is reinstalled on next use.This costs nothing.
pytest_collection_modifyitemsalready moves everysession_workertest to the end of the run, so--setup-planshows all 10 per-test worker fixtures created first and_session_worker_implonce after them, with nothing displacing it: 11 installs per run, one per worker fixture, and the reinstate path never fires. It earns its keep under-kor a reordering plugin, where the interleaving the ordering hook prevents becomes possible again.The registry is
monkeypatched rather than assigned in tests, and_MACOSis patched on rather than skipping off macOS, so the state machine is covered by the Linux and Windows runs that gate merges. 12 tests intest_worker_exclusivity.py.Coverage parity
macOS runs 62 of 96 against Linux 65 and Windows 71. Every remaining difference is a test the other non-native platform also skips, except one:
test_cap_kill(3)test_worker_shuts_down_host_machine_if_configured(1)Getting there needed more than widening skips:
== linux, but the bundles pinattr.worker.os.familytolinuxin all 10 steps across the two files, so widening the skip alone would have hung rather than passed: the worker reportsmacosand would never be assigned the job. Renamed_linuxto_posix.test_macos_worker_restarts_processis the launchd counterpart of the systemd and Windows-service tests. The installer writesKeepAlive { SuccessfulExit = false }, theRestart=on-failureanalog, so a SIGKILL is an unsuccessful exit and launchd respawns. BSDpgreprejects the Linux test's--count/--full, and launchd throttles respawns to one per 10s per job, so the wait is 120s rather than 30s.TestMacosJobUserOverrideports the four Linux override tests. The config-file case needs BSD sed's-i.bakplus agrep -qguard, since a sed matching nothing exits 0 and the test would then assert an override the agent never applied. The environment case has no systemd drop-in, so it writes the LaunchDaemon plist'sEnvironmentVariablesand pops the key in afinally. The three POSIX users it asserts on already exist, created byLocalMacWorkerfrom the sameworker_config.job_usersthe EC2 workers use.Rust fault injection is now reversible
TestRustUnavailableAndRecoverydeletedopenjd/model/_v1with no restore. On EC2 that dies with the instance; on a reserved-capacity Mac the venv is preserved across builds, so build N would leave build N+1 failingTestExplicitModeRouting[rust]— which runs before the class that broke it, reporting a missing adapter as a routing regression. The fixture now moves the tree aside and back, ande2e-macos.shrepairs a tree left behind by a build that died in between.Verified, and not
hatch run e2e:lintand the 13 exclusivity tests pass against the released 0.18.23; collection is 95 tests with 65/71/61 running across linux/windows/macos; thelaunchctl printparsing, BSDpgrepbehaviour, BSDsed -i, and/usr/bin/python3being 3.9.6 were all checked against a real Mac.Not verified: a suite run. Nothing here has executed against a real macOS worker. A dispatch is only possible from
mainline— branch creation is blocked on this repository and a fork PR gets no secrets — so the first run after merge is the validation run.An EC2 Mac was the alternative and is not available: Mac dedicated hosts are blocked account-wide on our dev accounts (every one of the 9 mac types refused with
UnsupportedHostConfigurationin all four us-east-1 AZs and in us-west-2, while a non-mac dedicated host allocated instantly — the quota reads 5, so the refusal is at a different layer). Lifting that needs an account exception request, tracked separately.What the first dispatch answers
e2e-macos.shlogs two things on the host specifically so this run settles them. Both are diagnostic: they print and cannot fail the build.IMDS reachability — this one can change code.
test_worker_requires_no_instance_profileis skipped on macOS on the grounds that a Mac has no IMDS, so_enforce_no_instance_profileraisesIMDSUnreachableErrorrather than the agent refusing work under an instance profile. But CodeBuild reserved capacity is EC2 Mac underneath, so IMDS may answer there and that skip may be hiding a test that would run. Read the output as:200oniam/info404/usr/bin/python3 --versionis the interpreterLocalMacWorkerbuilds the agent venv from, and the xfail onTestServiceSelectedFollowsServiceHintis justified by it being 3.9, which caps botocore below the release that modelsAssignedSession.metadata. If it reports newer, remove the xfail rather than leaving it to pass as an xpass.Beyond those two, the run measures what no amount of reading could: wall clock against the 180-minute timeout (11 installs, one per worker fixture), the 120s launchd respawn window in
test_macos_worker_restarts_process, the BSDsedpattern against the realworker.toml(guarded bygrep -q, so a mismatch fails loudly), and whether a macOS worker is assigned a job at all, which depends on it reportingattr.worker.os.family = macosagainst bundles that now admit it.Known residue on the reserved host, neither caused by this change:
_write_agent_credentialsleaves the agent user's~/.aws/credentialsin place, and a build killed mid-run leaks Deadline worker records that count against the fleet'smaxWorkerCount. Filed separately rather than bundled here.