Skip to content

feat: run the macOS e2e suite on CodeBuild - #1093

Open
andychoquette wants to merge 27 commits into
aws-deadline:mainlinefrom
andychoquette:macos-e2e-suite
Open

andychoquette wants to merge 27 commits into
aws-deadline:mainlinefrom
andychoquette:macos-e2e-suite

Conversation

@andychoquette

@andychoquette andychoquette commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Replaces the capability probe in pipeline/e2e-macos.sh with the suite. The probe answered its question on the reserved MAC_ARM fleet: passwordless sudo, a hidden sub-500 account via dscl, dseditgroup GID allocation, an /etc/sudoers.d rule and a LaunchDaemon bootstrapped and running as that account are all permitted there.

Reworked from #1080 rather than merged from it. That PR targeted a GitHub-hosted runner with no provisioned Deadline resources, so it carried a 366-line _scaffold_deadline_resources. On CodeBuild those resources already exist and arrive as the SSM-backed environment variables the Linux and Windows projects use, so that whole path is dropped. #1080 can be closed once this lands.

One host, one agent

LocalMacWorker installs onto the machine running the tests, so two workers cannot coexist: the second start() overwrites the first's worker.toml, worker.json and LaunchDaemon. Rather than skip every test that wants a differently-configured worker, macOS installs them one at a time. A new worker displaces the incumbent, and a displaced session worker is reinstalled on next use.

This costs nothing. pytest_collection_modifyitems already moves every session_worker test to the end of the run, so --setup-plan shows all 10 per-test worker fixtures created first and _session_worker_impl once after them, with nothing displacing it: 11 installs per run, one per worker fixture, and the reinstate path never fires. It earns its keep under -k or a reordering plugin, where the interleaving the ordering hook prevents becomes possible again.

The registry is monkeypatched rather than assigned in tests, and _MACOS is patched on rather than skipping off macOS, so the state machine is covered by the Linux and Windows runs that gate merges. 12 tests in test_worker_exclusivity.py.

Coverage parity

macOS runs 62 of 96 against Linux 65 and Windows 71. Every remaining difference is a test the other non-native platform also skips, except one:

Not run on macOS Why
test_cap_kill (3) Linux kernel capabilities. Windows skips these too.
test_worker_shuts_down_host_machine_if_configured (1) Needs to power the host off, and the macOS worker is the test host.

Getting there needed more than widening skips:

  • The POSIX job-attachment tests were gated on == linux, but the bundles pin attr.worker.os.family to linux in all 10 steps across the two files, so widening the skip alone would have hung rather than passed: the worker reports macos and would never be assigned the job. Renamed _linux to _posix.
  • test_macos_worker_restarts_process is the launchd counterpart of the systemd and Windows-service tests. The installer writes KeepAlive { SuccessfulExit = false }, the Restart=on-failure analog, so a SIGKILL is an unsuccessful exit and launchd respawns. BSD pgrep rejects the Linux test's --count/--full, and launchd throttles respawns to one per 10s per job, so the wait is 120s rather than 30s.
  • TestMacosJobUserOverride ports the four Linux override tests. The config-file case needs BSD sed's -i.bak plus a grep -q guard, since a sed matching nothing exits 0 and the test would then assert an override the agent never applied. The environment case has no systemd drop-in, so it writes the LaunchDaemon plist's EnvironmentVariables and pops the key in a finally. The three POSIX users it asserts on already exist, created by LocalMacWorker from the same worker_config.job_users the EC2 workers use.

Rust fault injection is now reversible

TestRustUnavailableAndRecovery deleted openjd/model/_v1 with no restore. On EC2 that dies with the instance; on a reserved-capacity Mac the venv is preserved across builds, so build N would leave build N+1 failing TestExplicitModeRouting[rust] — which runs before the class that broke it, reporting a missing adapter as a routing regression. The fixture now moves the tree aside and back, and e2e-macos.sh repairs a tree left behind by a build that died in between.

Verified, and not

hatch run e2e:lint and the 13 exclusivity tests pass against the released 0.18.23; collection is 95 tests with 65/71/61 running across linux/windows/macos; the launchctl print parsing, BSD pgrep behaviour, BSD sed -i, and /usr/bin/python3 being 3.9.6 were all checked against a real Mac.

Not verified: a suite run. Nothing here has executed against a real macOS worker. A dispatch is only possible from mainline — branch creation is blocked on this repository and a fork PR gets no secrets — so the first run after merge is the validation run.

An EC2 Mac was the alternative and is not available: Mac dedicated hosts are blocked account-wide on our dev accounts (every one of the 9 mac types refused with UnsupportedHostConfiguration in all four us-east-1 AZs and in us-west-2, while a non-mac dedicated host allocated instantly — the quota reads 5, so the refusal is at a different layer). Lifting that needs an account exception request, tracked separately.

What the first dispatch answers

e2e-macos.sh logs two things on the host specifically so this run settles them. Both are diagnostic: they print and cannot fail the build.

IMDS reachability — this one can change code. test_worker_requires_no_instance_profile is skipped on macOS on the grounds that a Mac has no IMDS, so _enforce_no_instance_profile raises IMDSUnreachableError rather than the agent refusing work under an instance profile. But CodeBuild reserved capacity is EC2 Mac underneath, so IMDS may answer there and that skip may be hiding a test that would run. Read the output as:

Output Meaning
token + 200 on iam/info The skip is wrong; a profile is attached and the test should run on macOS.
token + 404 IMDS answers but no profile is attached, so the agent would not refuse work and the test would fail rather than be inapplicable. The skip stays, but its reason needs rewording.
no token Confirms the reason as written.

/usr/bin/python3 --version is the interpreter LocalMacWorker builds the agent venv from, and the xfail on TestServiceSelectedFollowsServiceHint is justified by it being 3.9, which caps botocore below the release that models AssignedSession.metadata. If it reports newer, remove the xfail rather than leaving it to pass as an xpass.

Beyond those two, the run measures what no amount of reading could: wall clock against the 180-minute timeout (11 installs, one per worker fixture), the 120s launchd respawn window in test_macos_worker_restarts_process, the BSD sed pattern against the real worker.toml (guarded by grep -q, so a mismatch fails loudly), and whether a macOS worker is assigned a job at all, which depends on it reporting attr.worker.os.family = macos against bundles that now admit it.

Known residue on the reserved host, neither caused by this change: _write_agent_credentials leaves the agent user's ~/.aws/credentials in place, and a build killed mid-run leaks Deadline worker records that count against the fleet's maxWorkerCount. Filed separately rather than bundled here.

Platform gating, in both directions:

- The two Linux job-bundle tests were gated to skip only on Windows, so they ran
  against a macOS fleet and Deadline auto-cancelled them as NOT_COMPATIBLE. The other
  tests labelled "Linux specific" are left running, since nothing shows they are
  actually incompatible and skipping them would hide coverage.
- test_worker_shuts_down_host_machine_if_configured is skipped on macOS: it reads
  instance_id, which LocalMacWorker does not have, and expects the host to power off,
  which on a runner that hosts the worker ends the suite.
- test_session_root_dir shares the POSIX branch, under /private/tmp because the macOS
  system volume is sealed and /tmp is a symlink whose resolved form the agent reports.

Portability and correctness:

- The session_runtime switch used a bare `sed -i`, which is GNU-only; BSD sed takes
  the script as the backup suffix. The suffixed form works under both.
- The Metal probe piped system_profiler into grep under `set -e`, so a headless CI VM
  with no displays exited 1 and aborted the script before the Metal check it exists to
  perform. The inventory is informational and no longer fatal, and swift's stderr is
  redirected so a real failure is visible rather than looking like nothing ran.
- The runtime recovery test's stop confirmation is removed. Each branch named its own
  service manager, so on macOS the POSIX branch ran `systemctl is-active`, exited 127
  and satisfied both assertions immediately -- no settle time on the one platform where
  the restart raced. systemctl stop and Stop-Service are synchronous, so Linux and
  Windows lose nothing, and confirming the stop now sits behind stop_worker_service.
- The path mapping test asserts the job succeeded before checking outputs, since
  wait_until_complete treats FAILED as complete and the failure otherwise surfaced as
  an empty mapping with no indication of the cause.
- Worker annotations in the session runtime tests accept either concrete worker that
  can drive the agent's service, rather than only the EC2 one.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Replaces the capability probe in pipeline/e2e-macos.sh with the suite. The probe
answered its question on the reserved MAC_ARM fleet: passwordless sudo, a hidden
sub-500 account via dscl, dseditgroup GID allocation, an /etc/sudoers.d rule and a
LaunchDaemon bootstrapped and running as that account are all permitted there.

Reworked from aws-deadline#1080 rather than merged from it. That PR targeted a GitHub-hosted
runner with no provisioned Deadline resources, so it carried a 366-line
_scaffold_deadline_resources that created a farm, queues and fleets per run. On
CodeBuild those resources already exist and arrive as the SSM-backed environment
variables the Linux and Windows projects use, so that whole path is dropped along
with aws-deadline#1080's own workflow and its six diagnostic commits.

The bootstrap mirrors e2e.sh's TEST_TYPE block rather than shortening to
`hatch run e2e:test`. The reusable workflow sets TEST_TYPE=WHEEL, and with
WORKER_AGENT_WHL_PATH unset conftest installs the PUBLISHED agent instead of this
commit -- a green run that exercised nothing. The built wheel's existence is
asserted, because that fallback is silent.

conftest's create_worker gains the macOS branch, ahead of the EC2 one. This is the
half that is easy to miss: yielding LocalMacWorker is not enough, because the EC2
path asserts on SUBNET_ID and SECURITY_GROUP_ID and pulls a worker instance profile
out of bootstrap_resources, none of which exist for a worker that installs the agent
onto the host running the tests. It also drops the ec2_worker_type parameter, which
every call site passed the same fixture value for.

The probe's cleanup, stale gate and leftovers survey are replaced by a narrow
pre-build reset. LocalMacWorker's _reset_host_state clears only worker.toml and
worker.json within a run, so on reserved capacity a daemon left loaded by a killed
build and the worker id it registered would poison the next one. The account, group
and /opt/deadline venv are deliberately left: the installer is idempotent over them
and reusing them saves minutes. The bootout is polled rather than read once, since it
is asynchronous.

Requires deadline-cloud-test-fixtures 0.18.23, which wires MACOS to LocalMacWorker
(aws-deadline/deadline-cloud-test-fixtures#333). Until that releases, the e2e
requirement does not resolve and this cannot go green.

Verified locally as far as is possible without installing an agent on a developer
machine: 79 tests collect under OPERATING_SYSTEM=macos, e2e mypy and ruff are clean,
the macOS branch constructs with only configuration and deadline_client while AL2023
still requires SUBNET_ID, the CODEBUILD_BUILD_ID guard exits 1, and the wheel block
builds a real wheel and fails when the path does not resolve. Not verified: a suite
run, which needs the fixtures release and a dispatch from mainline.

The workflow stays dispatch-only. Adding the push trigger before one green run would
put a red macOS check on mainline.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
@github-actions github-actions Bot added the waiting-on-maintainers Waiting on the maintainers to review. label Sep 15, 2026
Comment thread test/e2e/test_session_runtime.py Outdated
Comment thread test/e2e/conftest.py Outdated
Comment thread test/e2e/test_session_runtime.py
Comment thread test/e2e/test_session_runtime.py
Comment thread test/e2e/conftest.py Outdated
Comment thread test/e2e/conftest.py
Comment thread test/e2e/test_worker_config.py
…hat need two

Review found that LocalMacWorker installs onto the host running the tests, so two
workers cannot coexist: the second start() overwrites the first's worker.toml,
worker.json and LaunchDaemon, and whichever configuration landed last is the one under
test. Every worker fixture in test_session_runtime.py creates its own worker, and
test_worker_config.py creates one per test, so on macOS all of them collide with the
session-scoped worker that 44 call sites depend on.

Skipping them was the obvious fix and it costs too much. macOS already runs 45 of the
79 collected tests against 56 on Linux and 62 on Windows, mostly for legitimate
reasons -- cap_kill is Linux capabilities, domain_user is Windows AD, and both
platforms' job-user and job-attachment variants are OS-specific. Skipping the
two-worker tests would have taken macOS to roughly 37 and removed session-runtime
routing entirely, which for a platform whose agent support is new is the more
interesting half.

So macOS installs them one at a time. Creating a worker stops whatever holds the host
first, and the session worker is reinstalled on next use if it was the one displaced --
lazily, so a run that never returns to it pays nothing. stop() leaves the venv and the
account, so reinstating is the installer plus a launchd bootstrap rather than a full
provision. Counts are unchanged at 56/62/45.

session_worker splits into a cached session-scoped impl and a function-scoped wrapper.
Without a per-use hook the displaced worker would stay stale: pytest hands back the
same object all session, so every later assertion would run against a host with no
agent. The wrapper is a no-op on Linux and Windows.

stop_worker gained a guard that turned out to be load-bearing. A displaced worker has
already been stopped, and stopping it again raises: stop() deletes the worker record
and the second DeleteWorker on the same id fails, so both the session worker's teardown
and any displaced per-test worker's would have thrown at the end of a run. Only the
worker still holding the host has anything to tear down.

Nothing changes for Linux or Windows. Both new call sites are behind `if _MACOS`, where
workers are separate instances and genuinely concurrent.

Six tests in test/e2e/test_worker_exclusivity.py cover the state machine directly,
including the double-stop that would have raised and a claim whose incumbent fails to
stop. They skip off macOS, where the registry is inert. Not verified end to end: that
needs the fixtures release and a dispatch from mainline, and the reinstall cost per
transition is unmeasured against the 180-minute timeout.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/conftest.py Fixed
Comment thread test/e2e/test_worker_exclusivity.py Outdated
Comment thread test/e2e/conftest.py Outdated
Comment thread test/e2e/conftest.py Outdated
Comment thread test/e2e/conftest.py Outdated
…w gaps

The rust fault injection moves openjd/model/_v1 aside instead of deleting it, and
moves it back in rust_unavailable_worker's teardown. On EC2 either would do,
because the instance is discarded with the class. LocalMacWorker installs into a
venv at /opt/deadline/worker on a reserved-capacity host, and e2e-macos.sh keeps
that venv deliberately, so a delete is permanent: build N would leave build N+1
failing TestExplicitModeRouting[rust], which runs before the class that broke it,
reporting a missing adapter as a routing regression.

Moved into the fixture rather than left in the test body, because that is the
scope that owns the host mutation, and the path is resolved once before the move:
openjd.model imports from _v1, so the import used to locate it stops working as
soon as it is gone.

e2e-macos.sh repairs a tree left aside by a build that died between the move and
the restore, which the fixture teardown cannot cover. It finds the directory
rather than asking python where the package is, for the same reason.

ec2_worker_type is resolved inside the EC2 branch. It raises ValueError for an
operating system it does not recognise, so resolving it up front would fail a
macOS run at fixture setup before the branch that wants LocalMacWorker could run.
That branch now names LocalMacWorker directly: ec2_worker_type is the override
point for swapping in an EC2 subclass, so a cast there asserted a relationship
the fixture does not guarantee and would have failed on the missing ec2_client
and subnet_id keywords.

_grab_bootstrap_log handles LocalMacWorker. macOS bootstrap -- account creation,
the sudoers rule, the LaunchDaemon -- is the part most likely to fail, and the
reusable workflow hides the CodeBuild log from GitHub, so without this the only
record of a start failure was the exception string. One command per file, because
LocalMacWorker runs under set -euo pipefail and a missing path would abort the
rest.

The session root moves from /private/tmp/mysessionroot to /opt/mysessionroot.
/private/tmp is mode 1777, so an unprivileged process could pre-create the
directory and have the agent adopt one it does not own. /Users/Shared is 1777 as
well; /opt is root:wheel 0755, on the writable data volume, and already holds the
agent's venv.

The stop confirmation before the service restart is restored with a macOS branch.
It was dropped because the POSIX branch names systemctl, which macOS has not got:
there `systemctl is-active` exits 127 with empty stdout and satisfies both
assertions immediately, contributing no settle time on the one platform where the
restart actually races. macOS asks launchd instead.

Collection is unchanged at 85 tests; runs/skips are 56/29 linux, 62/23 windows,
51/34 macOS.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/conftest.py
Comment thread test/e2e/test_session_runtime.py
Comment thread pipeline/e2e-macos.sh Outdated
…e registry on all platforms

_claim_host now stops the incumbent through stop_worker instead of calling
worker.stop() directly. A displaced worker was running jobs moments ago, which is
exactly when DeleteWorker returns ConflictException, and stop_worker is where the
retry for that lives. The bare call swallowed a transient conflict permanently:
_installed_worker moved to the newcomer regardless, so the incumbent's own
teardown took the already-displaced early return and nothing retried the delete.

reinstate_session_worker goes through _claim_host too, so it releases the host
before installing. A class-scoped per-test worker can still be alive when a
session_worker test reinstates; without this the session worker overwrote its
worker.toml in place and its teardown found nothing to stop, leaking the record.
_displaced_session_worker is now cleared after a successful start rather than
before it, so a failed reinstall stays displaced and the next use retries instead
of being handed a host with no agent.

stop_worker clears _installed_worker below the KEEP_WORKER_AFTER_FAILURE return,
not above it. A kept worker is installed and running, so releasing the slot let the
next _claim_host install straight over the state the flag exists to preserve --
destroying the post-mortem it promises while starting the new worker against a
LaunchDaemon and worker id it did not create. _claim_host now refuses that case
with a message naming the conflict, because only one agent fits on the host and
neither silent option is defensible.

The exclusivity tests use monkeypatch instead of assigning conftest's globals, and
patch _MACOS on rather than skipping off macOS. Assigning directly left a MagicMock
in _installed_worker for the rest of the session, which would have stopped the mock
in place of a real incumbent and sent real teardowns down the already-displaced
path. Patching _MACOS means the state machine is now covered by the Linux and
Windows runs that gate merges, not only by the platform it changes: 11 tests, up
from 6, passing on all three.

Four fixtures still declared ec2_worker_type without using it, so pytest resolved
it during setup and the unsupported-OS ValueError fired before create_worker could
branch. That hit test_session_root_dir, which is one of the tests this PR enables
on macOS, so the change would have gone untested behind a fixtures-compatibility
error.

Runs/skips: 67/23 linux, 73/17 windows, 56/34 macOS.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/test_job_attachments.py Outdated
Comment thread test/e2e/test_session_runtime.py Outdated
Comment thread test/e2e/conftest.py
Comment thread test/e2e/conftest.py Dismissed
Dropped when the module was rewritten to use monkeypatch, which failed
test_copyright_headers on every platform.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Mirrors the mutual-exclusion the fixtures package's own worker fixture applies.
Without it the Docker branch wins: the run reports macOS results from a Linux
container, and the macOS exclusivity registry claims and stops Docker workers
it was never written for.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/conftest.py Outdated
The third ordering bug in stop_worker, same shape as the KEEP_WORKER_AFTER_FAILURE
one: _installed_worker was cleared before _stop_with_retry(), so a stop that
failed left a live agent on a host the registry recorded as free, and the next
_claim_host installed straight over it -- the silent overwrite the registry
exists to prevent. Reachable on the first attempt, because the backoff gives up
on anything that is not a ConflictException: a throttle, an expired credential,
or a bootout failure all land there immediately.

The slot now clears in the try's else, so the registry says the host is free
exactly when it is. A failed stop keeps the worker recorded; the next claim sees
it and retries the stop -- covered by a test that fails the first stop, asserts
the worker stays recorded, then claims with a second worker and asserts the
incumbent was stopped again rather than overwritten.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
The header still described the capability probe this PR replaces: probe-specific
exit codes 1/2/3 that no longer exist, and a rationale saying nothing had
confirmed a CodeBuild macOS build can create users, write sudoers rules and
bootstrap a LaunchDaemon. A probe build did confirm all of that.

The workflow stays dispatch-only, but for the remaining reason rather than the
settled one: the suite has never completed a run on this fleet. The trigger note
now keys on a dispatched run passing instead of on the installer being proven.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/conftest.py Outdated
The launchctl stop gate was still vacuous. Keying on launchd's not-found message
OR the absence of a running state is satisfied by every error launchctl can emit,
because an error message has no state line either -- so any failure passed the
gate on the first attempt with no settle time, which is what the previous revision
claimed to fix. There are now three outcomes, not two: not-found is stopped, a
loaded job is stopped only if not running, and anything else is evidence of
neither and raises so the backoff keeps trying. Both call sites share one helper,
since both wanted the same predicate and both got it wrong the same way.

Verified the substring against a real launchctl: a missing label prints
'Could not find service "..." in domain for system', so the match itself was
right. A bad domain prints 'Unrecognized target specifier', which is the case the
old disjunction waved through.

KEEP_WORKER_AFTER_FAILURE no longer cascades. The refusal keyed on
'a test failed and the flag is set', and testsfailed never decreases, so one
failure anywhere converted every later macOS worker-creating test into a setup
error -- a wall of noise on exactly the run the flag exists to diagnose. It now
keys on _kept_worker, set only where stop_worker actually leaves a worker
installed. Displacement also passes honor_keep_after_failure=False, because the
first attempt at this moved the cascade rather than removing it: stop_worker
declined to stop during displacement, and _claim_host then recorded the newcomer
and installed it over a live agent.

Two mutations moved inside their try blocks, the same defect already fixed on the
rust fixture. _set_plist_env is a read-merge-write, so a failure after the plist
landed left the override in place with the finally unreached -- and bootout does
not rewrite the plist, so it survives into the next build and every job there runs
as the override user. test_config_file_user_override had the same gap around its
sed-then-grep, which poisons the rest of the current run.

BUILD_RESIDUE's comment claimed it covered logs, which it does not. It now states
why: _grab_bootstrap_log reads from /var/log/amazon/deadline after a start
failure, so wiping it first would destroy the only diagnostic for the failure
being hit. The plist and sudoers rule are absent for a different reason, that
every start() rewrites both wholesale.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Two facts the suite's correctness depends on and nothing has established.

test_worker_requires_no_instance_profile is skipped on macOS because a Mac has no
IMDS, so _enforce_no_instance_profile raises IMDSUnreachableError rather than the
agent refusing work under an instance profile. But CodeBuild reserved capacity is
EC2 Mac underneath, so IMDS may answer here and that skip may be hiding a test
that would run. The output distinguishes the three cases: a token plus a 200 on
iam/info means the skip is wrong, a 404 means IMDS answers with no profile so the
test would fail rather than be inapplicable, and no token confirms the skip's
stated reason.

/usr/bin/python3 --version is logged because it is the interpreter
LocalMacWorker builds the agent venv from, and the xfail on
TestServiceSelectedFollowsServiceHint is justified by it being 3.9 -- which caps
botocore below the release that models AssignedSession.metadata. If a future macOS
ships something newer the xfail should be removed, not left to pass as an xpass.

Diagnostic only: both print and neither can fail the build. --max-time 3 keeps a
silently dropped IMDS request from costing two minutes, and 169.254.169.254 is
link-local so nothing leaves the host. Verified off-EC2 that the whole block takes
3 seconds and reports unreachable rather than hanging.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
crowecawcaw
crowecawcaw previously approved these changes Oct 1, 2026
…test on macOS

The first dispatch failed after 16 seconds on PEP 668: the host's pip3 is
Homebrew's and refuses system-wide installs, so `pip3 install --upgrade hatch`
exits with externally-managed-environment. The Linux script gets away with the
same lines because it runs in a container.

--break-system-packages would work and is the wrong fix: this is reserved
capacity, so mutating the host's Homebrew python is a change every later build
inherits, which is the class of problem the rest of this script exists to avoid.
The tooling goes in a venv under the build directory, which CodeBuild already
cleans between builds, and PATH is prepended so hatch resolves to it.

test_worker_requires_no_instance_profile runs on macOS again. I had skipped it on
the assumption that a Mac host has no IMDS, so _enforce_no_instance_profile would
raise IMDSUnreachableError rather than the agent finding a profile and refusing
work. The diagnostic added for exactly this question disproved it on the real
host: CodeBuild reserved capacity is EC2 Mac underneath, IMDSv2 answers, and
iam/info returns 200, so a profile is attached and the test's premise holds the
same way it does on the Linux and Windows instances. The comment records that,
so the next person does not re-skip it on the same reasoning.

The other diagnostic confirmed /usr/bin/python3 is 3.9.6, which is what the
TestServiceSelectedFollowsServiceHint xfail depends on.

Runs are now 66 linux / 72 windows / 63 macOS of 95 collected.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…odeArtifact

The second dispatch got as far as LocalMacWorker.start() and failed in
_install_agent with AccessDenied on codeartifact:GetAuthorizationToken, as
arn:aws:sts::539109460720:assumed-role/customer-.../i-0c9b3233d0ad98a4b.

That identity is an instance profile, not the CodeBuild build role. send_command
runs every fixture command through sudo, whose env_reset strips AWS_* from the
environment, so the CLI falls back to IMDS -- which this host has, being EC2 Mac
underneath -- and picks up the CodeBuild host's own instance profile: an identity
in an AWS-owned account with no access to our CodeArtifact domain.
install_command_for_linux chains the login with && before pip install, so the
denial aborts the install.

On EC2 Linux the same lookup succeeds because the instance profile is one the
fixtures bootstrapped and it carries the access. The host's profile here is not
ours to configure, so the credentials go where a root shell will find them.

aws configure export-credentials rather than copying AWS_* directly: on CodeBuild
the build role usually arrives through the container credential endpoint rather
than as static keys, so there may be nothing in the environment to copy. The
transformation to a credentials file was verified locally end to end.

Written with mode 600, piped through stdin so no secret reaches a command line or
this log, removed by an EXIT trap, and listed in BUILD_RESIDUE so the next build
clears it too -- live credentials in root's home must not outlive the build that
needed them on reserved capacity.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…d them

The previous attempt wrote /var/root/.aws/credentials and root still resolved to
the host's instance profile. macOS sudoers carries env_keep+="HOME MAIL", so sudo
preserves the invoking user's HOME and the AWS CLI running as root looks in the
build user's home, not /var/root.

Root gets /var/root first, and the build user's home only as a fallback when root
still cannot see the build role. Writing [default] into the build user's home
shadows its own credentials, which on CodeBuild come from the container endpoint
and refresh, where a static snapshot does not -- and the suite is budgeted at up
to three hours. Paying that only when the first location fails keeps the refresh
intact on any host whose sudoers resets HOME.

Both locations are removed by the EXIT trap and listed in BUILD_RESIDUE.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
First real run: 33 passed, 3 failed, 11 errored. Both causes are fixed here.

The 11 errors were 'Worker agent install failed: exit_code: 254' -- the
CodeArtifact AccessDenied again, but only from 17:41 after succeeding from 17:24.
Copying the credentials to a file was the wrong shape: CodeBuild's build
credentials arrive through the container endpoint and rotate, so a file is a
snapshot that expires about 25 minutes in. A sudoers drop-in adding the AWS
credential variables to env_keep means root resolves credentials exactly as the
build user does, refresh included, and nothing secret is written to a disk that
outlives the build. Validated with visudo -cf before installing, because a
malformed file in /etc/sudoers.d breaks sudo for every user on reserved capacity.

The 3 failures were the tests this change enabled, all from one GNU-ism in the
shared bundles: BSD wc -l right-pads its count to 8 columns, so
[ "`ls ... | wc -l`" != "2" ] compares "       2" against "2" and fires on every
check. 24 of those across the two bundles are now numeric -ne, which ignores the
padding and behaves identically on GNU. Three other wc -l uses were already safe,
being unquoted in [ $COUNT -lt 3 ] where word-splitting strips it.

That is the third GNU-ism these bundles have yielded after /bin/env and the
os.family pin, all found only by running on the host.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
The run showed why it cannot work on macOS, and it is not the script.
complex_bundle/linux hardcodes /tmp/storageprofiletest in its attachment
manifests, while MacOSJobStorageProfile declares the resolved
/private/tmp/storageprofiletest -- correctly, since /tmp is a symlink there and
the agent reports resolved paths. Job attachment sync then fails with 'No path
mapping rule found for the source path /tmp/storageprofiletest' before any task
runs: step 1 FAILED, every dependent step CANCELED, the task NEVER_ATTEMPTED.

Enabling it on macOS needs a complex_bundle/macos variant carrying the resolved
paths, the same way the Windows suite has its own. That is an unwritten variant
rather than something this change can fix, so the test keeps its Linux-only skip
and the reason now says which difference blocks it.

The two dependency-data-flow tests stay enabled on macOS: they pass now that the
wc -l comparisons are numeric, and their bundle declares no storage profile
locations. macOS runs 62 of 95.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/conftest.py Outdated
Comment thread pipeline/e2e-macos.sh Outdated
…estart test

The run after the previous fixes came back 57 passed, 1 failed, and the failure
was not one of the new tests: test_worker_lifecycle_status_is_expected, which
predates this change and runs on all three platforms. is_worker_stopped backed off
to exhaustion, so the worker never reported STOPPED after stop_worker_service.

test_macos_worker_restarts_process runs immediately before it and shares the same
class-scoped worker. It killed the agent, then returned as soon as launchd showed
a new pid and pgrep found a process -- both true within seconds of the respawn and
well before the agent has registered. The next test then booted the service out
mid-registration, so the agent died without reporting STOPPED and that test failed
on a worker this one had left half-started.

It now waits for the worker to report started again, which is the state the next
test is entitled to assume from a class-scoped fixture. The worker id is bound to
a local first, mirroring how the other tests in this package narrow it.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
…doers drop-in

Two review findings.

_claim_host records the newcomer as holding the host before start() runs, which
is right while an install is in flight and wrong once it has failed. On a failed
reinstate the displaced incumbent was already stopped and this start never
replaced it, so no agent held the host, yet the slot still named the session
worker. The next claim routed that through stop_worker into a second
DeleteWorker on a record start() never created, which raises, and _claim_host
deliberately does not catch it: one failed reinstall refused every later claim.
Routed through stop_worker rather than clearing the slot directly, matching what
create_worker does on a failed start, so a start that registered before failing
still has its record deleted.

The existing test asserted _displaced_session_worker but not _installed_worker,
so it did not catch this. Both are asserted now, plus a second test stating the
consequence as behaviour.

The sudoers drop-in granting env_keep for the AWS credential variables could
outlive the build. A bare EXIT trap does not fire when the shell is killed
rather than exiting, which is the case the rest of this script is built around,
so a CodeBuild timeout or host reclaim left a host-wide Defaults env_keep on
reserved capacity that no later build removed. That applies to every sudo on the
host, not only the install it exists for, so a job launched through the agent's
sudo -u <job-user> would see whatever credential environment the invoking
process carried. The trap now also catches INT and TERM, and reset_agent_state
removes a leftover by name for the kill -9 no trap can catch.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
@andychoquette
andychoquette requested a review from a team as a code owner October 5, 2026 15:26
Comment thread test/e2e/test_session_runtime.py
create_worker's context manager wraps only start(); the yield is bare. So an
AssertionError raised from this fixture's finally propagates straight out of the
with, and the stop_worker that sat after it never ran. That made a failed restore
-- the one condition this fixture is most defensive about -- also the condition
that leaks the worker: on EC2 the instance and its record survive the rest of the
build, which is what stop_worker's ConflictException retry exists to handle, and
on macOS _installed_worker still names it so it is cleaned up only if some later
test happens to claim the host.

stop_worker now runs inside the finally, before the assert, and the trailing call
after the with is gone so it is not attempted twice on the normal path.

Also finishes the BUILD_RESIDUE half of the sudoers fix. The list still carried
/var/root/.aws, left over from the credential-copying approach that was replaced,
and did not carry the drop-in itself -- so a build killed by a signal no trap can
catch would leave a host-wide env_keep for AWS_SECRET_ACCESS_KEY and the container
authorization token on reserved capacity.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
Comment thread test/e2e/conftest.py Outdated
…eted

Both failed-start paths called stop_worker with the default
honor_keep_after_failure=True. Once any test had failed with the flag set, that
takes stop_worker's early return: it does not stop the worker, does not clear the
slot, and records it as kept. So after a failed reinstall _installed_worker named a
worker with no agent behind it, _kept_worker named it too, and every later claim
was refused with a message saying an agent was being kept for inspection -- which
was not true, because the install never completed. That is the wall of setup errors
_kept_worker was introduced to prevent, reached by a different route.

The flag preserves a failed *test's* worker. A start that raised left nothing worth
preserving, which is the same reasoning _claim_host already applies to its
displacement stop, so both paths now pass honor_keep_after_failure=False and all
three stop paths agree.

The existing failed-reinstall tests could not catch this: isolated_registry clears
KEEP_WORKER_AFTER_FAILURE, so they all ran the branch where honouring it is a
no-op. The new test sets the flag and a prior failure, then asserts the slot is
released, the worker is not recorded as kept, and the next claim proceeds.

Signed-off-by: andychoquette <78888816+andychoquette@users.noreply.github.com>
@andychoquette
andychoquette enabled auto-merge (squash) October 5, 2026 19:23

This branch was successfully deployed

1 active (outdated) deployment
mainline — a99a7374 Deployed Oct 2, 2026 by andychoquette via macOS E2E Test / E2E Tests #9
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on-maintainers Waiting on the maintainers to review.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants