Ship a C++ SDK in the wheel - #21639
Open
shoumikhin wants to merge 1 commit into
Open
Conversation
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21639
Note: Links to docs will display an error until the docs builds have been completed. ❌ 3 New Failures, 10 Pending, 6 Unrelated Failures, 8 Unclassified FailuresAs of commit 07628b2 with merge base fb5eedc ( NEW FAILURES - The following jobs have failed:
UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:
FLAKY - The following jobs failed but were likely due to flakiness present on trunk:
BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
shoumikhin
added a commit
that referenced
this pull request
Aug 7, 2026
## Why
The previous change split the runtime, kernels, delegate, thread pool and profiler out
of the Python extension into five prebuilt shared libraries. The wheel ships them, but
nothing outside Python can use them: the installed CMake package config names none of
the five, so a C++ application has no way to link them without hard-coding paths into
the wheel's private layout.
BEFORE AFTER
pip install executorch pip install executorch
| |
v v
executorch/lib/*.so executorch/lib/*.so
(shipped, but unnamed) |
v
a C++ app must find_package(executorch REQUIRED)
clone the repo and |
build from source v
target_link_libraries(app PRIVATE
executorch::runtime)
## What this change does
Gives the shipped libraries a public contract:
find_package(executorch 1.5 REQUIRED COMPONENTS kernels_optimized)
target_link_libraries(my_app PRIVATE executorch::runtime
executorch::kernels_optimized)
| target | library it resolves to |
| --- | --- |
| `executorch::runtime` | `libexecutorch.so` |
| `executorch::kernels_optimized` | `libexecutorch_kernels_optimized.so` |
| `executorch::backend_xnnpack` | `libexecutorch_backend_xnnpack.so` |
| `executorch::threadpool` | `libexecutorch_threadpool.so` |
| `executorch::etdump` | `libexecutorch_etdump.so` |
Namespaced rather than bare, because a name containing `::` must be an alias or
imported target, so CMake reports a missing one while configuring and names it. A bare
name is handed to the linker as `-lexecutorch`, which fails later with a worse message
or silently resolves to an unrelated system library. That matters more for a wheel than
for a source build: the wheel's contents depend on the options it was built with, so a
consumer asking for a delegate the wheel does not carry should be told during
configuration.
Each component target carries the retention its library needs. A registration-only
library has no symbol the application references, so the default `--as-needed` drops it
and its static initializer never runs, leaving a delegate that is linked and
unregistered. The options are scoped per library, because CMake removes duplicate
option text and a shared `--push-state` pair silently loses its scoping for the second
component.
## What to expect
Nothing changes for a Python user. This only adds a way to use the libraries the wheel
already shipped.
| | before | after |
| --- | --- | --- |
| C++ app links the runtime | build from source | `find_package(executorch)` |
| `find_package(executorch 1.5)` | any version accepted | version checked |
| headers for `Module` | not shipped | shipped |
The package also gains a version file, so `find_package(executorch 1.5 REQUIRED)`
answers correctly instead of accepting any request. Generated at packaging time rather
than checked in, because the version is only known then: `version.txt` gives the base
and a nightly overrides it. Without the file CMake reports the version as `unknown` and
accepts every request, so a consumer pinning a minimum silently gets whatever is
installed.
The headers move with the libraries. The package previously installed the subset a
custom-operator build needs, which does not include `extension/module`, the entry point
the documentation tells a C++ application to use. So the package shipped the libraries
to load and run a program and no way to call them. This adds `extension/module`, the
two directories holding the concrete allocator and loader a caller has to construct,
and `devtools/etdump`, whose library was already advertised as a component.
## Fixes from review of an earlier revision
The version file declared a variable for pinning an exact build that it never wrote, so
the config's own advice for that case compared against an empty string.
The thread pool switch sat on the thread pool target, while the header it guards is
exposed by every component and selects between a declaration and an inline definition.
A consumer naming that component in one translation unit and not another compiled two
definitions of the same function into one program, and the serial one silently won
wherever it was inlined. It now sits on the runtime, which every component depends on.
An interface link directory was carried with eleven lines defending it, while every
library already reaches the link line by absolute path. Removing it changes no build.
The relocation check skipped when `patchelf` was absent, which is indistinguishable
from a pass in the log. It now installs the tool and fails if it cannot.
Test plan:
A standalone application built from outside the wheel, in
`.ci/scripts/wheel/test_cpp_sdk.py`. It exports a real `.pte`, runs it from C++ through
`Module`, and compares the output against eager PyTorch, because a model that returns
wrong numbers without erroring satisfies every other check. Seven properties:
- `find_package` accepts the installed version and an older request, and rejects a
newer one
- linking only the runtime loads a program and reports every operator missing, which is
the split working rather than a defect, and it fails if the runtime starts carrying
kernels again
- adding the kernels component runs the model and matches eager PyTorch
- adding the delegate runs a delegated model and matches
- the same delegated program fails in an application that linked the kernels but not
the delegate, which is what shows the component is what registers it
- the application still runs after being copied away from the wheel with the absolute
search path removed, so the package is relocatable rather than only working where it
was built
- an application linking five components sees exactly one more backend than one linking
two, so there is one registry in the process rather than one per component
Ran against an installed wheel on x86_64 and aarch64. All seven pass, with the C++
output matching eager PyTorch to 2.4e-07 in every executing case.
ghstack-source-id: bca8af0
ghstack-comment-id: 5215967468
Pull-Request: #21639
This was referenced Aug 7, 2026
## The problem
The previous change split the runtime, kernels, delegate, thread pool and profiler into separate
shared libraries, and the wheel ships them. But nothing outside Python can use them, because the
installed CMake package names none of them. A C++ application would have to hard-code paths into
the wheel's private layout.
The headers have the same gap. The wheel installs only the subset a custom-operator build needs,
which leaves out `extension/module`, the entry point the documentation tells C++ callers to use. So
the wheel ships the libraries to run a model and no way to call them.
## The change
Name each shipped library as a CMake component, so `find_package` locates them, and ship the
headers a caller needs. A component is just a name a consumer can ask for, and CMake reports a
missing one while configuring rather than at link time.
```cmake
find_package(executorch 1.5 REQUIRED COMPONENTS kernels_optimized)
target_link_libraries(my_app PRIVATE executorch::runtime
executorch::kernels_optimized)
```
| component | library it resolves to |
| --- | --- |
| `executorch::runtime` | `libexecutorch.so` |
| `executorch::kernels_optimized` | `libexecutorch_kernels_optimized.so` |
| `executorch::backend_xnnpack` | `libexecutorch_backend_xnnpack.so` |
| `executorch::threadpool` | `libexecutorch_threadpool.so` |
| `executorch::etdump` | `libexecutorch_etdump.so` |
Each component records where the wheel keeps its libraries, so an application built against it
finds them without the caller setting a library search path.
CMake 3.28 or newer gets these targets. Older versions still get the long-standing
`EXECUTORCH_LIBRARIES` variable, carrying the runtime and the CPU kernels by path, because they write
the `$ORIGIN` marker (the "look next to me" token in a library search path) incorrectly:
```
3.24.3, 3.27.9 Makefiles double the dollar sign, Ninja drops the name
3.28.4, 3.31.8 both write the token correctly
```
## Test plan
Built the wheel, installed it into a clean environment, and built a C++ application against the
installed wheel alone:
- the application links the runtime, runs a model, and matches eager PyTorch, and still runs after
being copied away from the wheel.
- asking for a component the wheel does not ship fails while configuring, naming the component.
- a version request is honoured, including ranges.
- all 117 shipped headers can be included on their own, and one entry point per shipped component
also links against the shipped libraries. A header that compiles but has no implementation in any
shipped library would otherwise ship unnoticed: three such headers are excluded from the wheel,
found by linking every declaration they carry, and one dead declaration was removed.
- on real CMake 3.24 and 3.27, an application configures, builds and runs through
`EXECUTORCH_LIBRARIES`. The kernels need scoped retention there, because a registration-only
library exports nothing the application references and the linker drops it, which showed up as
"Missing operator" at run time rather than as a link error.
- `find_package` succeeds when the interpreter on PATH is not the one the wheel was built for. The
extension's own file name carries its suffix, so asking a different interpreter for it reported a
complete install as not found.
Ran on Linux x86_64 and aarch64. The macOS wheel keeps the fused extension and ships no separate
libraries, so these checks do not apply there and its smoke test does not run them.
ghstack-source-id: f39f080
ghstack-comment-id: 5215967468
Pull-Request: #21639
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
The previous change split the runtime, kernels, delegate, thread pool and profiler into separate
shared libraries, and the wheel ships them. But nothing outside Python can use them, because the
installed CMake package names none of them. A C++ application would have to hard-code paths into
the wheel's private layout.
The headers have the same gap. The wheel installs only the subset a custom-operator build needs,
which leaves out
extension/module, the entry point the documentation tells C++ callers to use. Sothe wheel ships the libraries to run a model and no way to call them.
The change
Name each shipped library as a CMake component, so
find_packagelocates them, and ship theheaders a caller needs. A component is just a name a consumer can ask for, and CMake reports a
missing one while configuring rather than at link time.
executorch::runtimelibexecutorch.soexecutorch::kernels_optimizedlibexecutorch_kernels_optimized.soexecutorch::backend_xnnpacklibexecutorch_backend_xnnpack.soexecutorch::threadpoollibexecutorch_threadpool.soexecutorch::etdumplibexecutorch_etdump.soEach component records where the wheel keeps its libraries, so an application built against it
finds them without the caller setting a library search path.
CMake 3.28 or newer gets these targets. Older versions still get the long-standing
EXECUTORCH_LIBRARIESvariable, carrying the runtime and the CPU kernels by path, because they writethe
$ORIGINmarker (the "look next to me" token in a library search path) incorrectly:Test plan
Built the wheel, installed it into a clean environment, and built a C++ application against the
installed wheel alone:
being copied away from the wheel.
also links against the shipped libraries. A header that compiles but has no implementation in any
shipped library would otherwise ship unnoticed: three such headers are excluded from the wheel,
found by linking every declaration they carry, and one dead declaration was removed.
EXECUTORCH_LIBRARIES. The kernels need scoped retention there, because a registration-onlylibrary exports nothing the application references and the linker drops it, which showed up as
"Missing operator" at run time rather than as a link error.
find_packagesucceeds when the interpreter on PATH is not the one the wheel was built for. Theextension's own file name carries its suffix, so asking a different interpreter for it reported a
complete install as not found.
Ran on Linux x86_64 and aarch64. The macOS wheel keeps the fused extension and ships no separate
libraries, so these checks do not apply there and its smoke test does not run them.