Split the runtime out of the Python bindings extension - #21610
Open
shoumikhin wants to merge 1 commit into
Open
Conversation
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21610
Note: Links to docs will display an error until the docs builds have been completed. ❌ 2 New Failures, 1 Pending, 1 Unrelated Failure, 8 Unclassified FailuresAs of commit cac57d5 with merge base cc853ae ( NEW FAILURES - The following jobs have failed:
UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:
FLAKY - The following job failed but was likely due to flakiness present on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
shoumikhin
added a commit
that referenced
this pull request
Aug 6, 2026
`pip install executorch` ships the C++ runtime welded inside the Python
bindings extension. That one file, `_portable_lib.so`, is about 11 MB and
contains the runtime, the operator registries, the CPU kernels, the XNNPACK
delegate, the profiler and several extensions, all statically linked together.
A Python user does not notice, but it has a correctness cost. Two other
extensions in the same wheel, the training bindings and the language model
runner, each carry their own private copy of the runtime. A process that loads
more than one of them ends up with more than one backend registry and more than
one thread pool. Backends register themselves from a static initialiser into a
table that is only process-wide if the runtime is loaded once, so a second copy
means a backend can land in a table nothing else reads.
This change builds each of those pieces once, as a shared library the wheel
ships, and makes the extensions link them instead of containing them:
executorch/lib/libexecutorch.so the runtime
executorch/lib/libexecutorch_kernels_optimized.so CPU operator kernels
executorch/lib/libexecutorch_backend_xnnpack.so the XNNPACK delegate
executorch/lib/libexecutorch_threadpool.so the shared thread pool
executorch/lib/libexecutorch_etdump.so the profiler
`_portable_lib.so` keeps only the Python binding glue and drops from about 11 MB
to well under 1 MB. Every other consumer in the tree that used to embed its own
copy now resolves it from the shared library, so a process has one registry and
one thread pool no matter how many of them it loads.
A `.pte` file produces the same outputs, the same operators are registered, and
the same backends are available. The existing header set the wheel installs is
untouched, and a custom-operator build against the Python extension keeps
working: the extension is at the same path and records a dependency on each
shared library.
That contract does need one addition. A custom-operator library references
runtime symbols and used to get them from the extension itself, which contained
the runtime. Now that the extension links the runtime instead, the installed
package config puts the shipped runtime on the extension's imported target, so
those symbols resolve from where the definitions actually moved to. Without it a
custom operator still links but fails to load with an undefined runtime
symbol.
Making these libraries linkable by a standalone C++ application needs a package
config that declares them, which is a separate change on top of this one. This
change is only the split, because it touches the most widely used artifact in
the wheel and should be reviewable on its own.
The libraries are built only where a shared runtime is possible, so macOS, iOS,
Android, Windows and embedded builds keep linking static libraries exactly as
before. Those platforms either have no dynamic loader at all, or in the case of
Windows a runtime with no export annotations for a DLL, so the static path stays
and is what the second branch of each build condition preserves. Asking for the
shared build on one of them is now an error at configure time rather than a
confusing packaging failure much later.
Two parts of this apply everywhere, because they are correct everywhere rather
than only in a wheel. A backend or kernel library that registers itself from a
static initialiser is now retained on a consumer's link line, which keeps a
shared backend from being dropped under `--as-needed`. And a Python extension no
longer links the embedding form of the Python library, which otherwise leaves a
hard dependency on an absolute path to the build machine's interpreter. Both
change link lines where they apply, and both fix a real problem rather than only
enabling this one.
Test plan:
Everything below runs against an installed wheel in a new virtual environment,
from a directory holding no checkout, so the checks cannot silently inspect the
source tree instead of the package.
- The wheel ships the five libraries above, each with a versioned file name and
a soname that matches it. A library without a version, which is what a
leftover from an earlier build looks like, fails.
- `_portable_lib.so` defines none of the registry, kernel, thread pool or
delegate symbols the shipped libraries own, and records a dependency on every
one of those libraries. Its size is reported rather than asserted, since the
number differs per architecture.
- For each of those symbol groups, exactly one shipped object defines it, and it
is the library that is supposed to own it. Naming the owner matters: the
monolithic layout also had exactly one definer of each, inside the extension.
- Every shipped library resolves its dependencies, and still resolves after
being copied into a different directory with the build-tree paths stripped, so
the wheel is relocatable rather than only working where it was built. No
shipped library carries an absolute runtime search path.
- Every Python extension the wheel ships imports in a clean environment. The
list is discovered from the installed package rather than written down, so an
extension added later is covered without anyone remembering to add it.
- A custom operator still compiles against the shipped Python extension, which
is the pre-existing contract for out-of-tree operators.
- The wheel's platform tag matches what it was built for.
- A model exported to a `.pte` and run through the Python bindings matches eager
PyTorch within tolerance, and so does the same model delegated to XNNPACK. The
delegated program is checked for the delegate's own identity, because a
partitioner that claimed nothing would fall back to the CPU kernels and still
match.
- All of the above on x86_64 and on aarch64. Both are necessary rather than
redundant: linker retention behaves differently between them, and x86_64 has
kept libraries that aarch64 dropped.
ghstack-source-id: d538f38
ghstack-comment-id: 5200527760
Pull-Request: #21610
shoumikhin
added a commit
that referenced
this pull request
Aug 6, 2026
`pip install executorch` ships the C++ runtime welded inside the Python
bindings extension. That one file, `_portable_lib.so`, is about 11 MB and
contains the runtime, the operator registries, the CPU kernels, the XNNPACK
delegate, the profiler and several extensions, all statically linked together.
A Python user does not notice, but it has a correctness cost. Two other
extensions in the same wheel, the training bindings and the language model
runner, each carry their own private copy of the runtime. A process that loads
more than one of them ends up with more than one backend registry and more than
one thread pool. Backends register themselves from a static initialiser into a
table that is only process-wide if the runtime is loaded once, so a second copy
means a backend can land in a table nothing else reads.
This change builds each of those pieces once, as a shared library the wheel
ships, and makes the extensions link them instead of containing them:
executorch/lib/libexecutorch.so the runtime
executorch/lib/libexecutorch_kernels_optimized.so CPU operator kernels
executorch/lib/libexecutorch_backend_xnnpack.so the XNNPACK delegate
executorch/lib/libexecutorch_threadpool.so the shared thread pool
executorch/lib/libexecutorch_etdump.so the profiler
`_portable_lib.so` keeps only the Python binding glue and drops from about 11 MB
to well under 1 MB. Every other consumer in the tree that used to embed its own
copy now resolves it from the shared library, so a process has one registry and
one thread pool no matter how many of them it loads.
A `.pte` file produces the same outputs, the same operators are registered, and
the same backends are available. The existing header set the wheel installs is
untouched, and a custom-operator build against the Python extension keeps
working: the extension is at the same path and records a dependency on each
shared library.
That contract does need one addition. A custom-operator library references
runtime symbols and used to get them from the extension itself, which contained
the runtime. Now that the extension links the runtime instead, the installed
package config puts the shipped runtime on the extension's imported target, so
those symbols resolve from where the definitions actually moved to. Without it a
custom operator still links but fails to load with an undefined runtime
symbol.
Making these libraries linkable by a standalone C++ application needs a package
config that declares them, which is a separate change on top of this one. This
change is only the split, because it touches the most widely used artifact in
the wheel and should be reviewable on its own.
The libraries are built only where a shared runtime is possible, so macOS, iOS,
Android, Windows and embedded builds keep linking static libraries exactly as
before. Those platforms either have no dynamic loader at all, or in the case of
Windows a runtime with no export annotations for a DLL, so the static path stays
and is what the second branch of each build condition preserves. Asking for the
shared build on one of them is now an error at configure time rather than a
confusing packaging failure much later.
Two parts of this apply everywhere, because they are correct everywhere rather
than only in a wheel. A backend or kernel library that registers itself from a
static initialiser is now retained on a consumer's link line, which keeps a
shared backend from being dropped under `--as-needed`. And a Python extension no
longer links the embedding form of the Python library, which otherwise leaves a
hard dependency on an absolute path to the build machine's interpreter. Both
change link lines where they apply, and both fix a real problem rather than only
enabling this one.
Test plan:
Everything below runs against an installed wheel in a new virtual environment,
from a directory holding no checkout, so the checks cannot silently inspect the
source tree instead of the package.
- The wheel ships the five libraries above, each with a versioned file name and
a soname that matches it. A library without a version, which is what a
leftover from an earlier build looks like, fails.
- `_portable_lib.so` defines none of the registry, kernel, thread pool or
delegate symbols the shipped libraries own, and records a dependency on every
one of those libraries. Its size is reported rather than asserted, since the
number differs per architecture.
- For each of those symbol groups, exactly one shipped object defines it, and it
is the library that is supposed to own it. Naming the owner matters: the
monolithic layout also had exactly one definer of each, inside the extension.
- Every shipped library resolves its dependencies, and still resolves after
being copied into a different directory with the build-tree paths stripped, so
the wheel is relocatable rather than only working where it was built. No
shipped library carries an absolute runtime search path.
- Every Python extension the wheel ships imports in a clean environment. The
list is discovered from the installed package rather than written down, so an
extension added later is covered without anyone remembering to add it.
- A custom operator still compiles against the shipped Python extension, which
is the pre-existing contract for out-of-tree operators.
- The wheel's platform tag matches what it was built for.
- A model exported to a `.pte` and run through the Python bindings matches eager
PyTorch within tolerance, and so does the same model delegated to XNNPACK. The
delegated program is checked for the delegate's own identity, because a
partitioner that claimed nothing would fall back to the CPU kernels and still
match.
- All of the above on x86_64 and on aarch64. Both are necessary rather than
redundant: linker retention behaves differently between them, and x86_64 has
kept libraries that aarch64 dropped.
ghstack-source-id: f879887
ghstack-comment-id: 5200527760
Pull-Request: #21610
This was referenced Aug 7, 2026
Contributor
Author
|
@claude review |
## The problem `pip install executorch` gives you the Python half of ExecuTorch and nothing a C++ program can link. Everything is fused into one large Python extension file, so a C++ developer has to clone the repository, sync submodules, and build from source. There is a correctness problem underneath the packaging one. Because the runtime is fused into the extension, anything else that needs it gets its own private copy, and two copies mean two registries. A backend registered in one is invisible to the other. ## The change Build the runtime and the pieces around it as separate shared libraries, and make the Python extension link them instead of embedding them. A shared library is a file a program loads at run time, so several programs can use one copy instead of each carrying its own. ``` executorch/ lib/libexecutorch.so the runtime lib/libexecutorch_kernels_optimized.so CPU operator kernels lib/libexecutorch_backend_xnnpack.so the XNNPACK delegate lib/libexecutorch_threadpool.so one thread pool per process lib/libexecutorch_etdump.so the profiler extension/pybindings/_portable_lib.so now under a megabyte, links the above ``` The extension is much smaller than before, because it no longer contains what it now links. Linux only, and only when the CUDA backend is off. macOS and Windows keep the fused extension because the split relies on ELF sonames, the `$ORIGIN` search-path token and GNU linker options, none of which apply there, and enabling it elsewhere now fails while configuring rather than much later. A CUDA build also keeps the fused extension for now, because the CUDA libraries are not yet shipped alongside the others. ## Test plan Built the wheel from source, installed it into a clean environment, and checked: - exactly one library defines each component, and it is the library that should own it. Counting owners alone would also pass on the old fused layout, which has exactly one too. - the Python extension defines none of them and resolves all of them from outside. - every shipped library loads with no unresolved dependency, and none of them searches a directory from the machine that built the wheel. The build-directory patterns are matched as whole path components, because a bare substring also matched an unrelated directory a user could really have, such as `/home/user/cmake-outputs/torchlibs`, and stripping that breaks a dependency the library legitimately resolves there. - a read-only build output is no longer archived world-writable. The previous code granted write to the group and to everyone via an absolute chmod mode; the fix only grants owner-write, so a `0555` file lands as `0755` instead of `0777`. - a custom operator library compiles and links against the shipped Python extension, which is the existing contract this must not break. Linking these libraries from a standalone C++ application needs the installed CMake package, which a later change adds. - the ahead-of-time quantized library records its route to the runtime through the same helper the other targets use. It was written by hand in two blocks that between them covered only the wheel layout and only when the wheel flag was set, so a plain `-DEXECUTORCH_BUILD_SHARED=ON` build left it with no route at all. Checked all four combinations of the shared build and the presence of the Python extension. Ran on Linux x86_64 and aarch64, including a Jetson device. Not fixed here: these libraries bundle third-party code that torch also links, and both keep it visible, so a process holds two definitions of symbols like `pthreadpool_create`. A caller reaches whichever the loader found first. Fixing it means hiding or dropping the bundled copies, which is a larger change. ghstack-source-id: 965d8df ghstack-comment-id: 5200527760 Pull-Request: #21610
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
pip install executorchgives you the Python half of ExecuTorch and nothing a C++ program canlink. Everything is fused into one large Python extension file, so a C++ developer has to clone
the repository, sync submodules, and build from source.
There is a correctness problem underneath the packaging one. Because the runtime is fused into the
extension, anything else that needs it gets its own private copy, and two copies mean two
registries. A backend registered in one is invisible to the other.
The change
Build the runtime and the pieces around it as separate shared libraries, and make the Python
extension link them instead of embedding them. A shared library is a file a program loads at run
time, so several programs can use one copy instead of each carrying its own.
The extension is much smaller than before, because it no longer contains what it now links.
Linux only, and only when the CUDA backend is off. macOS and Windows keep the fused extension
because the split relies on ELF sonames, the
$ORIGINsearch-path token and GNU linker options,none of which apply there, and enabling it elsewhere now fails while configuring rather than much
later. A CUDA build also keeps the fused extension for now, because the CUDA libraries are not yet
shipped alongside the others.
Test plan
Built the wheel from source, installed it into a clean environment, and checked:
owners alone would also pass on the old fused layout, which has exactly one too.
from the machine that built the wheel. The build-directory patterns are matched as whole path
components, because a bare substring also matched an unrelated directory a user could really
have, such as
/home/user/cmake-outputs/torchlibs, and stripping that breaks a dependency thelibrary legitimately resolves there.
to the group and to everyone via an absolute chmod mode; the fix only grants owner-write, so a
0555file lands as0755instead of0777.existing contract this must not break. Linking these libraries from a standalone C++ application
needs the installed CMake package, which a later change adds.
other targets use. It was written by hand in two blocks that between them covered only the wheel
layout and only when the wheel flag was set, so a plain
-DEXECUTORCH_BUILD_SHARED=ONbuild leftit with no route at all. Checked all four combinations of the shared build and the presence of the
Python extension.
Ran on Linux x86_64 and aarch64, including a Jetson device.
Not fixed here: these libraries bundle third-party code that torch also links, and both keep it
visible, so a process holds two definitions of symbols like
pthreadpool_create. A caller reacheswhichever the loader found first. Fixing it means hiding or dropping the bundled copies, which is a
larger change.