Skip to content

Arm backend: Add serialized xlarge TOSA model suite - #21492

Merged
bdemirb merged 2 commits into
pytorch:mainfrom
bdemirb:baris_mletorch_deepseek_qwen_ci_memory_gate
Jul 30, 2026
Merged

Arm backend: Add serialized xlarge TOSA model suite#21492
bdemirb merged 2 commits into
pytorch:mainfrom
bdemirb:baris_mletorch_deepseek_qwen_ci_memory_gate

Conversation

@bdemirb

@bdemirb bdemirb commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

The DeepSeek-R1-Distill-Qwen layer-test PRs were reverted because they caused OOMs in the no-driver Arm TOSA model pytest job.

The normal TOSA model shard should keep using pytest-xdist auto parallelism so unrelated model tests do not slow down. Add a separate serialized suite for xlarge TOSA model tests instead.

The reland of the DeepSeek layer tests can mark the memory-heavy cases as xlarge and route them through this suite. This also gives internal CI an explicit option for running any xlarge TOSA model coverage serially.

cc @digantdesai @freddan80 @per @zingo @oscarandersson8218 @mansnils @Sebastian-Larsson @robell @rascani

The DeepSeek-R1-Distill-Qwen layer-test PRs were reverted because they
caused OOMs in the no-driver Arm TOSA model pytest job.

The normal TOSA model shard should keep using pytest-xdist auto
parallelism so unrelated model tests do not slow down. Add a separate
serialized suite for xlarge TOSA model tests instead.

The reland of the DeepSeek layer tests can mark the memory-heavy cases
as xlarge and route them through this suite. This also gives internal CI
an explicit option for running any xlarge TOSA model coverage serially.

Change-Id: I83c09ce0ed73f131e1d02d4c6e526c3ce53188c0
Signed-off-by: Baris Demir <baris.demir@arm.com>
@bdemirb
bdemirb requested a review from digantdesai as a code owner July 30, 2026 12:04
@pytorch-bot

pytorch-bot Bot commented Jul 30, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21492

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 873b1b9 with merge base b26b9ac (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 30, 2026
@github-actions github-actions Bot added ciflow/trunk module: arm Issues related to arm backend labels Jul 30, 2026
@bdemirb

bdemirb commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator Author

@pytorchbot label "partner: arm"

@pytorch-bot pytorch-bot Bot added the partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm label Jul 30, 2026
@bdemirb

bdemirb commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator Author

@pytorchbot label "release notes: arm"

@pytorch-bot pytorch-bot Bot added the release notes: arm Changes to the ARM backend delegate label Jul 30, 2026
@bdemirb
bdemirb merged commit e798c4d into pytorch:main Jul 30, 2026
494 checks passed
bdemirb added a commit to bdemirb/executorch that referenced this pull request Jul 30, 2026
Relands the DeepSeek-R1-Distill-Qwen-1.5B layer tests that were
reverted by pytorch#21048 because the TOSA model shard could
run several checkpoint-shaped exports concurrently and OOM.

The prerequisite CI change in pytorch#21492 added a
serialized xlarge TOSA model suite while keeping the normal TOSA model
shard parallelized. This reland marks the DeepSeek TOSA layer tests as
xlarge so they are excluded from the normal shard and can be routed
through the serialized xlarge suite.

The tests use the checkpoint configuration from the Hugging Face model
and the upstream Qwen2 layer implementations that back this distilled
model. The covered layers include rotary embedding, rotary application,
KV repetition, attention, RMSNorm, MLP, decoder layer, and final norm.

Token embedding is excluded because the full checkpoint embedding
allocation is too large for regular CI.

Signed-off-by: Baris Demir <baris.demir@arm.com>
Change-Id: Ia28581bbb4ffe070bc35af060fcceef2ac90084a
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: arm Issues related to arm backend partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm release notes: arm Changes to the ARM backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants