Skip to content

Speed up permute propagation cleanup (#22161) - #22161

Open
apullin wants to merge 1 commit into
pytorch:mainfrom
apullin:export-D114224762
Open

Speed up permute propagation cleanup (#22161)#22161
apullin wants to merge 1 commit into
pytorch:mainfrom
apullin:export-D114224762

Conversation

@apullin

@apullin apullin commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Summary:

Speed up permute propagation cleanup.

ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:

  • Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping.
  • FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions.
  • PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform.

Behavior-preserving speedup, all existing pass tests pass.

On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%).

Differential Revision: D114224762

@apullin
apullin requested a review from digantdesai as a code owner August 25, 2026 21:14
@pytorch-bot

pytorch-bot Bot commented Aug 25, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22161

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d5cbd3e with merge base 4ce2ec2 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 25, 2026
@meta-codesync

meta-codesync Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

@apullin has exported this pull request. If you are a Meta employee, you can view the originating Diff in D114224762.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesync meta-codesync Bot changed the title Speed up permute propagation cleanup Speed up permute propagation cleanup (#22161) Aug 27, 2026
apullin added a commit to apullin/executorch that referenced this pull request Aug 27, 2026
Summary:

Speed up permute propagation cleanup.

ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:
- Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping.
- FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions.
- PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform.

Behavior-preserving speedup, all existing pass tests pass.

Differential Revision: D114224762
@apullin
apullin force-pushed the export-D114224762 branch from 858c82b to bb32005 Compare August 27, 2026 14:57
apullin added a commit to apullin/executorch that referenced this pull request Aug 27, 2026
Summary:

Speed up permute propagation cleanup.

ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:
- Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping.
- FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions.
- PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform.

Behavior-preserving speedup, all existing pass tests pass.

Differential Revision: D114224762
@apullin
apullin force-pushed the export-D114224762 branch from bb32005 to 86beae5 Compare August 27, 2026 15:21
apullin added a commit to apullin/executorch that referenced this pull request Aug 28, 2026
Summary:

Speed up permute propagation cleanup.

ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:
- Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping.
- FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions.
- PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform.

Behavior-preserving speedup, all existing pass tests pass.

On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%).

Differential Revision: D114224762
@apullin
apullin force-pushed the export-D114224762 branch from 86beae5 to 4985dfe Compare August 28, 2026 15:57
@rascani

rascani commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This needs a rebase with some of the other refactors we've been doing to these passes.

@apullin

apullin commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

@rascani Sadly, it seems there is a bit of a paradoxical situation with the stuck diff train. I haven't been able to find a place where I can rebase to where export will work and conflicts will be solved externally and on the oss internal signal.

But having the approvals in place is great bc I can queue this up if/when the internal diff train catches up.

apullin added a commit to apullin/executorch that referenced this pull request Sep 3, 2026
Summary:

Speed up permute propagation cleanup.

ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:

- Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping.
- FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions.
- PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform.

Behavior-preserving speedup, all existing pass tests pass.

On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%).

Differential Revision: D114224762
Summary:

Speed up permute propagation cleanup.

ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:

- Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping.
- FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions.
- PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform.

Behavior-preserving speedup, all existing pass tests pass.

On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%).

Differential Revision: D114224762
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported module: arm Issues related to arm backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants