Speed up permute propagation cleanup (#22161) - #22161
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22161
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit d5cbd3e with merge base 4ce2ec2 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@apullin has exported this pull request. If you are a Meta employee, you can view the originating Diff in D114224762. |
This PR needs a
|
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. Differential Revision: D114224762
858c82b to
bb32005
Compare
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. Differential Revision: D114224762
bb32005 to
86beae5
Compare
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%). Differential Revision: D114224762
86beae5 to
4985dfe
Compare
|
This needs a rebase with some of the other refactors we've been doing to these passes. |
|
@rascani Sadly, it seems there is a bit of a paradoxical situation with the stuck diff train. I haven't been able to find a place where I can rebase to where export will work and conflicts will be solved externally and on the oss internal signal. But having the approvals in place is great bc I can queue this up if/when the internal diff train catches up. |
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%). Differential Revision: D114224762
4985dfe to
d033cdf
Compare
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%). Differential Revision: D114224762
d033cdf to
d5cbd3e
Compare
Summary:
Speed up permute propagation cleanup.
ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:
Behavior-preserving speedup, all existing pass tests pass.
On an ensemble network of ~1M parameters, lowering produced identical output while running 102.5 seconds faster (26.7%).
Differential Revision: D114224762