Conversation
0z5a
force-pushed
the
uc/v02-b1-transfer-plan
branch
from
September 24, 2026 06:03
f4e1e16 to
b8f0f3b
Compare
0z5a
force-pushed
the
uc/v02-b2-transfer-executor
branch
from
September 24, 2026 06:05
8fcd090 to
9cecb11
Compare
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
0z5a
force-pushed
the
uc/v02-b2-transfer-executor
branch
from
September 26, 2026 23:35
9cecb11 to
e7ff688
Compare
0z5a
marked this pull request as ready for review
September 27, 2026 04:10
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Execute the transfer plan from B1 #8623 with bounded staging and batched P2P. This PR remains stacked on
uc/v02-b1-transfer-plan.Tested parent:
75fe58c09c390e19a948a25aed2410a31616f538. The fork PR diff contains only the executor, itsdeepspeed.commentry point and its tests.The executor uses the physical-copy descriptor without
logical_origin, includes a schema version in its schedule digest, and posts P2P throughdeepspeed.comm.execute_transfersagrees on a batch schedule once and packs small copies into one message per peer and direction within the staging budget. Each request carries a stable parameter key, so ranks that order otherwise identical parameters differently reject the batch before data moves. The committed tests cover FP32/FP16/BF16 copies, strided source packing, multiple rounds, replicated sources, digest disagreement, collective refusal, multi-tensor packing and key-order mismatch.E2E and speed comparison
Two RTX 5090 GPUs on one host, FP32 and NCCL. Routes were synchronized on each GPU; each sample is the slower rank, and the table reports medians. Direct means plan once and execute shard-to-shard; rebuild means gather source shards, reconstruct the full parameter, then extract target shards. The same source and target layouts and exact-value oracle were used for each comparison.
The model is
optimum-internal-testing/tiny-random-gpt_bigcode-multi_query-Falseat revision7e0d5dc36afe26cb54a250c39cde5b517b17a7be; the downloaded weight file SHA-256 is9c0f7e2ae40319c9fd8d39a1d671cd68e50fc753dcc24a4e5c6b56d123732e9c. The model route generated 344 resolved copies. The tiny-model row is five paired samples per route on the same two GPUs; the 67.1 MB row is the earlier four-sample single-parameter run. The weight file was removed after validation.The six-segment policy selects source 0 for targets 0 and 1, and source 1 for targets 2 and 3, matching each target's private Q rows. Picking the highest KV holder for every target still needs seven segments. The committed two-GPU test executes all three plans against the target-map
extractoracle.Validation at the final branch code: B1 CPU tests 98 passed; B2 manual two-rank suite 12 cases per rank passed. The full 64-tensor model and 67.1 MB E2E timing rows were measured on the unchanged production transfer path before the review-only test and spec updates. The 8-rank disjoint-endpoint case remains unrun because those GPUs were in use by other processes.
Scope
Single node, one device per process, synchronous on return. No CUDA graph, overlap with compute, online rollout integration or topology-aware replica selection. Refs deepspeedai#8230 and deepspeedai#8252.