Skip to content

[UC] Execute a shard-to-shard transfer within a staging budget - #1

Open
0z5a wants to merge 1 commit into
uc/v02-b1-transfer-planfrom
uc/v02-b2-transfer-executor
Open

0z5a wants to merge 1 commit into
uc/v02-b1-transfer-planfrom
uc/v02-b2-transfer-executor

Conversation

@0z5a

@0z5a 0z5a commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

Execute the transfer plan from B1 #8623 with bounded staging and batched P2P. This PR remains stacked on uc/v02-b1-transfer-plan.

Tested parent: 75fe58c09c390e19a948a25aed2410a31616f538. The fork PR diff contains only the executor, its deepspeed.comm entry point and its tests.

The executor uses the physical-copy descriptor without logical_origin, includes a schema version in its schedule digest, and posts P2P through deepspeed.comm. execute_transfers agrees on a batch schedule once and packs small copies into one message per peer and direction within the staging budget. Each request carries a stable parameter key, so ranks that order otherwise identical parameters differently reject the batch before data moves. The committed tests cover FP32/FP16/BF16 copies, strided source packing, multiple rounds, replicated sources, digest disagreement, collective refusal, multi-tensor packing and key-order mismatch.

E2E and speed comparison

Two RTX 5090 GPUs on one host, FP32 and NCCL. Routes were synchronized on each GPU; each sample is the slower rank, and the table reports medians. Direct means plan once and execute shard-to-shard; rebuild means gather source shards, reconstruct the full parameter, then extract target shards. The same source and target layouts and exact-value oracle were used for each comparison.

Workload Rebuild→extract Per-tensor direct Batched direct Batched vs rebuild Correctness
GPTBigCode tiny checkpoint, 64 tensors / 83,161 elements 15.512 ms 62.728 ms 11.440 ms 1.36×; 5.48× vs per-tensor All target shards exact; model output max absolute difference 0
4096×4096 FP32 parameter, 67.1 MB 3.242 ms 2.300 ms Not measured Per-tensor direct: 1.41× All target shards exact

The model is optimum-internal-testing/tiny-random-gpt_bigcode-multi_query-False at revision 7e0d5dc36afe26cb54a250c39cde5b517b17a7be; the downloaded weight file SHA-256 is 9c0f7e2ae40319c9fd8d39a1d671cd68e50fc753dcc24a4e5c6b56d123732e9c. The model route generated 344 resolved copies. The tiny-model row is five paired samples per route on the same two GPUs; the 67.1 MB row is the earlier four-sample single-parameter run. The weight file was removed after validation.

BigCode transfer policy Segments Written elements Two-GPU copy
Default lowest holder 7 768 Exact, including guard regions
Holder with the target's private Q rows 6 768 Exact, including guard regions
Highest-ranked KV holder 7 768 Exact target shards

The six-segment policy selects source 0 for targets 0 and 1, and source 1 for targets 2 and 3, matching each target's private Q rows. Picking the highest KV holder for every target still needs seven segments. The committed two-GPU test executes all three plans against the target-map extract oracle.

Validation at the final branch code: B1 CPU tests 98 passed; B2 manual two-rank suite 12 cases per rank passed. The full 64-tensor model and 67.1 MB E2E timing rows were measured on the unchanged production transfer path before the review-only test and spec updates. The 8-rank disjoint-endpoint case remains unrun because those GPUs were in use by other processes.

Scope

Single node, one device per process, synchronous on return. No CUDA graph, overlap with compute, online rollout integration or topology-aware replica selection. Refs deepspeedai#8230 and deepspeedai#8252.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the uc/v02-b2-transfer-executor branch from 9cecb11 to e7ff688 Compare September 26, 2026 23:35
@0z5a
0z5a marked this pull request as ready for review September 27, 2026 04:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant