feat: add GB300 MiniMax-M3 AgentX recipes / 新增 GB300 MiniMax-M3 AgentX 配置 - #3232
RohitNagraj wants to merge 9 commits into
Conversation
新增 GB300 MiniMax-M3 Dynamo-vLLM AgentX 聚合与分离式部署配置,固定运行依赖,并更新 srt-slurm 以支持各工作进程独立配置 Mooncake 存储。
|
Thanks for the contribution!
中文感谢你的贡献!
|
将 GB300 MiniMax-M3 配置的变更日志关联到 PR #3232。
| #!/usr/bin/env bash | ||
| set -eo pipefail | ||
| # Download, verify, and install the published wheel inside each backend container. | ||
| uv pip install --system --no-deps --reinstall --require-hashes \ |
There was a problem hiding this comment.
thanks for the contribution @RohitNagraj but any specific reason that mooncake nightly need to installed here. When will vllm docker be updated instead with this new version?
| set -eo pipefail | ||
| # Download, verify, and install the published wheel inside each backend container. | ||
| uv pip install --system --no-deps --reinstall --require-hashes \ | ||
| 'https://test-files.pythonhosted.org/packages/f3/00/fb2847f5564864be132d0f70384dde5688689d9f91b2c81d2609632f36e4/mooncake_transfer_engine_cuda13-0.3.14.dev20260910-cp312-cp312-manylinux_2_28_aarch64.whl#sha256=c55fcc42cf189fcdecdcdf9d776ef84b348a9970a9bd3aad2fad64cb2d9e6f82' |
There was a problem hiding this comment.
🟡 (optional) All four disaggregated GB300 MiniMax-M3 recipes fail their setup step once this wheel is pruned, since it is fetched from test-files.pythonhosted.org, TestPyPI's sandbox file host, not the production files.pythonhosted.org used elsewhere in this repo (e.g. benchmarks/benchmark_lib.sh:1457 for bfcl_eval). TestPyPI explicitly does not guarantee permanence of uploaded files, unlike PyPI proper. Fix: host the Mooncake wheel on a durable, project-controlled or production PyPI-backed location (or vendor it) so disagg jobs do not depend on a testing sandbox that can delete the file at any time.
Extended reasoning...
setup-disagg.sh runs uv pip install --require-hashes 'https://test-files.pythonhosted.org/packages/.../mooncake_transfer_engine_cuda13-0.3.14.dev20260910-...whl#sha256=...' inside every prefill/decode container for the four disagg recipes (dep4-tep4-c64, dep4-tp2-c90/c100/c110). test-files.pythonhosted.org backs test.pypi.org, which is documented as a sandbox with no durability guarantee for uploaded files. If the upload is removed or expires, the download 404s, set -eo pipefail aborts the script, and every disagg job using this setup script fails at startup with no fallback. Compare to benchmarks/benchmark_lib.sh:1457 which pins bfcl_eval from the production files.pythonhosted.org host for the same durability reason.
Verification: nit. Verified: benchmarks/multi_node/srt-slurm-recipes/configs/minimaxm3-mooncake/setup-disagg.sh:5 fetches the Mooncake wheel from https://test-files.pythonhosted.org/packages/.../mooncake_transfer_engine_cuda13-0.3.14.dev20260910-...whl#sha256=.... test-files.pythonhosted.org is TestPyPI's sandbox file host, not the production files.pythonhosted.org used everywhere else in the repo (e.g.…
仅对固定校验和的 Mooncake wheel 安装清除无关的软件包版本覆盖项,并保留哈希校验。同步主分支并保留变更日志历史内容。
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36457683154 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36457683154 |
更新 MiniMax-M3 TP2 分离式部署配方的并发度,并同步重命名配方及更新引用。
同步当前 main,并保留 MiniMax-M3 GB300 配方及追加式变更日志条目。
同步将 TEP4 max-num-seqs-2 与 TP2 max-num-seqs-2 分离式部署点调整至选定并发度,更新配方名称及引用,并在保留当前 main 内容的前提下解决配置、变更日志和 srt-slurm 版本冲突。
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
更新 MiniMax-M3 配方目录和 srt-slurm 配置字段,并保留追加式性能变更日志条目。
移除 MiniMax-M3 TP2 并发数 20 的数据点,并保留其余配置。
将 MiniMax-M3 GB300 配方指向共用的 SRT AgentX 基准测试脚本。
Add six GB300 MiniMax-M3 Dynamo-vLLM AgentX configurations: aggregate TP8 at C1 and TP2 at C24; disaggregated DEP4 prefill with four TEP4 decode workers at C50 or five TP2 decode workers at C50/C75/C110.
Pin the ARM64 vLLM nightly image and Dynamo package, use EAGLE3-GQA with three draft tokens, and install the checksum-verified Mooncake wheel for disaggregated serving. Throughput uses the repository's automatic golden acceptance selection; evals retain real speculative verification.
Use the current shared srt-slurm v2.23.2 dependency, keep process-local GPU device selection through
set_visible_devices, and place the recipes under the currentinferencex-e2e/project layout.Validation: exact configuration and filtered-family matrix generation, all six srt-slurm dry-runs, focused changelog tests, changelog validation, YAML parsing, and shell syntax checks pass. Hardware sweep and eval validation for the updated commit are pending.
AI model disclosure
AI assisted implementation and validation. The exact model identifier could not be verified.
中文
新增六个 GB300 MiniMax-M3 Dynamo-vLLM AgentX 配置:聚合部署采用 TP8、并发数 1,以及 TP2、并发数 24;分离式部署采用 DEP4 预填充,搭配四个 TEP4 解码工作进程、并发数 50,或五个 TP2 解码工作进程、并发数 50/75/110。
固定 ARM64 vLLM nightly 镜像和 Dynamo 包版本,采用 EAGLE3-GQA 与三个草稿 token,并为分离式部署安装经过校验和验证的 Mooncake wheel。吞吐测试由仓库自动选择 golden acceptance,eval 保留真实的投机解码验证。
使用当前共享的 srt-slurm v2.23.2 依赖,通过
set_visible_devices保留进程级 GPU 设备选择,并将配方放置在当前inferencex-e2e/项目目录下。验证:精确配置与过滤后的配置族矩阵生成、全部六个 srt-slurm dry-run、针对性变更日志测试、变更日志验证、YAML 解析和 Shell 语法检查均通过。更新后提交的硬件 sweep 和 eval 验证尚待完成。
AI 模型披露
AI 辅助了实现和验证,但无法核实确切的模型标识。