feat: add DSV4 FP4 B300 Dynamo-SGLang STP configuration / 新增 DSV4 B300 Dynamo-SGLang STP 配置 - #2362
Conversation
中文:新增 DSV4 FP4 B300 Dynamo-SGLang STP 配置,并接入共享 srt-slurm 配方矩阵。
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
中文:将性能变更日志链接更新为 #2362。
中文:将 B300 DSV4 Dynamo-SGLang 配置指向服务器上预置的 DeepSeek-V4-Pro-NVFP4 模型路径。
固定 srt-slurm 配方版本,确保 DSV4 B300 Dynamo-SGLang 运行使用可复现的配置。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30313591736 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30313591736 |
修复(dsv4):固定包含 UCX CUDA 传输配置的 srt-slurm 版本。
中文:将当前 main 合并到 DSV4 SGLang STP 配置分支。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30341010036 |
Use the stock DeepSeek-V4-Pro FP4 checkpoint, correct recipe labeling, and pin the srt-slurm checkout.\n\n中文:使用标准 DeepSeek-V4-Pro FP4 检查点,修正配置标注,并固定 srt-slurm 版本。
Pin NVIDIA/srt-slurm main, overlay the checked-in recipes, and keep STP chat-template formatting disabled. 修复:固定 NVIDIA/srt-slurm main 提交,覆盖仓库内置配方,并保持 STP 聊天模板格式化关闭。
同步最新 main 分支,并将 DSV4 STP 变更日志条目保留在文件末尾。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30431262611 |
|
/reuse-sweep-run 30431262611 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — |
|
/reuse-sweep-run |
# Conflicts: # perf-changelog.yaml
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — |
Summary
deepseek-ai/DeepSeek-V4-Procheckpoint staged at/scratch/models/DeepSeek-V4-ProNVIDIA/srt-slurmfrommain, pin commitc180328b98c3793ca84a1e24a030f90545eb7d5d, and overlay five recipes checked into this repositoryuse_chat_template: falsefor STP inputs and setUCX_TLS=cuda_copy,rcwithout aUCX_NET_DEVICESallowlist--no-preflightscoped to the compute-node-local/scratch/modelspathmainafter it is merged中文说明
deepseek-ai/DeepSeek-V4-Pro检查点,节点本地路径为/scratch/models/DeepSeek-V4-Promain克隆NVIDIA/srt-slurm,固定到提交c180328b98c3793ca84a1e24a030f90545eb7d5d,并覆盖本仓库内置的五个 recipeuse_chat_template: false,设置UCX_TLS=cuda_copy,rc,且不添加UCX_NET_DEVICES白名单/scratch/models路径启用--no-preflightmain中的 recipe