Skip to content

feat(vllm): update DP+EP to use --local-data-parallel-rank - #112

Draft
njhill wants to merge 2 commits into
NVIDIA:mainfrom
njhill:vllm-dp-update
Draft

feat(vllm): update DP+EP to use --local-data-parallel-rank#112
njhill wants to merge 2 commits into
NVIDIA:mainfrom
njhill:vllm-dp-update

Conversation

@njhill

@njhill njhill commented Apr 28, 2026

Copy link
Copy Markdown

This solves the same problem as #90 but keeps "external lb" mode (top-level vllm process per rank) rather than switching to "hybrid lb" mode which funnels all dp ranks' requests through a single frontend process and is undesirable.

Specifically this allows for the gpus assigned to each process to be a proper subset of its CUDA_VISIBLE_DEVICES values.

And for vLLM it also sets the --local-data-parallel-rank arg, which is needed to identify which local GPUs to use in this case.

Note that this requires a vLLM version which includes yet-to-be merged PR vllm-project/vllm#41164.

As such I'm assuming that the changes in this PR may need to be adjusted so that the new vllm DP launching behavior can be toggled based on the vLLM version being used?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant