feat(vllm): update DP+EP to use --local-data-parallel-rank - #112
Draft
njhill wants to merge 2 commits into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This solves the same problem as #90 but keeps "external lb" mode (top-level vllm process per rank) rather than switching to "hybrid lb" mode which funnels all dp ranks' requests through a single frontend process and is undesirable.
Specifically this allows for the gpus assigned to each process to be a proper subset of its
CUDA_VISIBLE_DEVICESvalues.And for vLLM it also sets the
--local-data-parallel-rankarg, which is needed to identify which local GPUs to use in this case.Note that this requires a vLLM version which includes yet-to-be merged PR vllm-project/vllm#41164.
As such I'm assuming that the changes in this PR may need to be adjusted so that the new vllm DP launching behavior can be toggled based on the vLLM version being used?