Add Kimi K2.5 B200 workload#43
Conversation
Follow the vLLM Blackwell NVFP4 recipe and retain the existing Kimi accuracy, BFCL, and serving benchmark coverage. Co-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
The current vLLM nightly rejects use_trtllm_ragged_deepseek_prefill as an unknown AttentionConfig field during startup. Co-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
|
GPU validation update:
|
Co-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
|
Buildkite #315 never reached the workload: agent-stack failed the job after its Kubernetes pod remained Pending for 15 minutes. I updated the workload to the official recipe's supported 4x B200 / TP4 topology in da44766, reran the parser/generator/shell/tests locally, and will retrigger GPU validation when capacity is available. |
|
Buildkite #316 passed in 46m56s on |
|
Recipe-alignment check (B200) Compared this config against the structured recipe that renders recipes.vllm.ai — Close match to the NVFP4 command: model ✓, Two notes:
Automated recipe-alignment review, AI-assisted (Claude). |
This PR was authored with assistance from Codex.
Summary
Recipe: https://recipes.vllm.ai/moonshotai/Kimi-K2.5
Local validation
lm_evalregistrybash -n lib/run.sh lib/server.sh lib/run_lm_eval.sh lib/run_vllm_bench.shpython3 .buildkite/test_generate_pipeline.py(6/6 passed)WORKLOADS=workloads/kimi_k2_5_b200.yamlpipeline generationgit diff --checkGPU validation