Skip to content

Latest commit

 

History

History
227 lines (178 loc) · 11.4 KB

File metadata and controls

227 lines (178 loc) · 11.4 KB

Building vllm.cpp

vllm.cpp uses CMake (>= 3.24) and a C++20 compiler (gcc 13/14 and clang are exercised; the tree builds -Werror-clean on gcc 14.2). The core has no ML dependencies; the OpenAI server uses a vendored header-only HTTP transport (cpp-httplib). The README carries the two-line quickstart; this page is the full build reference.

CPU build (the correctness / CI reference)

cmake -S . -B build
cmake --build build -j
ctest --test-dir build

The server is ON by default. Example binaries land under build/examples/: vllm-cli, server, vllm-bench, and tokenize.

CUDA build (NVIDIA GB10 / DGX Spark)

cmake -S . -B build-cuda \
  -DVLLM_CPP_CUDA=ON \
  -DVLLM_CPP_TRITON=ON \
  -DVLLM_CPP_CUTLASS_FETCH=ON
cmake --build build-cuda -j

Triton-AOT cubins for the fast GDN path are vendored: Python and Triton are needed only to regenerate them (VLLM_CPP_TRITON_REGEN), never to build or run them.

CUTLASS: the one external build dependency

CUTLASS (>= 4.5.0) is header-only, and it is the only thing a CUDA build fetches from the network. It feeds two independent consumers:

  • FlashAttention-2 prefill/decode, on every arch in 8.0 8.6 8.7 8.9 12.0a 12.1a.
  • The sm_12xa NVFP4 block-scaled GEMM, on Blackwell only.

Without it the FA2 kernels are not compiled and attention falls back to the portable path, which is slower. Nothing fails and no test goes red, so it is worth being deliberate about. Pick one:

-DVLLM_CPP_CUTLASS_FETCH=ON            # download CUTLASS 4.5.0 (~200 MB, needs network)
-DVLLM_CPP_CUTLASS_DIR=/path/to/cutlass  # reuse a checkout you already have

The default is neither, so that a disk- or network-constrained box configures without surprises. When you skip it on an arch that supports FA2, configure prints a CMake Warning saying so.

Confirm you got it from the configure output:

-- CUDA feature fa2: ENABLED for [86]
-- FlashAttention-2 prefill/decode: ENABLED for arch(es) [86] (runtime toggles VT_FA2_PREFILL, VT_FA2_DECODE)

The first line means the arch supports FA2; the second means it was actually built. If only the first appears, CUTLASS was not found.

Other CUDA families

Set the arch explicitly. It defaults to 121a (GB10), which will not load on anything else:

# Hopper H100/H200
cmake -S . -B build-cuda -DVLLM_CPP_CUDA=ON -DVLLM_CPP_CUDA_ARCHITECTURES=90a \
  -DVLLM_CPP_CUTLASS_FETCH=ON

# Ampere consumer (RTX 3090 = 86), Ada (89), Jetson Orin (87)
cmake -S . -B build-cuda -DVLLM_CPP_CUDA=ON -DVLLM_CPP_CUDA_ARCHITECTURES=86 \
  -DVLLM_CPP_CUTLASS_FETCH=ON

Of these, only sm_87 (Orin) and sm_110 (Thor) have been run on real hardware. sm_80/86/89, sm_90a and sm_100a/103a are build-verified: they compile -Werror-clean and emit the expected SASS, but no board here has executed them. See STATUS.md for what that label means and .agents/specs/cuda-arch-ampere-fastpath.md for the per-arch detail. Reports from those boards are welcome.

Metal build (Apple Silicon)

Metal is detected automatically on an Apple host with an ObjC++ compiler. The optional MLX GEMM provider is a separate opt-in and needs an MLX install:

cmake -S . -B build-metal -DVLLM_CPP_MLX=ON -DMLX_ROOT=/path/to/mlx
cmake --build build-metal -j

MLX is shape-gated to prefill, where it wins; it declines the m < 2 decode GEMV by design. Note that an MLX build produces a different greedy sequence than the default build (MLX's GEMM is not bit-identical), so goldens must not be re-anchored to an MLX build. Details in docs/BENCHMARKS.md.

Vulkan build

Headers are vendored and SPIR-V is committed, so no graphics toolchain is needed. It is off unless requested:

cmake -S . -B build-vulkan -DVLLM_CPP_VULKAN=ON
cmake --build build-vulkan -j

ROCm build (AMD GPUs) — never yet compiled

Read this before you file a bug. The HIP sources in this tree have never been compiled by anyone. There is no AMD GPU and no ROCm toolchain on any machine the maintainers use, so unlike every other backend here this one has no build report at all — not even "it compiles". If it fails for you, that is the expected first outcome and the most useful thing you can report. Please do, on issue #41.

cmake -S . -B build-hip -DVLLM_CPP_HIP=ON -DVLLM_CPP_HIP_ARCHITECTURES=gfx1100
cmake --build build-hip -j
ctest --test-dir build-hip -R 'rocm|cross_device'

VLLM_CPP_HIP_ARCHITECTURES is optional: leave it empty and hipcc targets the installed GPU, which is what you want when building on the machine you will run on. The validated names are upstream vLLM's HIP_SUPPORTED_ARCHS; anything else configures with a warning and is passed to hipcc anyway. If ROCm lives outside /opt/rocm, point at it with -DROCM_PATH=<prefix>.

-DVLLM_CPP_HIP=ON fails the configure when no HIP compiler is found rather than quietly producing a CPU-only build, for the same reason the CUTLASS note above exists: a silent downgrade is indistinguishable from success.

What exists today is the W0 skeleton — the vt::Backend, the Platform, one registered kernel (RmsNorm), and the tests that gate them. What that does and does not get you, and where to start on your specific board, is docs/ROCM.md.

Nix shells

The checked-in flake pins CMake, Ninja and the CUDA toolchain, so nothing has to be globally installed:

# CPU (correctness / CI reference)
nix develop .#default --command cmake -S . -B build-nix-cpu -G Ninja \
  -DVLLM_CPP_CUDA=OFF -DCMAKE_BUILD_TYPE=RelWithDebInfo
nix develop .#default --command cmake --build build-nix-cpu -j4

# CUDA (set the arch for your GPU)
nix develop .#cuda --command bash -lc \
  'cmake -S . -B build-nix-cuda -G Ninja -DVLLM_CPP_CUDA=ON \
    -DCMAKE_CUDA_COMPILER="$CMAKE_CUDA_COMPILER" \
    -DCMAKE_CUDA_HOST_COMPILER="$CMAKE_CUDA_HOST_COMPILER" \
    -DVLLM_CPP_CUDA_ARCHITECTURES=120a -DCMAKE_BUILD_TYPE=RelWithDebInfo'
nix develop .#cuda --command cmake --build build-nix-cuda -j4

On NixOS the CUDA shell exports the driver-library path and TRITON_LIBCUDA_PATH=/run/opengl-driver/lib so Triton finds libcuda without /sbin/ldconfig.

CMake options

Read from CMakeLists.txt. Defaults shown are the shipped defaults.

Option Default Purpose
VLLM_CPP_CUDA AUTO Build the CUDA backend: ON, OFF, or AUTO (on when a CUDA toolchain is found)
VLLM_CPP_CUDA_ARCHITECTURES 121a Target CUDA arch(s): 121a (GB10), 120a/120a;121a (consumer Blackwell), and cross-family targets 90a, 80/86/87/89, 100a/103a, 110. The a suffix is required for the native fp4 MMA
VLLM_CPP_METAL AUTO Build the Metal backend: ON, OFF, or AUTO (on for an Apple host with an ObjC++ compiler)
VLLM_CPP_VULKAN AUTO (= OFF) Build the Vulkan backend. Opt-in with -DVLLM_CPP_VULKAN=ON; headers are vendored and SPIR-V is committed
VLLM_CPP_HIP AUTO (= OFF) Build the ROCm/HIP backend. Opt-in with -DVLLM_CPP_HIP=ON, which fails loudly if no hipcc is found. Never compiled by anyone — see the ROCm section above
VLLM_CPP_HIP_ARCHITECTURES (empty) Target gfx arch(es), e.g. gfx1100 or gfx1100;gfx1151. Empty means hipcc targets the installed GPU
ROCM_PATH /opt/rocm ROCm installation prefix, for a nightly/TheRock install elsewhere
VLLM_CPP_MLX OFF Build the optional MLX GEMM provider for Metal (needs -DMLX_ROOT=<mlx install>)
MLX_ROOT (empty) Root of an MLX install (include/ + lib/) for VLLM_CPP_MLX
VLLM_CPP_SERVER ON Build the OpenAI HTTP server (needs third_party/httplib/httplib.h; disables itself with a warning if absent)
VLLM_CPP_TRITON OFF Consume the vendored per-arch Triton-AOT GDN cubins (CUDA only; no Python needed)
VLLM_CPP_TRITON_REGEN OFF Maintainer knob: regenerate the AOT cubins with Python + Triton
VLLM_CPP_CUTLASS_DIR third_party/cutlass CUTLASS source root (>= 4.5.0). Feeds the sm120a NVFP4 GEMM and FlashAttention-2 on 8.0/8.6/8.7/8.9/12.0a/12.1a. Absent on an FA2-capable arch, configure warns and FA2 is not built
VLLM_CPP_CUTLASS_FETCH OFF FetchContent CUTLASS 4.5.0 if not found locally (~200 MB, needs network)
VLLM_CPP_MARLIN ON Build the vendored Marlin NVFP4 W4A16 MoE GEMM (sm_12xa)
VLLM_CPP_BUILD_TESTS ON Compile and register ctest targets
VLLM_CPP_BUILD_EXAMPLES ON Build the example CLI, server, and bench binaries
VLLM_CPP_BENCH_PROFILE_CONTROL OFF Trace-only profiler replay control (never for production timing builds)

Backend and hardware state

Backend Hardware State
CPU x86-64 and arm64 Correctness / CI reference; at or ahead of llama.cpp on every GGUF axis, with an Arm i8mm quant-GEMM tier
CUDA GB10 / DGX Spark, sm_121a Gate-model correctness passes; 27B at/above vLLM throughput, 35B prefill-pending. The only runtime-gated CUDA target
CUDA Consumer Blackwell, sm_120a Build-supported (compiles, emits real sm_120a code, all fast paths resolve) but not runtime-proven here (no such card)
CUDA Hopper, sm_90a Build-supported; the fast GDN (Triton-AOT) path is build-verified, not runtime-proven here
CUDA Ampere/Ada (sm_80/86/87/89), datacenter Blackwell (sm_100a/103a), sm_110 Build-supported; the fast GDN path is build-verified per-arch on sm_80/86/89/100a (plus FA2 on Ampere, sm_100a NVFP4 GEMM), not runtime-gated here. sm_70/sm_75 unsupported (no bf16 tensor cores)
Metal Apple Silicon Two models run end to end and pass correctness; 18 of 75 ops native. Warm b=1 throughput is 95.9% of MLX-LM, or 97.6% with the optional MLX provider gated to prefill (where we are 1.5% ahead). Indicative
Vulkan Portable GPU Skeleton: 8 ops plus the fusion catalogue run and cross-check against CPU and CUDA. No model runs yet
Intel XPU Intel GPUs Spiked, hardware-blocked
ROCm AMD GPUs W0 skeleton committed (backend, platform, one op); its HIP sources are never compiled — no AMD hardware here. Open for contribution: ROCM.md
ANE Apple Neural Engine Post-parity roadmap

Only GB10 / sm_121a is a runtime-gated CUDA target today. Consumer Blackwell (120a) plus the cross-family targets are build-supported (they compile and emit real machine code, with the fast GDN path build-verified on several) but unproven at runtime here (no such board), and non-Apple / non-NVIDIA backends run a subset of operations. Per-op detail is in the backend matrix.

Quantization formats

Format State
NVFP4 W4A4 / W4A16 Both gate-model paths run on GB10, token-exact. FP4 tactics match vLLM; Marlin NVFP4 W4A16 grouped-MoE is the 35B expert path
compressed-tensors NVFP4A16 (W4A16), dense Correctness-complete via the Marlin weight-only path; speed not yet measured
GGUF F32 / F16 / Q4_0 / Q8_0 / Q3_K / Q4_K / Q5_K / Q6_K Supported. On CPU the six block encodings compute directly on the compressed blocks (VT_GGUF_KEEP_QUANT=0 disables it). GPU builds still expand GGUF weights
FP8 (W8A8) The 35B ModelOpt static per-tensor projection slice is implemented; generic FP8 modes and FP8 KV remain open
MXFP4 / MXFP8 Planned

Environment variables

Runtime knobs (op providers, keep-quant, profiling) are documented in docs/ENVIRONMENT.md.