Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .agents/NOW.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,8 @@ checkpoint on `upstream/main` at `59674cf1d`.
| Kimi-Linear-48B (KDA+NoPE-MLA+MoE) | **Full-model GB10 e2e RUNS** (bf16-resident §13): CPU+CUDA 13/13·656, no OOM. **Token gate NEAR-TIE 106/128** (6/8 token-exact) | device GDN/MLA islands + bf16 stream; 1.59 tok/s; default OFF |
| 35B fresh grid | **BOUND** @`1ea26427`: tput 0.93-1.03x, c16 0.93x. INTAKE + Option A both **RESOLVED NEGATIVE** (H2D-out-of-capture tput WASH) | Real lever left: prefill glue (task #61) |
| Qwen3.5-4B revalidation | 0.9971x @`59674cf1` (#35); TTFT/PSS pass, TPOT/ITL open | `docs/bench-evidence/` |
| MXFP4 parity | **TERMINAL (`QUANT-CT-MXFP4-FINAL-STACK`)**: c1 1.020x PASS + mem 2.63x; c2-c8 0.962-0.969 GPU-intrinsic. Both last levers exhausted: num_splits cap `VT_FA2_NSPLITS_CAP` gated-OFF (c1-only, green, 32B strict char-identical); glue folds via `vt::FusedChain`, residual out-of-catalog Inductor GEMM-epilogue fusion (#46). `VT_MARLIN_DENSE` on | Record; branch not merged |
| RPi5 A76 CPU | **R5 ASSEMBLY GREEN; llama floor MEASURED/NOT MET:** exact AAPCS64 beats SDOT, but vllm.cpp is 0.461x prefill / 0.653x decode+E2E vs llama.cpp; RSS 24.2% lower | W6: profile the 2.17x prefill / 1.53x decode gap, starting with BF16 GEMM; M1/T4 and concurrency remain |
| MXFP4 parity | **TERMINAL:** c1 1.020x pass; c2-c8 0.962-0.969 GPU-intrinsic; final cap/glue levers exhausted | Record; branch unmerged |
| ROW-SERVE-ASYNC-DENSE-MIRROR | **LANDED+dgx-VERIFIED** (`f9c969ae`): #31 async mirror on classic dense Qwen3; gate RED→GREEN, SACRED 184/184 | Residual: sibling scope one-liner |

In-flight branches (default-OFF, not pushed): `laguna-fp4proj-prod` (fp4),
Expand Down
4 changes: 2 additions & 2 deletions .agents/backend-matrix.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/feature-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -274,7 +274,7 @@ evidence.
|---|---|---|---|---|
| `BACKEND-CUDA-SM121` | GB10/sm121a | `PARTIAL` | gate workload built, traced, token/perf gated; full component-family coverage is open | [backend row](backend-matrix.md#cuda-target-rows) |
| `BACKEND-CUDA-OTHER` | vLLM sm70/75/80/86/87/89/90/100/101/103/110/120 targets | `ACTIVE` | 9 CUDA arches build-supported (sm80/86/87/89/90a/100a/103a/110/120a, single-arch portable-kernels-only, `-Werror` clean, SASS emitted); sm70/75/101 not build-supported; no non-121a target is runtime-validated | [backend matrix](backend-matrix.md), [CUDA inventory](specs/cuda-architecture-inventory.md) |
| `BACKEND-CPU` | production CPU | `PARTIAL` | persistent threadpool + chunked GEMM/row dispatch is 1/3/20-thread bit-identical and TSAN-clean; idle-host performance/RSS gate and compute-in-quant remain open | [backend matrix](backend-matrix.md) |
| `BACKEND-CPU` | production CPU | `PARTIAL` | shared CPU path is correctness-gated and the 20-core Arm/i8mm Qwen3.5-2B single-stream llama.cpp floor is closed; server concurrency is open. Raspberry Pi 5 Cortex-A76 R4-R5 is green: QEMU-built artifact, exact output, and an AAPCS64 Q8 leaf that beats compiler SDOT 3.66-5.08% on M1/T1 and M128. Its separate same-file llama.cpp floor is now measured/NOT MET on speed: vllm.cpp is 0.461x prefill and 0.653x decode/E2E, while using 24.2% less RSS; exact-prompt 64-token output matches. M1/T4 is −2.43%; BF16 GEMM, speed closure and concurrency remain open | [backend matrix](backend-matrix.md), [A76 Q8 dot](specs/cpu-a76-q8-dot.md), [assembly evidence](../docs/bench-evidence/rpi5-a76-q8-dot-20260806.md), [llama.cpp evidence](../docs/bench-evidence/rpi5-a76-llamacpp-20260806.md) |
| `BACKEND-ROCM` | ROCm | `INVENTORIED` | source/dispatch spike required; no "one flag" support claim | [backend matrix](backend-matrix.md) |
| `BACKEND-MLX` | Apple Metal through MLX | `ACTIVE` | Metal/MLX skeleton ACTIVE: two models (OPT-125m, Qwen3-0.6B) run e2e + pass correctness, native-MSL GEMM, batched command buffers 1.50x (compute-bound at 98%+); optional MLX GEMM provider | [backend matrix](backend-matrix.md) |
| `BACKEND-VULKAN` | Vulkan | `INVENTORIED` | runtime absent | [backend matrix](backend-matrix.md) |
Expand Down
8 changes: 5 additions & 3 deletions .agents/kernel-matrix.md

Large diffs are not rendered by default.

3 changes: 3 additions & 0 deletions .agents/parity-ledger.md
Original file line number Diff line number Diff line change
Expand Up @@ -914,3 +914,6 @@ Columns:
| 2026-08-06 (`row/KERNEL-FA2-GQA-SWAP`; `CLAIM-KERNEL-FA2-GQA-SWAP`; kernel `KERNEL-ATTN-FA2`; gated default-OFF, lifecycle unchanged) | Ports vLLM's FA2 `seqlenq_ngroups_swapped` decode optimization into the d128 varlen decode launcher (`LaunchDecodeVarlenFA2Bf16`, gate `VT_FA2_DECODE_GQA_SWAP`): the Qwen3-dense decode grid becomes `(batch, kv_heads)` not `(batch, hq)` — the ngroups query heads pack into seqlen_q, KV read once/group, presented WITHOUT a materialized transpose via kv-major-group-minor strides (a 1:1 mirror of the already-shipped d256 `LaunchDecodeFA2Bf16` swap). OFF path byte-identical to the prior plain-varlen reduction; ON is non-byte-exact only when num_splits>1 (split reduction order → near-tie, toward vLLM's own numerics). | Mirrors `flash-attention @ 2c839c33` `mha_fwd_kvcache` seqlenq_ngroups_swapped + `set_params_splitkv` and vLLM v0.25.0 `flash_attn.py flash_attn_varlen_func` decode (#47 measured vLLM's swapped grid `(1,6,16)` = batch×kv_heads vs ours `(1,3,64)` = batch×query_heads). The vendored `flash_fwd_kernel.h` `get_lse_tile`/combine already honor the flag in both the num_splits==1 direct-write and >1 combine paths (the d256 arm is the proof). | GB10 sm_121a CUDA 13.0: op RED-first test 280/280 (both GQA ratios × batch{1,2,4,8} × short+long ctx; `swap_launches==1` proves the grid engaged; swap-vs-plain near-tie; MHA-inert) — RED proven (wrong swapped stride → 26,528 violations); full binary 28/28·454,679 no regression; compute-sanitizer 0-err/0-leak; #44 MXFP4 e2e smoke swap-ON 3/3 deterministic TOKEN-EXACT + coherent, byte-identical to swap-OFF. `benchmark_binding=false` (c1-c8 x3 re-bench + default flip = recorded next step; #47 projects flash ~28%@c2 / ~55%@c8 of the gap). |
| 2026-08-06 (`row/KERNEL-MARLIN-DENSE-PORT`; `CLAIM-KERNEL-MARLIN-DENSE-PORT`; kernel `KERNEL-GEMM-MARLIN-W4A16`; gated default-OFF, lifecycle unchanged) | Vendors vLLM's OWN dense marlin W4A16 GEMM as a new `vt::MarlinDenseGemm` op (`VT_MARLIN_DENSE`, default OFF) and routes the E=1 dense NVFP4/MXFP4 projections (`dense_nvfp4_gemm.h` `MatmulNvfp4MarlinD`/`MatmulMxfp4W4A16D`/`GateUpFusedMarlinD`) through it. The dense kernel is direct-A + tile-per-CTA with vLLM's OWN dense fp32-C_tmp reduce, so at M<=8 it runs the sms-wide (48-CTA) grid WITHOUT the one-bf16-ULP shift the `VT_MARLIN_E1_PAR1` MoE-route par-regroup costs (#54: that ULP flips a strict 32B-NVFP4A16 token). Reuses the EXISTING marlin resident + workspace (same `marlin_permute` repack for dense and MoE — confirmed, no shim); rank-2 operand views, no moe_align gather. | 1:1 lift of vLLM @ `555967922` `csrc/libtorch_stable/quantization/marlin/`: `marlin.cu:326-541` (`marlin::marlin_mm` + config helpers) → `marlin_mm_dense.cu`; the torch::stable `marlin_gemm` wrapper (`:545-894`) → torch-free `cuda_marlin_dense.cu` launcher (mirrors `cuda_moe_marlin.cu`, dense c_tmp sizing `:713-716`); `kernel.h`/`marlin_template.h:1-2081` verbatim (the DENSE kernel — DISTINCT from the moe one, but SAME 12-param `Marlin<>` template so the generated `kernel_selector.h`+`sm80_*.cu` instantiation set is shared, namespace `marlin` from the local kernel.h). Shared `marlin.cuh`/`marlin_dtypes.cuh`/`dequant.h`/`marlin_mma.h` diff-verified byte-identical. Forced-Marlin a16 selection `kernels/linear/__init__.py:879-881`. | CPU `-fsyntax-only` CLEAN (`ops.cpp` + the `VT_MARLIN_NVFP4` routing header). GPU compile: all 3 new dense `.cu` compile CLEAN on dgx GB10 sm_121a under exact production flags (`-Werror=all-warnings`, `-static-global-template-stub=false`, `--generate-code=…sm_121a`). RED-first unit battery WRITTEN (`test_ops_moe_grouped.cpp`: NVFP4+MXFP4, M=1..8 × 3 shapes, dense-vs-CPU-ref AND dense-vs-grouped-route, row-shifted stride RED-injection). `benchmark_binding=false`; GPU EXEC gates (unit run + strict token battery dense-ON vs oracle incl. 32B-NVFP4A16:344 + launch-counter + nsys 48-CTA + binding c1..c8 x3) are the scoped dgx follow-up; default stays OFF until the strict battery proves oracle byte-match and the binding beats the MoE route (state `KERNEL-MARLIN-DENSE-PORT`). |
| 2026-08-06 (`row/H3-FP4-SPEED`; `ROAD-V1-H3`; model `MODEL-DIFFUSION-minimax-h3-mini-max-h3-dit`; lifecycle unchanged) | **MiniMax-H3 W-FP4a — fp4-RESIDENT NVFP4 routing for the device DiT forward (NO new quant code).** Until now both NVFP4 loaders dequantized the packed FP4 projections to bf16 and the device forward ran `vt::MatmulBT`, so the sm_121a FP4 tensor-core route never ran for H3. Adds `Nvfp4Weight` carriers to `MiniMaxH3DitBlockWeights`/`MiniMaxH3DitWeights`, a fp4-resident streamer `StreamMiniMaxH3Nvfp4ToDeviceFp4` (keeps the compressed-tensors triple host-resident; the shared dispatcher uploads+repacks lazily then frees the fp4 originals, so peak device memory is ~1/4 of the bf16 arm), and a `LinearDev` dispatch that routes a non-Empty fp4 projection through `dense_nvfp4::MatmulNvfp4W4A16D`. | The routing is vLLM's OWN forced-Marlin-for-a16 selection: the checkpoint is weight-only NVFP4 (no `input_activations`, `IsTrueW4A4()==false`), so `kernels/linear/__init__.py:879-881` forces the Marlin W4A16 kernel, mirrored by `include/vllm/model_executor/models/dense_nvfp4_gemm.h:12-22,505-549` (`MatmulNvfp4W4A16D` -> single-expert `vt::MoeGroupedGemmNvfp4Marlin`, the SAME kernel Laguna routed-experts + dense Qwen3-32B NVFP4 use). Not cutlass-fp4/W4A4 (needs fp4 activations, private to `qwen3_5.cpp`). fc1 is already merged `[gate;up]` -> one W4A16 GEMM + `vt::SiluAndMul`. | **CPU-GATED (wiring), verified**: `test_minimax_h3` 62/62 cases / 30039 assertions, 0 failed. The synthetic-NVFP4 case streams the fp4 twin, asserts the loader kept the projections PACKED (fp4 slot set / bf16 slot Empty, and the inverse for the bf16 loader), runs fp4 + bf16 device forwards on the SAME file, asserts the W4A16 dispatcher executed ALL 11 quantized GEMMs (the `Nvfp4W4A16Stats` this-path-ran counter), and bounds fp4-vs-bf16 <= 2e-3. On CPU the dispatcher has no Marlin op so it falls to the bf16 arm's own dequant+matmul (hence a WIRING gate here); the Marlin kernel numerics are CUDA-gated independently by `test_ops_nvfp4_matmul` / `test_linear_method` (2e-3 f32-out / 8e-3 bf16-out vs a bf16 reference). `benchmark_binding=false`. PENDING: GB10 CUDA build + the fp4-vs-bf16 numeric delta and steady per-step timing (disk/build window); real-checkpoint t2va e2e DISK-BLOCKED (~41 GB working set). Comparability: vLLM-Omni serves NO quantized H3 (BF16-only in practice; source-audited `a4ea67a2`, spec §8.3) -> HW/loader-forced-indirect. |
| 2026-08-06 (`row/BACKEND-CPU`; `BACKEND-CPU` R1; PR #65; lifecycle remains `PARTIAL`) | Adds `vllm-cpu-kernel-bench`, a developer-only vt-op benchmark substrate: deterministic quant-GEMM fixtures, calibrated batched timing, cache-pressure profiles, affinity, JSON, checksums, system metadata, and grouped generic/Cortex-A76 `perf_event_open` counters with explicit multiplex/unsupported status. Production dispatch and numerics are unchanged. | No vLLM behavior counterpart; vLLM remains the x86 semantic oracle and llama.cpp `237ad9b96` remains the Pi performance floor. The quant fixture invokes the existing `vt::MatmulBTQuant` contract unchanged. | **CPU-GATED, `benchmark_binding=false`.** GCC 15.2 `-Wall -Wextra -Werror` build; clang-format clean; `test_cpu_kernel_bench_cli` deterministic JSON schema/checksum + invalid-input + structured-counter cases; direct 1/4-thread x86 runs and real generic PMU counts. X86 timings are non-binding tool validation. Pi PMU execution, model correctness, throughput and memory all remain `PENDING`. |
| 2026-08-06 (`KERNEL-CPU-A76-Q8-DOT` R4-R5; `CLAIM-KERNEL-CPU-A76-Q8-DOT`; physical RPi5 Cortex-A76; closing commit: this checkpoint) | Adds an exact-order ACLE SDOT control and an original AAPCS64 two-block Q8_0×Q8_0 leaf behind Linux DotProd/MIDR dispatch. `auto` selects assembly only on Cortex-A76+DotProd; x86, non-DotProd and other Arm CPUs retain portable dispatch. Explicit `portable`/`sdot`/`a76-asm` same-binary controls remain. ARM64 builds/tests locally through buildx/QEMU; Pi is execution-only. | The integer structure is informed by llama.cpp `237ad9b96` `ggml/src/ggml-cpu/arch/arm/quants.c:1076-1160`, while this port deliberately retains the local portable function's per-block f32 reduction order. vLLM `555967922` supplies Qwen3.5 semantics, not a corresponding CPU microkernel. Local anchors: `src/vt/cpu/cpu_quant_dot_{sdot.cpp,a76.S}`, Q8 dispatch in `cpu_quant_dot.cpp`, direct tests in `tests/vt/test_ops_quant_dot.cpp`, and [immutable evidence](../docs/bench-evidence/rpi5-a76-q8-dot-20260806.md). | **PASS for the compiler-gap/component gate; row `GATING`, `benchmark_binding=true`.** Final binaries `9eb57cf...`/`a94dad30...`; QEMU focused suite 20/20, 150258 assertions; physical-Pi checksums exact. Assembly vs compiler SDOT wall/cycles/instructions: M1/T1 +3.66%/+3.17%/+10.10%, M128/T1 +5.08%/+4.61%/+10.24%, M128/T4 +3.69%/+3.69%/+9.74%. M1/T4 is an explicit −2.43% wall/−4.32% cycles residual despite 8.77% fewer instructions. All 64 Qwen tokens equal the x86 golden in all nine runs; median assembly vs SDOT TTFT −1.55%, TPOT −0.05% neutral, E2E −0.13%. Disassembly proves GCC's framed dependent one-block loop versus the stack-free independent two-block schedule. Same-file Pi llama.cpp, peak memory and concurrency remain `PENDING`; no competitor-floor binding is claimed. |
| 2026-08-06 (`KERNEL-CPU-A76-Q8-DOT` R6 competitor checkpoint; physical RPi5 Cortex-A76; lifecycle remains `GATING`; closing commit: this checkpoint) | Measures the separate four-core A76 same-file llama.cpp floor after the assembly leaf became default. No production code changes. The vllm.cpp nominal p16 request measures 17 input tokens, so the binding competitor uses pp17/tg64/pp17+tg64. A same-text CLI arm verifies 64-token greedy output equality. | Official llama.cpp tag b9892 `ee445f93d` reconstructed under QEMU because historical recorded fork object `237ad9b96` is unavailable; exact recorded anchors match (`quants.c:400`, `arch/arm/quants.c:1076`, `repack.cpp:2725`, `qwen35.cpp`). Local evidence: [Pi competitor record](../docs/bench-evidence/rpi5-a76-llamacpp-20260806.md). | **CORRECTNESS PASS, PERFORMANCE NOT MET, `benchmark_binding=true`.** Three clean unthrottled vllm.cpp reps: prefill 12.81 tok/s, decode 2.55 tok/s, output-equivalent E2E 2.46 tok/s, E2E 26,018.39 ms. llama.cpp three-sample p17/tg64/combined: 27.77 / 3.91 / 3.77 tok/s, E2E 16,998.49 ms. vllm.cpp ratios 0.461x prefill / 0.653x decode+E2E; peak RSS wins 2.841 vs 3.747 GiB (24.2% less). Same-text normalized output SHA `a5a630d7...` equal; all vllm performance tokens retain golden SHA `0ec98e...`. Intrusive 50 ms forked sampler run VOID; accepted timing has no sampler, RSS sampled separately at 1 Hz. Next lever: fresh both-engine profile, then BF16 GEMM; M1/T4 and concurrency remain. |
Loading