Skip to content

Commit 581d335

Browse files
committed
Merge Qwen GDN BA projection
Pack the 27B GDN b/a weights into one owner, issue one default BA projection, and teach the fused and unfused gate consumers to read row-strided F32/BF16 views. Keep the exact BF16 rounding and performance gates open. FOLLOWING_AGENTS_PROTOCOL Assisted-by: Codex:GPT-5 [Codex]
1 parent fe2ecc7 commit 581d335

24 files changed

Lines changed: 765 additions & 144 deletions

.agents/coordination.md

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -113,7 +113,7 @@ time owns the GB10. Results without the lock for their entire run are discarded.
113113
| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
114114
|---|---|---|---|---|---|---|---|
115115
| `CLAIM-PR3` | `KERNEL-GDN-AOT-BF16`, `KERNEL-GDN-SCRATCH` | root takeover of stopped `validate_pr3` / `complete_pr3` stream | primary recovery tree `/home/mudler/_git/vllm.cpp-pr3-validate`; 27B default/component integration in `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; DGX evidence `~/work/vllm.cpp-noPy` plus `~/work/vllm.cpp-nvfp4-small-m/debug/gdn-out-bf16-c16-ab-20260712` | recovery branch integrated into `main` by `a767188`; current checkpoint on `codex/nvfp4-small-m` | PR #3 files and rows `KERNEL-GDN-AOT-BF16`+`KERNEL-GDN-SCRATCH`; 27B-only `GdnOutDType` default/f32 override in `qwen3_5.cpp`; ledger/inventory/roadmap evidence. No 35B default change. GPU lock: residual trace/pool classification remains separate from FP4 W3 | `ACTIVE` | 2026-07-13 (vendored BF16 H32/H48 AOT/safety evidence and native 16/16 correctness remain green. Binding `3f256ab` has c16 total at 1.027889× but mean TPOT/ITL at 0.987450×. The dynamic scan ranks packed pure-decode GDN fusion after W3-C removes tactic-selection confounding. All 35B paths keep f32) |
116-
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356` | current checkpoint `codex/nvfp4-small-m` | Online gate/evidence surfaces through the accepted merged-GDN spike. `KERNEL-GEMM-BF16` remains unclaimed while `READY`; implementation, loader/memory repair, exact grid, and 35B performance are excluded until the next claim transition. Any A/B or trace series uses one `flock` | `ACTIVE` | 2026-07-14 (`179a0fc` status `9e0143fa…7b57` proves exact **193 vs 97** BF16 projection calls. [Merged-projection spike](specs/gdn-merged-input-projections.md) accepted/`READY`; BA implementation is the next claim. Binding 55/124; no speed credit) |
116+
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; W1 preflight `~/work/vllm.cpp-gdn-ba/evidence/` | current checkpoint `codex/nvfp4-small-m` | Online-gate execution, including immutable evidence for the now-unclaimed `GATING` W1 BA implementation. qkvz code, loader-memory repair, exact grid, and 35B performance are excluded. Any A/B or trace series uses one `flock`; further `KERNEL-GEMM-BF16` code requires a fresh `ACTIVE` claim | `ACTIVE` | 2026-07-14 (W1 implementation claim released at `GATING`. F32-output merged/split 27B each pass **235/235 + 16/16**, packed replay/memcheck and 35B inertness pass; BF16 output fails **233/235**. Immutable build/trace/component and exact rounding repair remain pending; binding 55/124, no speed credit) |
117117
| `CLAIM-NVFP4-SMALL-M-2` | `KERNEL-GEMM-NVFP4-W4A4` (`W2`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-small-m/b5c6e4fd65cdacea8f378e18ae101ebf521e8f01/w2` plus exact online gate | `codex/nvfp4-small-m` | W2 only: 32 tactics, merged CT gate/up and one-input `SiluAndMulFp4Quant`, with independent fallbacks and full correctness/safety/component/oracle gates | `RELEASED` | 2026-07-12 (implementation/correctness complete; strict old-oracle acceptance failed. The vLLM 0.24.0 ratios are historical diagnostics; W3 remains the trace-driven repair track under the new v0.25.0 denominator) |
118118
| `CLAIM-NVFP4-SMALL-M-3` | `KERNEL-GEMM-NVFP4-W4A4` (`W3-C`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; oracle cache fixture under `tests/fixtures/nvfp4_flashinfer_v025_gb10/`; immutable C3/C3R/corrected component `~/work/vllm.cpp-nvfp4-persistent/d211b8f80fff831a712f0bfafa4f65f1abe1892d/evidence` | `codex/nvfp4-small-m` | W3-C document/import/atomicity, ready-map/lifecycle/5,000-us wiring and corrected same-plan gates only; other levers excluded | `RELEASED` | 2026-07-13 (W3-C reproduction control complete: six-process and corrected 12-leg components use identical 64/64 maps with zero tuning/misses. W3-E strict-fails 39/40 timing + 1/8 memory; no exact grid/35B performance) |
119119
| `CLAIM-NVFP4-SMALL-M-4` | `KERNEL-GEMM-NVFP4-W4A4` (`W3-F`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-device-alpha/7517af4f983fe322ac88ce2d9869e1441b7be3fd/` | `codex/nvfp4-small-m` | W3-F only: model-owned device F32 alpha, tensor op ABI/validation, `VT_FP4_DEVICE_ALPHA=0` fallback, ported tests, safety/model/node-trace and frozen-plan c2/c16 gates per [spike](specs/nvfp4-device-alpha.md). Quant/GDN/attention/host-weight changes excluded | `RELEASED` | 2026-07-13 (local/CUDA/operator/memcheck/model/trace gates pass. Completed 12-leg/612-request component is c2/c16 1.001967×/1.000144× but strict-fails 27/40 timing + 3/8 memory. No speed credit/exact grid/35B performance; the completed scan moves W3-G FA2 decode under CLAIM-SERVE-GATE-1) |
@@ -123,15 +123,16 @@ time owns the GB10. Results without the lock for their entire run are discarded.
123123
`CLAIM-SERVE-GATE-1` owns the binding grid and finalized exact-c2 evidence.
124124
Root `179a0fc`, status `9e0143fa…7b57`, proves the selected **193 vs 97** BF16
125125
projection topology; detailed precursor chronology remains only in the
126-
append-only record. `KERNEL-GEMM-BF16` is unclaimed while `READY`. Its next
127-
legal transition is an explicit BA-only `ACTIVE` claim; qkvz, exact-grid and
128-
35B work remain excluded until their preceding component gates close.
126+
append-only record. `KERNEL-GEMM-BF16` BA-only W1 is implemented/`GATING` and
127+
its implementation claim is released; the serving claim owns immutable
128+
evidence execution only. Exact BF16 rounding, trace and component result remain
129+
open. qkvz, exact-grid and 35B work stay excluded until those gates close.
129130

130131
## Handoff queue
131132

132133
| Priority | Row/block | Dependency | Next handoff | State |
133134
|---|---|---|---|---|
134-
| 1 | `KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE` | `3f256ab` remains 55/124; accepted [spike](specs/gdn-merged-input-projections.md) explains all +96 projection launches without speed credit | claim and gate W1 merged BA; only then claim qkvz, followed conditionally by the exact grid | `READY` |
135+
| 1 | `KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE` | `3f256ab` remains 55/124; W1 F32-output BA code is preflight-correct, BF16 output fails, and no trace/component speed credit exists | repeat the pushed W1 build/model/safety gates, prove 145-vs-193 graph structure, resolve BF16 rounding, then complete c2/c16 before qkvz | `GATING` |
135136
| 2 | `SERVE-ASYNC-LLM` block (`ENG-CORE-BUSY-LOOP`, `SERVE-ASYNC-LLM`, `ENG-ASYNC-SCHED`, `ENG-PRIORITY-SCHED`) | joint spike accepted ([async-serving.md](specs/async-serving.md)); W1/W2/W4 implemented/`GATING`, W3 `READY`; complete control shows W3 is neutral for order-0 speed | retain W3 as unclaimed parity work until order 0 closes; do not fold it into the active kernel claim | `GATING` |
136137
| 3 | `SERVE-E2E-NIGHTLY` | `SERVE-GATE-ONLINE` evidence where benchmarks overlap | write spike and CI/nightly split | `INVENTORIED` |
137138
| 4 | C1 kernel drop-in alignment | accepted kernel-family inventory + [drop-in ABI spike](specs/dropin-kernel-abi.md) | `BACKEND-ABI-VT` W0 is implemented and CPU-green; run cross-build + flock-held GB10 runtime/capture/model/trace handoff, close scalar-forwarder/backend-shim debts, then migrate one family per independently gated checkpoint | `GATING` |

.agents/engine-matrix.md

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -18,7 +18,9 @@ Finalized exact-c2 root `179a0fc` (status `9e0143fa…7b57`) proves matching FP4
1818
tactics and explains all 96 extra BF16 projection launches: qkv+z+b+a versus
1919
vLLM qkvz+ba across 48 GDN layers. The accepted
2020
[merged-projection spike](specs/gdn-merged-input-projections.md) makes
21-
`KERNEL-GEMM-BF16` `READY`; BA is the first component gate and qkvz the second.
21+
`KERNEL-GEMM-BF16` W1 `GATING`. The one-owner F32-output BA merge is
22+
preflight-correct and safety-green; BF16 output fails the token near-tie, while
23+
immutable trace/component evidence remains pending. qkvz stays excluded.
2224
No trace duration earns speed credit. Host PSS/RSS separately retains a
2325
**22.920 GiB** CPU weight mirror plus source mmap residency.
2426

@@ -155,7 +157,7 @@ claims it.
155157
| `SERVE-C-ABI` | Stable LocalAI-style C FFI (17 exported symbols; blocking and nonblocking request handles) | T0 | Original project ABI; pinned vLLM has no C ABI | `include/vllm.h:143,181,207`; `src/capi/vllm_c.cpp:229,264,327,391` | `tests/capi/test_capi.cpp:320,428,505,574,606,640`; `tests/capi/test_dlopen.cpp:77,86`; `tests/capi/c_header_compile.c:1` | `planned: specs/c-api-library.md` | `ANCHOR-BACKFILL` | - |
156158
| `SERVE-CPP-API` | Rich `LLM` and `AsyncLLM` C++ API | T1 | `vllm/entrypoints/llm.py:66,422`; `vllm/v1/engine/async_llm.py:70` | - | - | `planned: specs/cpp-api.md` | `INVENTORIED` | - |
157159
| `SERVE-CLI-BENCH` | Serve and latency/throughput/serve benchmark modes | T0 | `vllm/entrypoints/cli/serve.py:44`; `vllm/entrypoints/cli/benchmark/main.py:29` | separate binaries + explicit scheduler-capacity flags `examples/server/main.cpp:63,96,116,170`; `examples/bench/main.cpp:40,109`; `examples/bench/bench_core.h:96,468` | server help contract `examples/CMakeLists.txt:34`; benchmark `tests/examples/test_bench.cpp:15,48` | `planned: specs/cli-serve-bench.md` | `PARTIAL` | - |
158-
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Existing schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), [driver](../scripts/dgx-online-serving.sh#L14), batch-keyed [validator](../tools/bench/online_gate.py#L99), and fail-closed [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` status `9e0143fa…7b57` proves **12/12** invariant local B=2 ranges, exact oracle topology, matching FP4 tactics, and **193 vs 97** BF16 projection GEMMs. The [merged-projection spike](specs/gdn-merged-input-projections.md) is accepted: BA then qkvz are pending component gates with no speed credit yet. Host memory remains **22.920 GiB** CPU weights plus mmap pages; exact grid and 35B performance remain blocked | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
160+
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Existing schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), [driver](../scripts/dgx-online-serving.sh#L14), batch-keyed [validator](../tools/bench/online_gate.py#L99), and fail-closed [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` status `9e0143fa…7b57` proves exact before-state topology and matching FP4 tactics. W1 BA code is `GATING`: F32-output merged/split 27B each pass **235/235 + 16/16**, packed capture/memcheck and 35B inertness pass, but BF16 output fails **233/235** and immutable 145-vs-193/component evidence is pending. No speed credit. Host memory remains **22.920 GiB** CPU weights plus mmap pages; exact grid and 35B performance remain blocked | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
159161
| `SERVE-E2E-NIGHTLY` | Server conformance and real-model nightly suites for all release gates | T0 | `tests/entrypoints/openai/`; `tests/v1/e2e/`; `.buildkite/test-pipeline.yaml` | current unit/conformance tests only; no scheduled DGX suite | `tests/vllm/entrypoints/openai/test_conformance.cpp:1`; `tests/parity/test_qwen36_paged_engine.cpp:78`; `tests/parity/test_qwen27_paged_engine.cpp:110` | `planned: specs/server-e2e-nightly.md` | `INVENTORIED` | - |
160162
| `SERVE-CLI-CHAT` | Interactive chat and complete commands | T1 | `vllm/entrypoints/cli/main.py:18-34` has no direct chat/complete command at the pin; project extension | - | - | `planned: specs/cli-chat-complete.md` | `INVENTORIED` | - |
161163
| `SERVE-POOLING-ENDPOINTS` | Embeddings, pooling, score, rerank | T2 | `vllm/entrypoints/pooling/embed/api_router.py:25`; `vllm/entrypoints/pooling/scoring/api_router.py:1` | - | - | `planned: specs/pooling-endpoints.md` | `INVENTORIED` | - |

.agents/environment.md

Lines changed: 10 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -80,17 +80,16 @@ inner 4096, state 128; context 262144.
8080

8181
## TODO
8282

83-
- Binding immutable `3f256ab` vLLM v0.25.0 evidence completes the exact
84-
27B cache-off grid at **55/124 axes pass, 69 fail**. The post-W3-I scan is
85-
complete: accepted async root `3812d8` contains all six timing legs and both
86-
shape-neutral traces. ON/OFF total ratio is **1.002153×**, direction-
87-
normalized TPOT/TTFT are **1.010291× / 0.862159×**, and traced GPU time is
88-
**1.002004×**; W3 stays uncredited and leaves the immediate speed path.
89-
Capture exact c2 ours/vLLM kernels and map the RMSNorm/generated partitions
90-
plus resolved FP4 tactics before coding. The independent host-memory repair must remove the measured
91-
**22.920 GiB** persistent CPU weight mirror and overlapping load-time source
92-
pages without increasing final GPU allocations. Close all 69 failed axes
93-
before 35B performance.
83+
- Binding immutable `3f256ab` remains **55/124 axes pass, 69 fail** against
84+
vLLM v0.25.0. Finalized c2 root `179a0fc` already maps the executed path and
85+
selects the complete **193 vs 97** GDN projection mismatch. W1 merged BA is
86+
implemented/`GATING` in `~/work/vllm.cpp-gdn-ba`: production CUTLASS 4.5
87+
build, packed F32/BF16 capture/replay, strict memcheck, merged/split 27B and
88+
inert 35B preflights pass; BF16 projection output fails the token near-tie.
89+
Rebuild the pushed SHA, close exact 145-vs-193 trace, rounding parity and the
90+
c2/c16 component before qkvz. Independently remove the measured **22.920
91+
GiB** host-weight mirror and overlapping source pages. No 35B performance
92+
command runs before all 27B axes pass.
9493
- Keep the existing SGLang v0.5.13 P1 evidence immutable. The distinct
9594
shared-prefix gate pins v0.5.15 `f63458b` and image digest `d0a667e`; its PX1
9695
deterministic 64k/256k harness/counter work is ready after the priority

.agents/feature-matrix.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -108,7 +108,7 @@ and reference-engine performance.
108108

109109
| ID | Block | State | Grounded summary | Detailed evidence / spike |
110110
|---|---|---|---|---|
111-
| `QUANT-CUDA-GATES` | NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 | `DONE` | support/correctness stays closed; performance remains `ACTIVE` at `3f256ab` **55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md) is accepted/`READY` under `KERNEL-GEMM-BF16` | quant matrix §2 + [coverage spike](specs/quantization-coverage.md) |
111+
| `QUANT-CUDA-GATES` | NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 | `DONE` | support/correctness stays closed; performance remains `ACTIVE` at `3f256ab` **55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md) has BA-only W1 code `GATING`: F32 output is token-correct, BF16 output is not, and structure/performance remain pending | quant matrix §2 + [coverage spike](specs/quantization-coverage.md) |
112112
| `QUANT-GGUF` | llama.cpp encodings and output presets | `PARTIAL` | F32/Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K materialize; F16 was corrected to reader-only; CPU threadpool/chunked dispatch is correctness-gated but its B4 speed/RSS checkpoint is pending; no direct compute-in-quant or llama.cpp speed parity | quant matrix §1 |
113113
| `QUANT-VLLM-BREADTH` | generic FP8/MX, AWQ/GPTQ, CT integer, vendor methods, KV | `PARTIAL` | gate-specific implementations exist; generic dispatch/modes remain inventoried | quant matrix §§2-3 |
114114
| `QUANT-MLX` | affine Q2-8, MXFP4/MXFP8/NVFP4, QQ, mixed recipes/imports | `INVENTORIED` | required for Apple backend; no MLX runtime yet | quant matrix §4 |

0 commit comments

Comments
 (0)