You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Pack the 27B GDN b/a weights into one owner, issue one default BA projection, and teach the fused and unfused gate consumers to read row-strided F32/BF16 views. Keep the exact BF16 rounding and performance gates open.
FOLLOWING_AGENTS_PROTOCOL
Assisted-by: Codex:GPT-5 [Codex]
Copy file name to clipboardExpand all lines: .agents/coordination.md
+6-5Lines changed: 6 additions & 5 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -113,7 +113,7 @@ time owns the GB10. Results without the lock for their entire run are discarded.
113
113
| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
114
114
|---|---|---|---|---|---|---|---|
115
115
| `CLAIM-PR3` | `KERNEL-GDN-AOT-BF16`, `KERNEL-GDN-SCRATCH` | root takeover of stopped `validate_pr3` / `complete_pr3` stream | primary recovery tree `/home/mudler/_git/vllm.cpp-pr3-validate`; 27B default/component integration in `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; DGX evidence `~/work/vllm.cpp-noPy` plus `~/work/vllm.cpp-nvfp4-small-m/debug/gdn-out-bf16-c16-ab-20260712` | recovery branch integrated into `main` by `a767188`; current checkpoint on `codex/nvfp4-small-m` | PR #3 files and rows `KERNEL-GDN-AOT-BF16`+`KERNEL-GDN-SCRATCH`; 27B-only `GdnOutDType` default/f32 override in `qwen3_5.cpp`; ledger/inventory/roadmap evidence. No 35B default change. GPU lock: residual trace/pool classification remains separate from FP4 W3 | `ACTIVE` | 2026-07-13 (vendored BF16 H32/H48 AOT/safety evidence and native 16/16 correctness remain green. Binding `3f256ab` has c16 total at 1.027889× but mean TPOT/ITL at 0.987450×. The dynamic scan ranks packed pure-decode GDN fusion after W3-C removes tactic-selection confounding. All 35B paths keep f32) |
116
-
|`CLAIM-SERVE-GATE-1`|`SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2`| root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`| current checkpoint `codex/nvfp4-small-m`| Online gate/evidence surfaces through the accepted merged-GDN spike. `KERNEL-GEMM-BF16` remains unclaimed while `READY`; implementation, loader/memory repair, exact grid, and 35B performance are excluded until the next claim transition. Any A/B or trace series uses one `flock`|`ACTIVE`| 2026-07-14 (`179a0fc` status `9e0143fa…7b57` proves exact **193 vs 97** BF16 projection calls. [Merged-projection spike](specs/gdn-merged-input-projections.md) accepted/`READY`; BA implementation is the next claim. Binding 55/124; no speed credit) |
116
+
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; W1 preflight `~/work/vllm.cpp-gdn-ba/evidence/` | current checkpoint `codex/nvfp4-small-m` | Online-gate execution, including immutable evidence for the now-unclaimed `GATING` W1 BA implementation. qkvz code, loader-memory repair, exact grid, and 35B performance are excluded. Any A/B or trace series uses one `flock`; further `KERNEL-GEMM-BF16` code requires a fresh `ACTIVE` claim | `ACTIVE` | 2026-07-14 (W1 implementation claim released at `GATING`. F32-output merged/split 27B each pass **235/235 + 16/16**, packed replay/memcheck and 35B inertness pass; BF16 output fails **233/235**. Immutable build/trace/component and exact rounding repair remain pending; binding 55/124, no speed credit) |
117
117
|`CLAIM-NVFP4-SMALL-M-2`|`KERNEL-GEMM-NVFP4-W4A4` (`W2`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-small-m/b5c6e4fd65cdacea8f378e18ae101ebf521e8f01/w2` plus exact online gate |`codex/nvfp4-small-m`| W2 only: 32 tactics, merged CT gate/up and one-input `SiluAndMulFp4Quant`, with independent fallbacks and full correctness/safety/component/oracle gates |`RELEASED`| 2026-07-12 (implementation/correctness complete; strict old-oracle acceptance failed. The vLLM 0.24.0 ratios are historical diagnostics; W3 remains the trace-driven repair track under the new v0.25.0 denominator) |
118
118
|`CLAIM-NVFP4-SMALL-M-3`|`KERNEL-GEMM-NVFP4-W4A4` (`W3-C`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; oracle cache fixture under `tests/fixtures/nvfp4_flashinfer_v025_gb10/`; immutable C3/C3R/corrected component `~/work/vllm.cpp-nvfp4-persistent/d211b8f80fff831a712f0bfafa4f65f1abe1892d/evidence`|`codex/nvfp4-small-m`| W3-C document/import/atomicity, ready-map/lifecycle/5,000-us wiring and corrected same-plan gates only; other levers excluded |`RELEASED`| 2026-07-13 (W3-C reproduction control complete: six-process and corrected 12-leg components use identical 64/64 maps with zero tuning/misses. W3-E strict-fails 39/40 timing + 1/8 memory; no exact grid/35B performance) |
119
119
|`CLAIM-NVFP4-SMALL-M-4`|`KERNEL-GEMM-NVFP4-W4A4` (`W3-F`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-device-alpha/7517af4f983fe322ac88ce2d9869e1441b7be3fd/`|`codex/nvfp4-small-m`| W3-F only: model-owned device F32 alpha, tensor op ABI/validation, `VT_FP4_DEVICE_ALPHA=0` fallback, ported tests, safety/model/node-trace and frozen-plan c2/c16 gates per [spike](specs/nvfp4-device-alpha.md). Quant/GDN/attention/host-weight changes excluded |`RELEASED`| 2026-07-13 (local/CUDA/operator/memcheck/model/trace gates pass. Completed 12-leg/612-request component is c2/c16 1.001967×/1.000144× but strict-fails 27/40 timing + 3/8 memory. No speed credit/exact grid/35B performance; the completed scan moves W3-G FA2 decode under CLAIM-SERVE-GATE-1) |
@@ -123,15 +123,16 @@ time owns the GB10. Results without the lock for their entire run are discarded.
123
123
`CLAIM-SERVE-GATE-1` owns the binding grid and finalized exact-c2 evidence.
124
124
Root `179a0fc`, status `9e0143fa…7b57`, proves the selected **193 vs 97** BF16
125
125
projection topology; detailed precursor chronology remains only in the
126
-
append-only record. `KERNEL-GEMM-BF16` is unclaimed while `READY`. Its next
127
-
legal transition is an explicit BA-only `ACTIVE` claim; qkvz, exact-grid and
128
-
35B work remain excluded until their preceding component gates close.
126
+
append-only record. `KERNEL-GEMM-BF16` BA-only W1 is implemented/`GATING` and
127
+
its implementation claim is released; the serving claim owns immutable
128
+
evidence execution only. Exact BF16 rounding, trace and component result remain
129
+
open. qkvz, exact-grid and 35B work stay excluded until those gates close.
129
130
130
131
## Handoff queue
131
132
132
133
| Priority | Row/block | Dependency | Next handoff | State |
133
134
|---|---|---|---|---|
134
-
| 1 |`KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE`|`3f256ab` remains 55/124; accepted [spike](specs/gdn-merged-input-projections.md) explains all +96 projection launches without speed credit | claim and gate W1 merged BA; only then claim qkvz, followed conditionally by the exact grid |`READY`|
135
+
| 1 |`KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE`|`3f256ab` remains 55/124; W1 F32-output BA code is preflight-correct, BF16 output fails, and no trace/component speed credit exists | repeat the pushed W1 build/model/safety gates, prove 145-vs-193 graph structure, resolve BF16 rounding, then complete c2/c16 before qkvz |`GATING`|
135
136
| 2 |`SERVE-ASYNC-LLM` block (`ENG-CORE-BUSY-LOOP`, `SERVE-ASYNC-LLM`, `ENG-ASYNC-SCHED`, `ENG-PRIORITY-SCHED`) | joint spike accepted ([async-serving.md](specs/async-serving.md)); W1/W2/W4 implemented/`GATING`, W3 `READY`; complete control shows W3 is neutral for order-0 speed | retain W3 as unclaimed parity work until order 0 closes; do not fold it into the active kernel claim |`GATING`|
136
137
| 3 |`SERVE-E2E-NIGHTLY`|`SERVE-GATE-ONLINE` evidence where benchmarks overlap | write spike and CI/nightly split |`INVENTORIED`|
137
138
| 4 | C1 kernel drop-in alignment | accepted kernel-family inventory + [drop-in ABI spike](specs/dropin-kernel-abi.md)|`BACKEND-ABI-VT` W0 is implemented and CPU-green; run cross-build + flock-held GB10 runtime/capture/model/trace handoff, close scalar-forwarder/backend-shim debts, then migrate one family per independently gated checkpoint |`GATING`|
No trace duration earns speed credit. Host PSS/RSS separately retains a
23
25
**22.920 GiB** CPU weight mirror plus source mmap residency.
24
26
@@ -155,7 +157,7 @@ claims it.
155
157
|`SERVE-C-ABI`| Stable LocalAI-style C FFI (17 exported symbols; blocking and nonblocking request handles) | T0 | Original project ABI; pinned vLLM has no C ABI |`include/vllm.h:143,181,207`; `src/capi/vllm_c.cpp:229,264,327,391`|`tests/capi/test_capi.cpp:320,428,505,574,606,640`; `tests/capi/test_dlopen.cpp:77,86`; `tests/capi/c_header_compile.c:1`|`planned: specs/c-api-library.md`|`ANCHOR-BACKFILL`| - |
156
158
|`SERVE-CPP-API`| Rich `LLM` and `AsyncLLM` C++ API | T1 |`vllm/entrypoints/llm.py:66,422`; `vllm/v1/engine/async_llm.py:70`| - | - |`planned: specs/cpp-api.md`|`INVENTORIED`| - |
157
159
|`SERVE-CLI-BENCH`| Serve and latency/throughput/serve benchmark modes | T0 |`vllm/entrypoints/cli/serve.py:44`; `vllm/entrypoints/cli/benchmark/main.py:29`| separate binaries + explicit scheduler-capacity flags `examples/server/main.cpp:63,96,116,170`; `examples/bench/main.cpp:40,109`; `examples/bench/bench_core.h:96,468`| server help contract `examples/CMakeLists.txt:34`; benchmark `tests/examples/test_bench.cpp:15,48`|`planned: specs/cli-serve-bench.md`|`PARTIAL`| - |
158
-
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Existing schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), [driver](../scripts/dgx-online-serving.sh#L14), batch-keyed [validator](../tools/bench/online_gate.py#L99), and fail-closed [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` status `9e0143fa…7b57` proves **12/12** invariant local B=2 ranges, exact oracle topology, matching FP4 tactics, and **193 vs 97** BF16 projection GEMMs. The [merged-projection spike](specs/gdn-merged-input-projections.md) is accepted: BA then qkvz are pending component gates with no speed credit yet. Host memory remains **22.920 GiB** CPU weights plus mmap pages; exact grid and 35B performance remain blocked | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
160
+
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Existing schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), [driver](../scripts/dgx-online-serving.sh#L14), batch-keyed [validator](../tools/bench/online_gate.py#L99), and fail-closed [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` status `9e0143fa…7b57` proves exact before-state topology and matching FP4 tactics. W1 BA code is `GATING`: F32-output merged/split 27B each pass **235/235 + 16/16**, packed capture/memcheck and 35B inertness pass, but BF16 output fails **233/235** and immutable 145-vs-193/component evidence is pending. No speed credit. Host memory remains **22.920 GiB** CPU weights plus mmap pages; exact grid and 35B performance remain blocked | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
159
161
|`SERVE-E2E-NIGHTLY`| Server conformance and real-model nightly suites for all release gates | T0 |`tests/entrypoints/openai/`; `tests/v1/e2e/`; `.buildkite/test-pipeline.yaml`| current unit/conformance tests only; no scheduled DGX suite |`tests/vllm/entrypoints/openai/test_conformance.cpp:1`; `tests/parity/test_qwen36_paged_engine.cpp:78`; `tests/parity/test_qwen27_paged_engine.cpp:110`|`planned: specs/server-e2e-nightly.md`|`INVENTORIED`| - |
160
162
|`SERVE-CLI-CHAT`| Interactive chat and complete commands | T1 |`vllm/entrypoints/cli/main.py:18-34` has no direct chat/complete command at the pin; project extension | - | - |`planned: specs/cli-chat-complete.md`|`INVENTORIED`| - |
Copy file name to clipboardExpand all lines: .agents/feature-matrix.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -108,7 +108,7 @@ and reference-engine performance.
108
108
109
109
| ID | Block | State | Grounded summary | Detailed evidence / spike |
110
110
|---|---|---|---|---|
111
-
|`QUANT-CUDA-GATES`| NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 |`DONE`| support/correctness stays closed; performance remains `ACTIVE` at `3f256ab`**55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md)is accepted/`READY` under `KERNEL-GEMM-BF16`| quant matrix §2 + [coverage spike](specs/quantization-coverage.md)|
111
+
|`QUANT-CUDA-GATES`| NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 |`DONE`| support/correctness stays closed; performance remains `ACTIVE` at `3f256ab`**55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md)has BA-only W1 code `GATING`: F32 output is token-correct, BF16 output is not, and structure/performance remain pending| quant matrix §2 + [coverage spike](specs/quantization-coverage.md)|
112
112
|`QUANT-GGUF`| llama.cpp encodings and output presets |`PARTIAL`| F32/Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K materialize; F16 was corrected to reader-only; CPU threadpool/chunked dispatch is correctness-gated but its B4 speed/RSS checkpoint is pending; no direct compute-in-quant or llama.cpp speed parity | quant matrix §1 |
0 commit comments