Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/parity-ledger.md
Original file line number Diff line number Diff line change
Expand Up @@ -426,7 +426,7 @@ Columns:
| 2026-07-14 (`SERVE-GATE-ONLINE` third async-credit execution + direct-script bootstrap repair; `FAILED / VOID`) | Executes clean `b8681ac` through all six explicit ON/OFF timing legs under one uncontended lock. The first required Torch trace then fails before profiler startup because an absolute script invocation cannot import repository-local `tools`; neither trace or completion marker exists, so all timings remain diagnostic-only. The repair bootstraps the repository root only for direct script execution and adds an outside-repository absolute-path regression. Production inference is unchanged; live surfaces replace the preceding async narratives with this checkpoint. | Root `~/work/vllm-async-credit/b8681ac80b3f84af71955cf3a20cece2a118ea1f`; series/raw-set/log-set/corpus-set SHA `e8c7a4b7…86b0` / `65bff32f…6e51` / `79fd3836…f9c8` / `9fb8027a…d30`. Anchors: repaired [profiler](../tools/bench/profile_vllm_online_gate.py#L21), [absolute-path contract](../tests/tools/test_online_gate_trace.py#L29), [async spike](specs/async-serving.md), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Six × **6/6** requests pass. Provisional ON/OFF medians are **160.798982 / 160.582996 total tok/s = 1.001345×**, TPOT **106.279834 / 107.333011 ms = 1.009909×**, and TTFT **812.002454 / 696.263685 ms = 0.857465×**; total CV ≤0.137978%. The whole series is **VOID** and earns no credit because both traces/metadata/summaries/marker are absent. Cleanup returns GPU/lock/port idle; focused contracts pass **6/6**. Binding remains **55/124**, W3 stays unowned `READY`, and a fresh commit/root must repeat every timing and trace arm. |
| 2026-07-14 (`SERVE-GATE-ONLINE` fourth async-credit execution; six timings + ON trace, `FAILED / VOID`) | Executes clean `9b1774c` through all six timing legs and a complete explicit-ON Torch trace under one lock. The driver then applies the accepted 48-prompt H1d `--model-key 27` annotation-count contract to the six-prompt c2 diagnostic; the valid trace has 1,539 rather than 1,588 generation annotations, so summary fails closed before OFF. Shape-neutral re-read succeeds; the corrected recipe omits the model key only for this c2 diagnostic. Production inference and H1d validation are unchanged; live surfaces replace prior async narratives with this checkpoint. | Root `~/work/vllm-async-credit/9b1774c014880a0039545ea1be0fa01426cbd900`; series/raw-set/log-set/trace-set SHA `0d204e91…0310` / `9ce64024…4bc` / `6956cb19…7cf0` / `acb402a9…d419`; ON selected trace SHA `ad071f36…bf1`. Anchors: [profiler](../tools/bench/profile_vllm_online_gate.py#L21), shape-neutral [summarizer](../tools/bench/summarize_torch_kernels.py#L23), [async spike](specs/async-serving.md), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Six × **6/6** timing requests pass. Provisional ON/OFF medians are **160.287860 / 160.213485 total tok/s = 1.000464×**, TPOT **106.648618 / 107.594484 ms = 1.008869×**, and TTFT **809.941298 / 697.928448 ms = 0.861703×**; total CV ≤0.249954%. Shape-neutral ON aggregation is **1,803,708 kernels / 171.130066 s**. The whole series remains **VOID** because OFF/summary/manifest/marker are absent. Cleanup returns GPU/lock/port idle; binding remains **55/124**, W3 stays unowned `READY`, and a fresh no-model-key c2 series must repeat every arm. |
| 2026-07-14 (`SERVE-GATE-ONLINE` fifth async-credit execution + durable finalizer; complete diagnostic, neutral for speed) | Executes clean `3812d8` through all six explicit async ON/OFF timing legs and both shape-neutral Torch traces under one uncontended lock. Replaces the fragile inline derived step with a standard-library fail-closed finalizer that validates six raw legs, requested/resolved mode metadata, trace contracts and per-kernel totals, records output non-invariance, hashes every immutable artifact and writes the completion marker last. Production inference and local async defaults are unchanged; live surfaces collapse all precursor chronology to this accepted result. | Root `~/work/vllm-async-credit/3812d8d2b4a68d2e501007d01fe10cdf17751d02`; finalizer/summary/manifest/marker/artifact SHA `b3082a6e…1633` / `35b7344a…c323` / `e757b4ad…86c6` / `aa1e410b…369c` / `ead68397…8e56`; selected trace SHA ON/OFF `57413dd1…1cba` / `89bb9900…3c4`. Anchors: [profiler](../tools/bench/profile_vllm_online_gate.py#L21), [finalizer](../tools/bench/finalize_async_credit.py#L1), [ported contracts](../tests/tools/test_async_credit_summary.py#L1), [async spike](specs/async-serving.md), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Six × **6/6** timing requests pass. ON/OFF medians are **160.347697 / 160.003134 tok/s = 1.002153×** total, **106.642353 / 107.739836 ms = 1.010291×** TPOT, and **807.657803 / 696.329685 ms = 0.862159×** TTFT. Shape-neutral traces are **1,798,044 / 170.819890 s** ON and **1,810,902 / 170.478267 s** OFF: **1.002004×** GPU time, so no 1.04× speed credit. ON digests are stable; OFF/pairs vary diagnostically with batch shape. Cleanup passes; focused **26/26**, all tool tests **55/55**, record checker green. Binding stays **55/124**; W3 remains unowned `READY`, and speed work moves to exact low-batch RMSNorm/generated-partition + FP4-tactic mapping. |
| 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 trace contract; local capture pending) | Reconstructs the accepted c2 oracle raw trace into an exact ordered executed-path contract, then parameterizes only the trace-build CUDA observer/server/driver for an explicitly requested B=S=2 graph while retaining B=16 as the accepted default. The c2 driver keeps the model gate, frozen plans, three local Nsight sessions × four ranges, six-prompt closed-loop corpus, fresh paired oracle trace, cache/lifecycle evidence and one lock. It deliberately emits no accepted low-batch status until a dedicated finalizer lands; production builds and inference dispatch are unchanged. | Accepted oracle root `~/work/vllm-async-credit/3812d8d2b4a68d2e501007d01fe10cdf17751d02`, selected trace SHA `57413dd1…1cba`; local anchors [controller](../include/vt/cuda/cuda_profiler_control.h#L13), [CUDA observer](../src/vt/cuda/cuda_backend.cu#L226), [server flag](../examples/server/main.cpp#L90), [driver](../scripts/dgx-online-serving.sh#L14), [validator](../tools/bench/online_gate.py#L1761), [requested-batch contract](../tests/tools/test_online_gate_client.py#L860), [online-gate spike](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Oracle evidence is exact: **1,524 clean B=2 windows / 1,160 kernels each**, identical ordered-name/signature SHA `858915dd…fad0` / `b5c6fcac…dd7b`, **177 generated RMSNorm/quant calls / 0.442805 ms** and FP4 tactics **128 Stream-K 128x64x256 + 80 static-persistent 128x32x256** per window. CPU **106/106**, tool **57/57**, focused **35/35**, and record/mutation/doc **18/18** pass; no GPU command ran. Local c2 capture/final status is **PENDING**, so binding remains **55/124**, no residual/speed credit/implementation leaf is claimed, and exact-grid/35B performance stay blocked. |
| 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 trace contract; local capture pending) | Reconstructs the accepted c2 oracle raw trace into an exact ordered executed-path contract, then parameterizes only the trace-build CUDA observer/server/driver for an explicitly requested B=S=2 graph while retaining B=16 as the accepted default. The c2 driver keeps the model gate, frozen plans, three local Nsight sessions × four ranges, six-prompt closed-loop corpus, fresh paired oracle trace, cache/lifecycle evidence and one lock. It deliberately emits no accepted low-batch status until a dedicated finalizer lands; production builds and inference dispatch are unchanged. | Accepted oracle root `~/work/vllm-async-credit/3812d8d2b4a68d2e501007d01fe10cdf17751d02`, selected trace SHA `57413dd1…1cba`; local anchors [controller](../include/vt/cuda/cuda_profiler_control.h#L13), [CUDA observer](../src/vt/cuda/cuda_backend.cu#L226), [server flag](../src/vllm/entrypoints/openai/server_main.cpp#L692), [driver](../scripts/dgx-online-serving.sh#L14), [validator](../tools/bench/online_gate.py#L1761), [requested-batch contract](../tests/tools/test_online_gate_client.py#L860), [online-gate spike](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Oracle evidence is exact: **1,524 clean B=2 windows / 1,160 kernels each**, identical ordered-name/signature SHA `858915dd…fad0` / `b5c6fcac…dd7b`, **177 generated RMSNorm/quant calls / 0.442805 ms** and FP4 tactics **128 Stream-K 128x64x256 + 80 static-persistent 128x32x256** per window. CPU **106/106**, tool **57/57**, focused **35/35**, and record/mutation/doc **18/18** pass; no GPU command ran. Local c2 capture/final status is **PENDING**, so binding remains **55/124**, no residual/speed credit/implementation leaf is claimed, and exact-grid/35B performance stay blocked. |
| 2026-07-14 (`SERVE-GATE-ONLINE` first exact local-c2 execution + batch-keyed validator repair; `FAILED / VOID`) | Executes clean `ad8b58f` through the real 27B model gate and the first of three local B=2 Nsight sessions under one lock. The four-replay controller, six-prompt client, frozen plan map and graceful target lifecycle pass; range validation then applies c16's 1,107-kernel contract to the valid 1,011-kernel B=2 graph and fails closed before sessions 2/3 or the oracle. The repair retains c16 and keys an independently tested 27B/B=2 contract through validator, summarizer and driver. Production inference is unchanged. | Root `~/work/vllm.cpp-executed-path-c2/ad8b58f8708ce9bdf32aa9043611b3f6049be7fd`; run/execution/model-gate/control/raw-report-set/evidence-set SHA `9f285fd6…0aec` / `2a3d326f…6b56` / `bc9dc95b…da17` / `ee1589df…c719` / `f3aa3ca9…64c6` / `0da532ac…ac35`. Anchors: batch contracts and [validator](../tools/bench/online_gate.py#L99), [driver](../scripts/dgx-online-serving.sh#L666), [ported requested-batch cases](../tests/tools/test_online_gate_client.py#L1083), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Model gate passes **1/1**; local client/probe pass **6/6 + 2/2** and profile control records B=2, four replays, 64 frozen plans, zero tuning. Read-only reconstruction proves all four reports lossless and identical at **1,011 kernels + 7 memcpy + 1 memset**, multiset SHA `6b75bcff…1ce3`; mean traced kernel time is 109.722456 ms, FP4 is the oracle-matched **128+80** split, and RMSNorm-family structure is **177 calls / 2.237944 ms**. The series is **VOID**: no sessions 2/3, fresh oracle, final status, or binding ratio. Focused **35/35**, all tool **57/57**, policy **18/18**, and record/doc checks pass; cleanup returns GPU/lock/port idle. Binding remains **55/124** and full repaired retry is required. |
| 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 complete raw capture + durable-finalizer implementation; raw complete / finalizer `PENDING`) | Executes clean `179a0fc` through the whole repaired paired trace series under one uncontended lock, then adds a standard-library fail-closed finalizer and exact-B=2 ported contracts. The finalizer reuses the batch-keyed raw validator, proves local range invariance, accepts only the observed steady-oracle topology plus bounded drains and launch-signature allowlist, resolves complete family/tactic counts, hashes the immutable artifact set and writes the status marker last. Production inference and accepted c16 behavior are unchanged; live status surfaces retain only this current snapshot. | Raw root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; fresh oracle trace SHA `2b3bf412…785c`; local node SHA `44fcf31f…b93d`; anchors: batch-aware [validator](../tools/bench/online_gate.py#L99), [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1), [ported finalizer contracts](../tests/tools/test_low_batch_trace_summary.py#L1), [online-gate spike](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). A full read-only finalizer preflight against the raw root passes; committed durable summary/manifest/marker hashes remain pending. | Model gate passes. All **12/12** local ranges are lossless/invariant at **1,011+7+1**; the oracle has **1,522×1,160** invariant steady B=2 windows plus two bounded drains. Diagnostic median local/oracle kernel time is **111.076528 / 105.520831 ms = 1.052650×**. BF16 GEMMs lead at **193/51.662672 ms vs 97/48.798042 ms** (+96 launches/+2.864630 ms), followed by equal-count RMSNorm partitions at **2.249728 vs 0.439491 ms** (+1.810237 ms); FP4 tactics match **128+80** and FP4 time is non-positive. Cross-profiler timing is non-binding; durable final status, speed credit, exact-grid rerun and 35B performance remain **PENDING**. Binding stays **55/124**; cleanup returns GPU/lock/port idle. |
| 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 durable finalization; `COMPLETE DIAGNOSTIC`) | Runs the exact pushed `fe28003` finalizer bytes once against immutable raw `179a0fc`, revalidates the full raw chain, and writes the completion marker last. This changes evidence lifecycle only: production inference, c16 validation and the binding performance grid are unchanged. Live surfaces replace the pending checkpoint with this durable snapshot while detailed chronology remains here and in state. | Root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; summary / manifest / status / artifact-set / finalizer SHA `0ef6a124…0273` / `2556cfd0…2f21` / `9e0143fa…7b57` / `cc248ad2…823a` / `45dbf28a…3311`; run-log SHA `362f0f1e…1cef`. Anchors: committed [finalizer](../tools/bench/finalize_low_batch_trace.py#L1), [ported contracts](../tests/tools/test_low_batch_trace_summary.py#L1), [serving-gate spec](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Status is **`complete-diagnostic`**. It retains **12/12** invariant local **1,011+7+1** ranges, **1,522×1,160** steady oracle B=2 windows plus two drains, and matching **128+80** FP4 tactics. Diagnostic local/oracle medians remain **111.076528 / 105.520831 ms = 1.052650×**; BF16 GEMMs lead at +96 launches/+2.864630 ms, RMSNorm follows at +1.810237 ms, and FP4 is non-positive. Cross-profiler timing earns no speed credit; binding remains **55/124**. No GPU command ran, GPU/lock/port are idle, and the next legal checkpoint is the BF16-GEMM whole-chain spike before code. |
Expand Down
2 changes: 1 addition & 1 deletion docs/BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -380,7 +380,7 @@ built on it rather than keeping the flattering one.
| Laguna NVFP4 decode | `flock $HOME/gpu.lock ./build-cuda/examples/laguna-gen --model ~/laguna-xs-nvfp4 --gpu` (that directory holds the S-2.1 checkpoint); `drop_caches` first, create the CUDA context before loading weights |
| DeepSeek-V4-Flash decode | `deepseek-v4-gen --gpu --kv-cache` on `ds4flash.gguf`, captured under tmux |
| Metal vs MLX-LM | Paired A/B harness, interleaved runs, cold legs discarded |
| Vulkan vs llama.cpp Vulkan | Not yet runnable (no model runs on Vulkan). Planned harness in `.agents/specs/vulkan-full-support.md` §5.2 |
| Vulkan vs llama.cpp Vulkan | Same GGUF both arms: ours `-DVLLM_CPP_VULKAN=ON`, llama.cpp `-DGGML_VULKAN=ON` at `237ad9b96` via `llama-bench`; clean legs only, one `flock $HOME/gpu.lock`. GEMV sweep: `benchmarks/vulkan_gemv_ab.cpp` |

Build flags, environment variables, and the full gate list are in
[BUILD.md](BUILD.md) and [ENVIRONMENT.md](ENVIRONMENT.md).
1 change: 1 addition & 0 deletions docs/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -155,6 +155,7 @@ on CUDA/CPU builds beyond the documented behavior.
| `VT_GEMMA4_RESIDENT_EXPERTS` | unset | `=1` preloads the Gemma-4 MoE experts resident on the GPU(s) after the first use instead of streaming them per step (discrete-ROCm optimization). No-op (with a stderr note) on a binary built without `-DVLLM_CPP_HIP` |
| `VT_GEMMA4_RESIDENT_GPUS` | `2` | Number of GPUs across which resident Gemma-4 experts are spread; clamped to the ROCm device count. Read only when `VT_GEMMA4_RESIDENT_EXPERTS=1` |
| `VT_GEMMA4_RESIDENT_MAX_LAYERS` | (all) | Caps how many MoE layers get resident-preloaded, to fit a smaller VRAM budget. Read only when `VT_GEMMA4_RESIDENT_EXPERTS=1` |
| `VLLM_CPP_HTTP_FIXED_POOL` | `1` (fixed) | `=0` reverts the HTTP worker pool to the legacy dynamic mode. Production uses the capacity-derived fixed pool; the opt-out exists for same-binary A/B attribution |
| `VT_ROCM_ATTN_CPU_REF` | unset | `=1` routes ROCm paged attention through the CPU reference kernel instead of the HIP kernel — a correctness A/B for the ROCm attention bring-up |
| `VT_DEBUG_SAMPLED` | unset | `=1` prints the per-step sampled token id(s) to stderr (sampling-loop debug). Read-only; does not change output. Read once per token, so it does not stall the hot loop |

Expand Down
Loading
Loading