perf(dspark): propose two tokens where the draft already drafts four, and widen NVFP4 activation loads - #896
Conversation
… and widen NVFP4 activation loads
dd050b3 to
dd62c4a
Compare
Closed — the reported baseline does not reconcile with this repo'sNot a judgement on the code. The problem is the measurement, and it is specific. You report benchmarking against pristine
That matters for this PR specifically, because the headline change is "propose two tokens where the Worth checking on your side:
If you can produce a before/after where the main baseline lands near 112 tok/s at tau ~1.66, reopen |
Summary
The headline change: propose two tokens in the narrow band instead of one. The measurement that pinned it at one is stale, and the draft is already paying for the rows it throws away.
Proof of speedup
sm_120)Decode tok/s (
dspark-decode@4k,dspark_tau_check, pristineorigin/main@30ddb88vs this PR in separate build trees, run alternately on the same box):Guards:
ctest19/19 ·LOSSLESS=1verified across repeats · tau unchanged · AR improved (91.77 -> 92.94, floor bar 0.98x).