Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
54 commits
Select commit Hold shift + click to select a range
597ee89
perf(gemma4): pack resident MoE on GPU0 first then spill GPU1
Aug 8, 2026
357583a
fix(gemma4): stream FP8 resident upload — no host BF16 cache OOM
Aug 8, 2026
84ee0bf
fix(gemma4): full dual-GPU FP8 resident without host/GPU OOM
Aug 8, 2026
ce8bfb1
perf(gemma4): peer-copy off-device resident experts (stay on GPU)
Aug 8, 2026
4fe4ed7
perf(gemma4): device MulScalar+Add for MoE expert mix (stay on GPU)
Aug 8, 2026
fd894a6
perf(gemma4): reuse MoE top-k scratch buffers
Aug 8, 2026
ac1d669
perf(gemma4): allocate peer MoE scratch only when needed
Aug 8, 2026
b4a994b
feat(cli): --repeat N for load-once warm tok/s benches
Aug 8, 2026
18d74bf
perf(gemma4): parallel FP8 expert dequant (cold path)
Aug 8, 2026
64b3876
perf(gemma4): pin host FP8 expert BF16 caches + device-upload diag
Aug 8, 2026
876eecb
perf(gemma4): reuse ExpertGeGLU scratch across top-k
Aug 8, 2026
38028ea
feat(rocm): strided/pointer batched MatmulBT for MoE (opt-in)
Aug 8, 2026
40166fb
perf(gemma4): fused Gelu across top-k in batched MoE (opt-in)
Aug 8, 2026
c80838a
perf(gemma4): hoist MoE scratch outside token loop
Aug 8, 2026
959d7f6
feat(rocm): device MoeRouterTopK + Gemma4 uses it
Aug 8, 2026
8a5464e
perf(gemma4): first-expert writes ysum directly (skip Zero+Add)
Aug 8, 2026
e3bdce0
perf(gemma4): separate GeLU*up + TLS MoE scratch across layers
Aug 8, 2026
f2cc17c
perf(gemma4): fuse expert mix into down-proj GEMM (alpha/beta)
Aug 8, 2026
5d239fe
perf(gemma4): PLE GeGLU via GeluMulSeparate (no 2*ple pack)
Aug 8, 2026
ad09bb7
perf(gemma4): TLS reuse of per-layer ForwardBody scratch buffers
Aug 8, 2026
dd2463a
perf(gemma4): contiguous PLE-by-layer layout [L,T,ple]
Aug 8, 2026
ae2fd08
perf(gemma4): TLS dense MLP gate_up temps + CLI fflush multi-run stats
Aug 8, 2026
25c9216
perf(gemma4): async expert H2D + optional VRAM LRU (no per-expert sync)
Aug 8, 2026
4895378
perf(gemma4): fuse RmsNorm+residual residual joins (ROCm)
Aug 8, 2026
7bd4b70
perf(gemma4): fuse expert gate+up into one MatmulBT (3→2 GEMMs/expert)
Aug 8, 2026
5fa5f20
perf(gemma4): prefetch top-k expert H2D + fused gate_up in batch path
Aug 8, 2026
2de2457
chore(rocm): GemmCompute selector (default 32F; 16bf/16f unsupported …
Aug 8, 2026
31eeb00
perf(gemma4): DualRmsNormPlusRes fuse MoE post-FF (3 norms+add+res)
Aug 8, 2026
86be92c
perf(rocm): thread-local hipBLAS handle (no per-GEMM mutex/SetStream)
Aug 8, 2026
4e8ada6
fix(rocm): restore mutex/unordered_map includes for batched GEMM tables
Aug 8, 2026
ed25009
perf(gemma4): TLS QKV/attn temps in Gemma4AttnBlock
Aug 8, 2026
cddf4fe
perf(gemma4): TLS router RMS ones-weight + VT_GEMMA4_PROFILE breakdown
Aug 8, 2026
43da184
feat(rocm): opt-in M=1 BF16 GEMV path (default off; hipblas faster)
Aug 8, 2026
381030b
perf(rocm): LDS-cached M=1 BF16 GEMV (still opt-in; hipblas faster)
Aug 8, 2026
61f3502
feat(gemma4): fused top-k ExpertGeGLU HIP path (opt-in, correct, slow)
Aug 8, 2026
ba36e9e
docs: PR draft — Vulkan Q8 ~98 t/s bar vs ROCm FP8 ~34
Aug 8, 2026
ab7b124
feat(rocm): hipBLASLt MatmulBT path (VT_ROCM_HIPBLASLT=1; default off)
Aug 8, 2026
e383804
fix(rocm): hipBLASLt MatmulBT via COL layout (matches GemmEx; no hang)
Aug 8, 2026
cd2e032
chore(gemma4): disable broken expert batch paths (gather/strided wron…
Aug 8, 2026
d97290a
fix(gemma4): resident mode — no FP8 device cache when packs exist; 12…
Aug 8, 2026
c6c8531
feat(gemma4): native FP8 expert path (VT_GEMMA4_FP8_NATIVE=1; default…
Aug 8, 2026
a897f60
perf(gemma4): layer profile attn/mlp/moe under VT_GEMMA4_PROFILE
Aug 8, 2026
f119e9c
perf(gemma4): top-k fused GeluAndMul (one gelu for all experts)
Aug 8, 2026
dd3f701
perf(gemma4): fused top-k path uses MatmulBTAlphaBetaRocm directly
Aug 8, 2026
1d1e93b
docs: gfx1201 FP8 Lt matrix — W8A8 works but slower than BF16 M=1
Aug 8, 2026
f4436f8
feat(gemma4): custom RDNA4 BF16 Expert GeGLU kernels (VT_GEMMA4_CUSTO…
Aug 8, 2026
f42f4ff
feat(server): load chat_template.jinja + verbose chat request logs
Aug 8, 2026
f162d93
feat(server): chat-dbg stage heartbeats while processing prompts
Aug 8, 2026
5f70f66
feat(server): vLLM-style request logger + metrics for Hermes serve
Aug 8, 2026
98b2bbd
fix(server): reject huge Hermes SOUL prompts; clamp max_tokens
Aug 8, 2026
0483ef9
feat(server): Hermes long-ctx defaults — 200k prompt cap, APC note
Aug 8, 2026
0bae7cb
feat(sched): log chunked prefill progress (INFO prefill computed=N/M %)
Aug 8, 2026
3cbecb5
fix(engine): generate waits on own request_id; chat enable_thinking flag
Aug 8, 2026
e086435
fix(gemma4): portable vt::fused_ops seam — no vt::rocm in models/
Aug 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -840,6 +840,7 @@ add_library(vllm STATIC
src/vllm/entrypoints/openai/serving_utils.cpp
src/vllm/entrypoints/openai/serving_completion.cpp
src/vllm/entrypoints/openai/serving_chat.cpp
src/vllm/entrypoints/openai/request_logger.cpp
src/vllm/entrypoints/openai/serving_models.cpp
src/vllm/entrypoints/openai/run_batch.cpp
src/vllm/entrypoints/openai/tool_parsers/abstract.cpp
Expand Down Expand Up @@ -912,6 +913,7 @@ add_library(vllm STATIC
src/vt/op_provider.cpp
src/vt/communicator.cpp
src/vt/ops.cpp
src/vt/fused_ops.cpp
src/vt/merged_gemm.cpp
src/vt/cuda/nvfp4_persistent_cache.cpp
src/vt/cpu/cpu_backend.cpp
Expand Down Expand Up @@ -1145,6 +1147,10 @@ if(VLLM_CPP_HIP)
src/vt/rocm/rocm_matmul_hipblaslt.hip
src/vt/rocm/rocm_paged_attn.hip
src/vt/rocm/rocm_gemma4_experts.hip
src/vt/rocm/rocm_gemma4_fused_experts.hip
src/vt/rocm/rocm_gemma4_expert_geglu.hip
src/vt/rocm/rocm_fp8_channel_gemv.hip
src/vt/rocm/rocm_moe_router.hip
src/vt/rocm/rocm_ops.hip)
if(VLLM_CPP_HIP_ARCHITECTURES)
set_source_files_properties(
Expand All @@ -1155,6 +1161,10 @@ if(VLLM_CPP_HIP)
src/vt/rocm/rocm_matmul_hipblaslt.hip
src/vt/rocm/rocm_paged_attn.hip
src/vt/rocm/rocm_gemma4_experts.hip
src/vt/rocm/rocm_gemma4_fused_experts.hip
src/vt/rocm/rocm_gemma4_expert_geglu.hip
src/vt/rocm/rocm_fp8_channel_gemv.hip
src/vt/rocm/rocm_moe_router.hip
src/vt/rocm/rocm_ops.hip
PROPERTIES HIP_ARCHITECTURES "${VLLM_CPP_HIP_ARCHITECTURES}")
endif()
Expand Down
9 changes: 9 additions & 0 deletions docs/BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -373,3 +373,12 @@ built on it rather than keeping the flattering one.

Build flags, environment variables, and the full gate list are in
[BUILD.md](BUILD.md) and [ENVIRONMENT.md](ENVIRONMENT.md).

## 2026-08-08 — Gemma4 FP8 stream lab (gfx1201 R9700)

| Path | Warm tok/s | Notes |
|------|------------|--------|
| vllm-cli Paris HIP=0 stream experts | ~38 | `--repeat` after cold expert fill |
| server `/v1/completions` | ~38 | exclusive |
| server `/v1/chat` thinking off | ~32 | after expert cache |
| llama.cpp Vulkan Q8 tg128 (bar) | ~98 | separate stack |
2 changes: 2 additions & 0 deletions docs/FEATURES.md
Original file line number Diff line number Diff line change
Expand Up @@ -307,3 +307,5 @@ backends in scope as inventoried rows. Neither changed a capability, so **no
mark on this page moved**. An inventoried backend is not a supported one, and the same
holds for the 31 architectures inventoried on 2026-08-05. A row's lifecycle state and its support mark
are independent: see [STATUS.md](STATUS.md). Parakeet ASR (encoder + CTC/RNN-T/TDT) runs natively on CPU, 4 checkpoints token-exact vs HF.

| Gemma4 MoE ROCm fused helpers (`vt::fused_ops`) | partial | Portable seam; ROCm fast path; CPU/Vulkan link |
8 changes: 8 additions & 0 deletions docs/STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2210,3 +2210,11 @@ and outside this repo.
**Agent onboarding:** [session](../.agents/specs/session-onboarding.md) +
[entry](../.agents/specs/developer-agent-protocol-entrypoint.md) implemented;
documentation-only.

## 2026-08-08 — Gemma4 ROCm fused helpers via portable vt:: seam (#154)

Model files (`gemma4.cpp`, `gemma4_moe.cpp`) no longer call `vt::rocm::*` directly.
Fused paths go through `include/vt/fused_ops.h` (`vt::RmsNormPlusAdd`,
`DualRmsNormPlusRes`, `GeluMulSeparate`, `MatmulBTAlphaBeta`, `MatmulBTFp8Channel`,
`ExpertGeGLUBf16TopKM1`). ROCm fast path under `VLLM_CPP_HIP`; non-HIP stubs for
peer/pin/resident upload. `check-device-leakage` holds baseline.
55 changes: 39 additions & 16 deletions examples/cli/main.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -8,14 +8,16 @@
// vllm-cli --model <dir> --prompt "<text>"
// [--tokenizer-config <path>] [--device auto|cpu|cuda]
// [--max-tokens N] [--temperature T] [--top-p P] [--top-k K]
// [--seed S] [--stream]
// [--seed S] [--stream] [--repeat N]
// [--gpu-memory-utilization F] [--kv-cache-memory BYTES]
//
// <dir> holds config.json, tokenizer.json and the *.safetensors shards (T0:
// safetensors only). Loading a real checkpoint is a GPU/dgx concern; on a CPU
// box `--help` / bad-args still work without a model (smoke-tested in CI).
// --repeat N runs N completions after one load (warm bench / decode tok-s).
#include "vllm.h"

#include <chrono>
#include <cstdio>
#include <cstdlib>
#include <cstring>
Expand All @@ -36,6 +38,7 @@ struct Args {
unsigned long long seed = 0;
bool have_seed = false;
bool stream = false;
int repeat = 1; // load once, complete N times (warm tok/s)
std::string speculative_config; // vLLM --speculative-config JSON; "" => off.
// --device (ABI v14): "auto" (default probe), "cpu", or "cuda" — the names
// of vLLM's DeviceConfig.device this build serves. Mapped to the int the ABI
Expand All @@ -54,12 +57,13 @@ void Usage(const char* argv0, std::FILE* out) {
"usage: %s --model <dir> --prompt \"<text>\"\n"
" [--tokenizer-config <path>] [--device auto|cpu|cuda]\n"
" [--max-tokens N] [--temperature T] [--top-p P] [--top-k K]\n"
" [--seed S] [--stream]\n"
" [--seed S] [--stream] [--repeat N]\n"
" [--gpu-memory-utilization F] [--kv-cache-memory BYTES]\n"
" [--speculative-config '<json>']\n"
"\n"
"Runs one completion over the vllm.cpp C ABI (libvllm). <dir> holds\n"
"config.json, tokenizer.json and the *.safetensors shards.\n",
"Runs completion(s) over the vllm.cpp C ABI (libvllm). <dir> holds\n"
"config.json, tokenizer.json and the *.safetensors shards.\n"
"--repeat N: load once, run N blocking completions (for warm tok/s).\n",
argv0);
}

Expand Down Expand Up @@ -99,6 +103,9 @@ bool ParseArgs(int argc, char** argv, Args& a, int& exit_code) {
a.have_seed = true;
} else if (flag == "--stream") {
a.stream = true;
} else if (flag == "--repeat") {
a.repeat = std::atoi(NextArg(argc, argv, i));
if (a.repeat < 1) a.repeat = 1;
} else if (flag == "--speculative-config") {
a.speculative_config = NextArg(argc, argv, i);
} else if (flag == "--gpu-memory-utilization") {
Expand Down Expand Up @@ -218,6 +225,9 @@ int main(int argc, char** argv) {
int rc = 0;
if (args.stream) {
// ── Streaming: print deltas as they arrive. ──────────────────────────────
if (args.repeat != 1) {
std::fprintf(stderr, "vllm-cli: --repeat with --stream not supported; using 1\n");
}
st = vllm_complete_stream(engine, args.prompt.c_str(), &sp, &StreamPrintCb,
nullptr);
std::fputc('\n', stdout);
Expand All @@ -227,20 +237,33 @@ int main(int argc, char** argv) {
rc = 1;
}
} else {
// ── Blocking: run to completion, then print the whole text. ──────────────
vllm_completion out;
st = vllm_complete(engine, args.prompt.c_str(), &sp, &out);
if (st != VLLM_OK) {
std::fprintf(stderr, "vllm-cli: completion failed (status %d): %s\n",
static_cast<int>(st), vllm_last_error());
rc = 1;
} else {
std::fputs(out.text != nullptr ? out.text : "", stdout);
std::fputc('\n', stdout);
// ── Blocking: load once, optionally repeat for warm tok/s. ───────────────
for (int r = 0; r < args.repeat; ++r) {
vllm_completion out{};
const auto t0 = std::chrono::steady_clock::now();
st = vllm_complete(engine, args.prompt.c_str(), &sp, &out);
const auto t1 = std::chrono::steady_clock::now();
const double secs =
std::chrono::duration<double>(t1 - t0).count();
if (st != VLLM_OK) {
std::fprintf(stderr, "vllm-cli: completion failed (status %d): %s\n",
static_cast<int>(st), vllm_last_error());
rc = 1;
break;
}
if (r == 0 || args.repeat == 1) {
std::fputs(out.text != nullptr ? out.text : "", stdout);
std::fputc('\n', stdout);
}
const int ct = out.completion_tokens;
const double tps = (secs > 0.0 && ct > 0) ? (static_cast<double>(ct) / secs) : 0.0;
std::fprintf(stderr,
"vllm-cli: finish_reason=%s prompt_tokens=%d completion_tokens=%d\n",
"vllm-cli: run=%d/%d finish_reason=%s prompt_tokens=%d "
"completion_tokens=%d secs=%.3f tok_s=%.3f\n",
r + 1, args.repeat,
out.finish_reason != nullptr ? out.finish_reason : "(none)",
out.prompt_tokens, out.completion_tokens);
out.prompt_tokens, ct, secs, tps);
std::fflush(stderr);
vllm_completion_free(&out);
}
}
Expand Down
75 changes: 72 additions & 3 deletions examples/server/main.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -62,9 +62,11 @@
#include <fstream>
#include "vllm/entrypoints/openai/api_server.h"
#include "vllm/entrypoints/openai/chat_mm.h"
#include "vllm/entrypoints/openai/request_logger.h"
#include "vllm/entrypoints/openai/serving_chat.h"
#include "vllm/entrypoints/openai/serving_completion.h"
#include "vllm/entrypoints/openai/serving_models.h"
#include "vllm/v1/metrics/loggers.h"
#include "vllm/entrypoints/openai/reasoning_parsers/detect.h"
#include "vllm/entrypoints/openai/tool_parsers/detect.h"
#include "vllm/model_executor/model_loader/safetensors_reader.h"
Expand Down Expand Up @@ -178,6 +180,15 @@ struct Args {
// routers only under `if envs.VLLM_SERVER_DEV_MODE` (api_server.py:238). Off by
// default → /abort_requests 404s. Enables the /abort_requests production wiring.
bool enable_server_dev_mode = false;
bool verbose = false;
// Gemma4 HF/vLLM: --default-chat-template-kwargs enable_thinking (default OFF).
bool enable_thinking = false;
// Request logging (Python vLLM --enable-log-requests parity). Default ON.
bool enable_log_requests = true;
bool enable_log_outputs = false;
int max_log_len = 256;
// Attach Prometheus logger + GET /metrics (default ON for solid Hermes serve).
bool enable_metrics = true;
// Scheduling policy: "fcfs" (default), "priority" (mirrors vLLM's
// --scheduling-policy / SchedulerConfig.policy), or "lpm" (SGLang's
// cache-aware longest-prefix-match admission ordering, ENG-SGLANG-BEHAVIOR-FLAG;
Expand Down Expand Up @@ -225,6 +236,11 @@ struct Args {
" [--enable-force-include-usage]\n"
" [--enable-tokenizer-info-endpoint]\n"
" [--enable-server-dev-mode]\n"
" [--verbose]\n"
" [--enable-thinking|--no-enable-thinking]\n"
" [--enable-log-requests|--disable-log-requests]\n"
" [--enable-log-outputs] [--max-log-len N]\n"
" [--enable-metrics|--disable-metrics]\n"
" [--[no-]enable-prefix-caching]\n"
" [--[no-]enable-radix-attention]\n"
" [--scheduling-policy fcfs|priority|lpm]\n"
Expand Down Expand Up @@ -319,6 +335,24 @@ Args ParseArgs(int argc, char** argv) {
a.video_dequant_bf16 = true;
} else if (flag == "--enable-server-dev-mode") {
a.enable_server_dev_mode = true;
} else if (flag == "--verbose" || flag == "-v") {
a.verbose = true;
} else if (flag == "--enable-thinking") {
a.enable_thinking = true;
} else if (flag == "--no-enable-thinking") {
a.enable_thinking = false;
} else if (flag == "--enable-log-requests") {
a.enable_log_requests = true;
} else if (flag == "--disable-log-requests") {
a.enable_log_requests = false;
} else if (flag == "--enable-log-outputs") {
a.enable_log_outputs = true;
} else if (flag == "--max-log-len") {
a.max_log_len = std::stoi(NextArg(argc, argv, i, argv[0]));
} else if (flag == "--enable-metrics") {
a.enable_metrics = true;
} else if (flag == "--disable-metrics") {
a.enable_metrics = false;
} else if (flag == "--enable-prefix-caching" ||
flag == "--no-enable-prefix-caching" ||
flag == "--enable-radix-attention" ||
Expand Down Expand Up @@ -416,6 +450,26 @@ Args ParseArgs(int argc, char** argv) {
int main(int argc, char** argv) {
try {
const Args args = ParseArgs(argc, argv);
if (args.verbose) {
setenv("VT_SERVER_VERBOSE", "1", /*overwrite=*/1);
std::cerr << "server: verbose stage logging enabled (debug_stages)\n";
}
{
vllm::entrypoints::openai::RequestLogConfig log_cfg;
log_cfg.enable_log_requests = args.enable_log_requests;
log_cfg.enable_log_outputs = args.enable_log_outputs || args.verbose;
log_cfg.max_log_len = args.max_log_len;
log_cfg.debug_stages = args.verbose ||
(std::getenv("VT_SERVER_VERBOSE") &&
std::getenv("VT_SERVER_VERBOSE")[0] == '1');
vllm::entrypoints::openai::ConfigureRequestLogger(log_cfg);
std::cerr << "server: request logging "
<< (log_cfg.enable_log_requests ? "ON" : "OFF")
<< " outputs=" << (log_cfg.enable_log_outputs ? "ON" : "OFF")
<< " max_log_len=" << log_cfg.max_log_len
<< " debug_stages=" << (log_cfg.debug_stages ? "ON" : "OFF")
<< "\n";
}

const fs::path dir(args.model_dir);
const std::string config_path = (dir / "config.json").string();
Expand Down Expand Up @@ -666,9 +720,13 @@ int main(int argc, char** argv) {
const std::string eos =
tokenizer.EosId() >= 0 ? tokenizer.Decode({tokenizer.EosId()}) : "";
chat_prompt_fn =
vllm::entrypoints::MakeChatTemplatePromptFn(chat_template, bos, eos);
std::cerr << "server: using chat template from " << tokenizer_config_path
<< "\n";
vllm::entrypoints::MakeChatTemplatePromptFn(
chat_template, bos, eos, args.enable_thinking);
std::cerr << "server: using chat template (" << chat_template.size()
<< " chars) from " << tokenizer_config_path
<< " or sibling chat_template.jinja"
<< " enable_thinking="
<< (args.enable_thinking ? "true" : "false") << "\n";
} catch (const std::exception& e) {
std::cerr << "server: no chat template (" << e.what()
<< "); falling back to the simple role-join prompt\n";
Expand Down Expand Up @@ -855,6 +913,17 @@ int main(int argc, char** argv) {
: "")
<< "\n";

// Prometheus /metrics (Python vLLM always-on family names).
std::unique_ptr<vllm::v1::metrics::PrometheusStatLogger> prom_logger;
if (args.enable_metrics) {
prom_logger = std::make_unique<vllm::v1::metrics::PrometheusStatLogger>(
served_model_name, loaded->max_model_len(), /*engine_index=*/0);
// Sync engine path records on step(); async may under-report until fully wired.
loaded->engine().set_stat_logger(prom_logger.get());
server.set_metrics_logger(prom_logger.get());
std::cerr << "server: GET /metrics enabled (PrometheusStatLogger)\n";
}

std::cerr << "server: listening on http://" << args.host << ":" << args.port
<< " (model '" << served_model_name << "', HTTP worker pool ";
if (server.http_worker_count() == 0) {
Expand Down
15 changes: 9 additions & 6 deletions include/vllm/entrypoints/chat_template.h
Original file line number Diff line number Diff line change
Expand Up @@ -63,19 +63,22 @@ class ChatTemplateError : public std::runtime_error {
// branch. Empty => the `tools` variable is an empty list
// (falsy). The `tojson` filter is a minja builtin.
// Throws ChatTemplateError on any parse or evaluation error.
// chat_template_kwargs: optional Jinja variables (vLLM
// --default-chat-template-kwargs). Supported keys today: enable_thinking (bool).
std::string apply_chat_template(
const std::string& template_str,
const std::vector<openai::ChatMessage>& messages, bool add_generation_prompt,
const std::string& bos_token = "", const std::string& eos_token = "",
const std::vector<openai::ChatCompletionToolsParam>& tools = {});
const std::vector<openai::ChatCompletionToolsParam>& tools = {},
bool enable_thinking = false);

// Adapt a chat template to Task 2's ChatPromptFn seam (serving_chat.h). The
// returned callable renders `template_str` for the messages + generation flag it
// is handed, so an OpenAIServingChat constructed with it applies the real chat
// template instead of DefaultChatPromptFallback.
// Adapt a chat template to Task 2's ChatPromptFn seam (serving_chat.h).
// enable_thinking defaults false for agent/Hermes latency (Gemma4 empty thought
// block when false — HF/vLLM recipe parity).
openai::ChatPromptFn MakeChatTemplatePromptFn(std::string template_str,
std::string bos_token = "",
std::string eos_token = "");
std::string eos_token = "",
bool enable_thinking = false);

// Load the `chat_template` string out of a tokenizer_config.json file. Handles
// both the plain-string form and the list-of-{name,template} form (picks the
Expand Down
49 changes: 49 additions & 0 deletions include/vllm/entrypoints/openai/request_logger.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
// OpenAI serve request logging — shaped like Python vLLM --enable-log-requests.
// Ported concepts from: vllm/entrypoints/logger.py + api_server access logs.
#pragma once

#include <cstdint>
#include <string>

namespace vllm::entrypoints::openai {

struct RequestLogConfig {
// Mirrors --enable-log-requests / --disable-log-requests (default ON for serve).
bool enable_log_requests = true;
// Mirrors --enable-log-outputs (requires enable_log_requests).
bool enable_log_outputs = false;
// Mirrors --max-log-len (chars of prompt/output preview).
int max_log_len = 256;
// Lab deep stages (former VT_SERVER_VERBOSE chat-dbg).
bool debug_stages = false;
};

// Process-wide config (set once at server startup).
void ConfigureRequestLogger(const RequestLogConfig& cfg);
const RequestLogConfig& GetRequestLogConfig();

// Truncate for log lines (max_log_len); escapes newlines.
std::string LogPreview(const std::string& s, int max_len);

// HTTP ingress (api_server).
void LogHttpIngress(const char* method, const char* path, size_t body_bytes);

// After chat/completions parse + template.
void LogRequestReceived(const std::string& request_id, const std::string& endpoint,
const std::string& model, bool stream, int max_tokens,
int n_messages, int n_tools, size_t prompt_chars,
const std::string& prompt, const std::string& roles_summary);

// Mid-flight stages (debug_stages only) — heartbeats during prefill/decode.
void LogRequestStage(const std::string& request_id, const std::string& stage);

// Completion of a request.
void LogRequestFinished(const std::string& request_id, int prompt_tokens,
int completion_tokens, const std::string& finish_reason,
double elapsed_sec, const std::string& output_text);

// Errors.
void LogRequestError(const std::string& request_id, const std::string& endpoint,
const std::string& what);

} // namespace vllm::entrypoints::openai
4 changes: 4 additions & 0 deletions include/vllm/model_executor/model_loader/nvfp4_dequant.h
Original file line number Diff line number Diff line change
Expand Up @@ -84,4 +84,8 @@ void DequantFp8ToBf16(const uint8_t* weight_f8, float weight_scale,
void DequantFp8ChannelToBf16(const uint8_t* weight_f8, const uint16_t* scale_bf16,
int64_t N, int64_t K, uint16_t* out_bf16);

// Nesting guards: expert-level parallel prefetch serializes row-parallel dequant.
void Fp8DequantBeginOuterParallel();
void Fp8DequantEndOuterParallel();

} // namespace vllm
Loading
Loading