Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
150 commits
Select commit Hold shift + click to select a range
5c7c22c
opencl: flush profiling batch at shutdown for incomplete batches (#25…
shaofeiqi Jun 26, 2026
960d628
mamba2: remove hardcoded 2x expansion factor and invalid d_inner % d_…
limloop Jun 26, 2026
f818065
CUDA: batch out_prod broadcast (dps2>1) path with cublasSgemmBatched …
leonardHONG Jun 26, 2026
b11f7c1
mtmd: add more validations (#25013)
ngxson Jun 26, 2026
e7e3f35
sycl : clamp softmax input to avoid underflow (#24941)
Jassieluo Jun 26, 2026
1a87dcd
server + ui: SSE Replay Buffer (#23226)
ServeurpersoCom Jun 26, 2026
c16c35b
ggml-cpu: fix SVE leftover path in ggml_vec_dot_f32 (#24699)
tdakhran Jun 26, 2026
2f18fe1
CUDA: add cublasSgemmBatched mapping for HIP/MUSA vendor headers (#25…
leonardHONG Jun 26, 2026
9df0680
vulkan: Workaround compiler bug in conv2d coopmat2 path (#24924)
jeffbolznv Jun 26, 2026
ded1561
ui: fix accessibility for hover-gated interactive elements assisted b…
sanjayahari Jun 26, 2026
5a6a0dd
vulkan: add INTEL_XE1 arch enum and enable coopmat1 on Intel Xe-LPG P…
fish-jiang Jun 26, 2026
487a6cc
vulkan: opt mul_mat_vecq for mi50 (#22933)
chraac Jun 26, 2026
96183e9
ggml : bump version to 0.15.3 (ggml/1550)
ggerganov Jun 26, 2026
e7ea94a
sync : ggml
ggerganov Jun 26, 2026
5397c36
openvino: Update to OV 2026.2.1, self-contained release packages, ope…
ravi9 Jun 26, 2026
024930c
arg: fix handling --spec-draft-hf and --hf-repo-v (#25043)
ngxson Jun 26, 2026
5d8ccdf
devops : add llama in all docker images (#25035)
angt Jun 26, 2026
3fc4e10
sched : reintroduce less synchronizations during split compute (#20793)
aendk Jun 26, 2026
050ee92
app : allow --version, --licenses & --help (#25054)
angt Jun 26, 2026
83d385b
tests : fix test-chat-template --no-common option (#25075)
CISC Jun 27, 2026
0275c0f
ci : add windows-openvino to check-release (#25022)
CISC Jun 27, 2026
c299a92
binaries : Improve rpc-server and export-graph-ops names. (#25045)
ckastner Jun 27, 2026
0b6529d
vulkan: fix step operator for 0 input (#25036)
0cc4m Jun 27, 2026
9bebfcb
sycl : fix failed ut cases of norm (#25044)
arthw Jun 27, 2026
0ed235e
[CUDA] Added a cudaMemcpy2DAsync fast path to ggml_cuda_cpy (#25057)
gaugarg-nv Jun 27, 2026
ebd048f
opencl: flash attention improvement (#25069)
wanghqc Jun 27, 2026
27c8bb4
logs : reduce v2 (#25078)
ggerganov Jun 28, 2026
c1a1c8e
common : allow --offline in llama download (#25091)
angt Jun 28, 2026
d1b3425
spec : add DFlash support (#22105)
ruixiang63 Jun 28, 2026
f68a788
jinja: add --dump-prog for debugging (#25086)
ngxson Jun 28, 2026
c818263
chat : implement minicpm5 parser (#24889)
aldehir Jun 28, 2026
fa72bc6
dflash: refactor draft model conversion (#25110)
ruixiang63 Jun 28, 2026
7cb8576
ui: fix stop and reasoning skip in single-model mode (#25084)
ServeurpersoCom Jun 28, 2026
dbdaece
Revert "ui: fix accessibility for hover-gated interactive elements as…
allozaur Jun 28, 2026
b3fed31
jinja, chat: add --reasoning-preserve flag (#25105)
ngxson Jun 28, 2026
277a105
common : remove unused regex-partial (#25118)
o7si Jun 29, 2026
6cb18b2
tools/ui: restore Tailwind scanning in ignored worktrees (#24879)
seryogakovalyov Jun 29, 2026
8c146a8
DeepSeek V4 (#24162)
am17an Jun 29, 2026
25a1d63
vulkan: use flops instead of weight tensor size for submission heuris…
0cc4m Jun 29, 2026
6f4f53f
common : dedup preset and cached model entries in /v1/models (#25131)
angt Jun 29, 2026
86b9470
Revert "sched : reintroduce less synchronizations during split comput…
ORippler Jun 30, 2026
6c5de1c
ggml-webgpu: add support for NVFP4 (#25143)
yomaytk Jun 30, 2026
d9df110
HIP: use hipBLAS for dense prefill on gfx900, keep MMQ for MoE (#24588)
DEV-DUFORD Jun 30, 2026
f708a5b
vulkan: roll bk loop in matmul for asahi linux (#24663)
xingjianll Jun 30, 2026
e495d1e
CUDA: fix Gemma E4B MTP FlashAttention (#25148)
JohannesGaessler Jun 30, 2026
931eb37
CUDA: fix get_rows_back for tables with more than 65535 rows (grid-y …
mattjallo Jun 30, 2026
799fcc0
common,server: handle bracketed IPv6 literals in URL authority (#25140)
ServeurpersoCom Jun 30, 2026
4f31eed
model : register t_layer_inp for qwen3next (#25141)
jschmied Jun 30, 2026
0eca4d4
cuda : prevent integer truncation and overflow errors when using KQ m…
fairydreaming Jun 30, 2026
fd1a057
opencl: initial q1_0 support (#25160)
lhez Jul 1, 2026
7af4279
ui: Remove PWA navigate fallback to prevent caching API endpoint requ…
allozaur Jul 1, 2026
9d88e7c
ui Prevent tool messages from incorrectly appending to other conversa…
allozaur Jul 1, 2026
6dbc117
ggml-cpu: add AVX2 optimization for nvfp4 dot product and use UE4M3 L…
ragz4125 Jul 1, 2026
b820cc8
CUDA: consistent use of __restrict__ + PDL for FA (#25185)
JohannesGaessler Jul 1, 2026
13e6738
hexagon: flash attention rework (optimizations, accuracy improvements…
max-krasnyansky Jul 1, 2026
a6647b1
common : use hf primary split as model path (#25194)
angt Jul 1, 2026
4fc4ec5
opencl: allow loading precompiled binary kernels from library (#23042)
lhez Jul 1, 2026
fdb1db8
llama : add llama_model_ftype_name() (#25134)
angt Jul 2, 2026
c8ae9a7
vendor : update cpp-httplib to 0.49.0 (#25218)
cabelo Jul 3, 2026
5a460de
Remove redundant CUDA copies after gated_delta_net. (#23940)
gaugarg-nv Jul 3, 2026
9487528
ui: Add MCP Servers Opt-In for first time visitors (#25239)
allozaur Jul 3, 2026
b5315e1
server + ui: ping silent SSE streams every 1s and kick only after 3s …
ServeurpersoCom Jul 3, 2026
067de93
ui: align persisted config with strict server schema and enable think…
ServeurpersoCom Jul 3, 2026
75a48a9
cuda: enable topk-moe fusion for 288 experts (#25267)
pwilkin Jul 3, 2026
152d337
spec: support spec-draft-p-min in DFlash (#25246)
ruixiang63 Jul 3, 2026
f113e02
ui: strip path and weight extension from model id in single model mod…
ServeurpersoCom Jul 3, 2026
d4cff11
ui: Improve performance when streaming (#25225)
ntowle Jul 3, 2026
2d97363
chat: trim messages sent to StepFun parser (fixes long reasoning loop…
pwilkin Jul 3, 2026
ef2d770
ggml : fix broken CPU concat implementation for quantized types (#25247)
fairydreaming Jul 4, 2026
6658925
ui: add sync blocks so display/behavior settings can be set via --ui-…
ServeurpersoCom Jul 4, 2026
a410713
llama : add guard for K/V rotation input when buffer is unallocated (…
liminfei-amd Jul 4, 2026
78d2f52
cuda : concat implementation for quantized types (#25303)
fairydreaming Jul 5, 2026
7a63fde
ggml: Update VMM Pool allocation ggml-cuda.cu - Turing P2P access fix…
VexxieCode Jul 5, 2026
4b2a0cd
ggml : fix tensor-parallel + -ncmoe crash on MoE models (#25028)
liminfei-amd Jul 5, 2026
3e5036f
abort if we see a multi buffer (#25276)
netrunnereve Jul 5, 2026
2da6686
Fix stale tensor-split params for draft models (#24814)
RedToasty Jul 5, 2026
72874f5
ggml-cuda: optimize conv_transpose_1d indexing (#25310)
adavyas Jul 6, 2026
898b088
ui: fake 200 for proxy DELETE req (#25298)
ngxson Jul 6, 2026
d06ddd3
ggml-hip: enable -ffast-math for HIP builds (#23862)
a-huk Jul 6, 2026
4871961
scripts : use HF_TOKEN when downloading UI assets (#25280)
angt Jul 6, 2026
d80e878
ui: restore Ctrl+B sidebar toggle shortcut (#25307)
ServeurpersoCom Jul 6, 2026
86961ef
vulkan: fix 32-bit integer overflow in CEIL_DIV (#25245)
hokanosekai Jul 6, 2026
3b4fca1
ggml-cpu: Enable tiled matmul on AIX (#25199)
shalinib-ibm Jul 6, 2026
20a04b2
ggml-cpu: use UE4M3 LUT in ARM NVFP4 dot product (#25331)
ragz4125 Jul 6, 2026
bfdf581
server: temporary skip model downloading API test (#25355)
ngxson Jul 6, 2026
cb295bf
CUDA: extend K-type validation to V-types for flash attention (#24403)
sanmai Jul 6, 2026
9abce74
server: fix deadlock in load_models() when erasing a finished downloa…
ServeurpersoCom Jul 6, 2026
74976e1
CUDA: remove -sm row, refactor cuBLAS (#24216)
JohannesGaessler Jul 6, 2026
f36e5c3
metal: add col2im_1d op (f32/f16/bf16) (#25176)
ServeurpersoCom Jul 6, 2026
ee445f9
common: Set optimal default thread count for ppc ( linux as well as A…
shalinib-ibm Jul 6, 2026
6f8895f
opencl: general flash attention decode performance optimizations (#25…
wanghqc Jul 7, 2026
a8cfdbb
vulkan : check src0 type in GGML_OP_SET_ROWS to avoid failures due to…
fairydreaming Jul 7, 2026
defa95c
speculative : fix out-of-bounds read in ngram-map on prompt shrink (#…
o7si Jul 7, 2026
1a7c25b
ggml : make ggml_time_init idempotent (#24422)
aisk Jul 7, 2026
26145b3
sycl : rename the env vars from "disable" to "enable" (#25042)
arthw Jul 7, 2026
3d4cbdf
sycl : use sycl func to fix AOT double type issue (#25081)
arthw Jul 7, 2026
9e5ef0d
sycl : enhance argsort to support all UT cases (#25125)
arthw Jul 7, 2026
95e5254
[SYCL] fix unsupport ACC UT cases for noncontiguous (#25124)
arthw Jul 7, 2026
d209086
sycl : set K_QUANTS_PER_ITERATION to 1 on DMMV path (#25063)
malsbat Jul 7, 2026
55edb2d
[SYCL] support OP cross_entropy_loss, cross_entropy_loss_back (#25236)
arthw Jul 7, 2026
47e1de7
[SYCL] support op col2im_1d (#25264)
arthw Jul 7, 2026
108f186
[SYCL] fix unsupported UT cases of CONT & CPY (#25231)
arthw Jul 7, 2026
024c46a
llama: fix quantized kv-cache for dsv4 (#25202)
am17an Jul 7, 2026
33ca0dc
ggml-hip : add -fno-finite-math-only alongside -ffast-math (#25373)
asf0 Jul 7, 2026
c1a411f
common : add missing <fstream> include in common.h (#25220)
zhangrunda Jul 7, 2026
6c487e2
server: enforce prompt cache RAM limit (#25070)
tarruda Jul 7, 2026
5eca4e3
server : add timings and progress to /responses API stream (#25348)
thom-dev-fr Jul 7, 2026
f5525f7
server : fix draft model fit vs load inconsistency (#25056)
wadealexc Jul 7, 2026
3899b39
CUDA: Fuse MMVQ post-scale for NVFP4 (#24481)
ORippler Jul 7, 2026
c198af4
spec : fix naming, spacing (#25410)
ggerganov Jul 7, 2026
bec4772
Add Q2_0 quantization: type definition and CPU backend (#24448)
khosravipasha Jul 7, 2026
931ca30
opencl: fix potential crash in aos reconstruct (#25383)
lhez Jul 8, 2026
68a521b
ggml : add support for CPU f16->f16 GGML_OP_SET_ROWS (#25344)
fairydreaming Jul 8, 2026
57b50e1
ggml : fix A indexing in simd_gemm scalar tail-column path (#25390)
tyronecai Jul 8, 2026
4a7ee31
fix: OOB reads in UGM tokenizer (precompiled_charsmap handling) (#18750)
hourhl Jul 8, 2026
0512ef1
metal : add set_rows with src0 f16 (#25434)
fairydreaming Jul 8, 2026
da46e59
llama-eval : fix crash when answer is None in HTML dump (#25435)
ggerganov Jul 8, 2026
f1161b1
ui: Context usage gauge and panel (#25340)
allozaur Jul 8, 2026
f296fdf
common: auto-create prompts-log-dir at argument parsing, so all tools…
rankaiyx Jul 8, 2026
230ea9d
llama-batch: add n_keep_tail in split_equal for recurrent models (#25…
am17an Jul 8, 2026
bbebeec
server-stream: follow-up on SSE Replay Buffer (#23226) (#25047)
ServeurpersoCom Jul 8, 2026
90e0f5c
llama: refactor fused ops (#24646)
am17an Jul 8, 2026
ed8c261
cuda : add support for f16->f16 GGML_OP_SET_ROWS (#25367)
fairydreaming Jul 8, 2026
07e012a
Make hip quality check run on all changes (#25403)
ORippler Jul 8, 2026
c264f65
cli : move to HTTP-based implementation (#24948)
ngxson Jul 8, 2026
81ff7ab
hexagon: new vtcm layouts and improved pipelines for MUL_MAT, MUL_MAT…
max-krasnyansky Jul 8, 2026
0bbc87b
vulkan: for small AMD GPUs, reduce submission threshold based on CU c…
0cc4m Jul 8, 2026
1ee0939
llama-batch: fix allowed decreasing pos in a seq (#25449)
am17an Jul 8, 2026
167d057
opencl: ragged-tile MoE prefill FP16 GEMM optimization (skip padded e…
wanghqc Jul 8, 2026
a646006
vulkan: disable FA mask_opt on GCN to improve performance (#24362)
0cc4m Jul 8, 2026
92366df
opencl: Q6_K GEMM/GEMV fix for ne01 of weights that are not multiple…
wanghqc Jul 8, 2026
32e41fa
ggml-webgpu: tune subgroup split (d_split) in flash_attn_vec (#25418)
yomaytk Jul 8, 2026
f2d1c2f
hexagon: add VISION RoPE support (#25216)
aparmp-quic Jul 9, 2026
64c8b7d
server : respect min-step when splitting prompt batches (#25420)
aldehir Jul 9, 2026
2021515
cuda: align snake fusion matcher with the other backends (#25460)
ServeurpersoCom Jul 9, 2026
ccb0c34
ggml-hip: enable -funsafe-math-optimizations (#24668)
RapidMark Jul 9, 2026
92b187c
metal : add CONV_2D_DW (depthwise convolution) support (#21565)
Sou-ly Jul 9, 2026
259f2e2
llama-bench : init params.offline (#25476)
angt Jul 9, 2026
683f0c7
Only index by compile times + always multiply/add (#25445)
ORippler Jul 9, 2026
f84a519
Refactor: Consistently use smart pointers in `test-backend-ops` (#25440)
ORippler Jul 9, 2026
c15c5c7
meta: add hard emphasis on agents not writing descriptions/comments (…
pwilkin Jul 9, 2026
5c3a586
ggml : fix conv 2d dw (#25490)
ggerganov Jul 9, 2026
82fce65
server : move chat-template thinking probe inside the init try/catch …
palios-taey Jul 9, 2026
fb30ba9
hexagon: tiling, tracing and optimizations for unary ops (#25474)
aparmp-quic Jul 9, 2026
3de7dd4
cli: add --output option (#25484)
ngxson Jul 9, 2026
0749449
ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_a…
fairydreaming Jul 9, 2026
049326a
opencl: cluster-parallel decode FA for Adreno (#25473)
wanghqc Jul 9, 2026
7dea0c3
Merge remote-tracking branch 'upstream/master' into jimwu.gfx11-sync-…
Jul 9, 2026
d2312fc
ci(gfx11): skip empty verbose stop chunk in test-gfx answer parse
Jul 10, 2026
e583071
ci(hip): allowlist renamed Q2_K mmq + rwkv_wkv_f32 VGPR spills
Jul 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
4 changes: 2 additions & 2 deletions .devops/cann.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,7 @@ ENTRYPOINT ["/app/tools.sh"]
# ==============================================================================
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

ENTRYPOINT [ "/app/llama-cli" ]

Expand All @@ -156,7 +156,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

HEALTHCHECK --interval=5m CMD [ "curl", "-f", "http://localhost:8080/health" ]

Expand Down
4 changes: 2 additions & 2 deletions .devops/cpu.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -115,7 +115,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
4 changes: 2 additions & 2 deletions .devops/cuda.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -113,7 +113,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -124,7 +124,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
4 changes: 2 additions & 2 deletions .devops/intel.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -141,7 +141,7 @@ ENTRYPOINT ["/app/tools.sh"]
FROM base AS light

COPY --from=build /app/lib/ /app
COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -153,7 +153,7 @@ FROM base AS server
ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/lib/ /app
COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
4 changes: 2 additions & 2 deletions .devops/musa.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -126,7 +126,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
16 changes: 8 additions & 8 deletions .devops/openvino.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
ARG OPENVINO_VERSION_MAJOR=2026.2
ARG OPENVINO_VERSION_FULL=2026.2.0.21903.52ddc073857
ARG OPENVINO_VERSION_MAJOR=2026.2.1
ARG OPENVINO_VERSION_FULL=2026.2.1.21919.ede283a88e3
ARG UBUNTU_VERSION=24.04

# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
ARG IGC_VERSION=v2.34.4
ARG IGC_VERSION_FULL=2_2.34.4+21428
ARG COMPUTE_RUNTIME_VERSION=26.18.38308.1
ARG COMPUTE_RUNTIME_VERSION_FULL=26.18.38308.1-0
ARG IGC_VERSION=v2.36.3
ARG IGC_VERSION_FULL=2_2.36.3+21719
ARG COMPUTE_RUNTIME_VERSION=26.22.38646.4
ARG COMPUTE_RUNTIME_VERSION_FULL=26.22.38646.4-0
ARG IGDGMM_VERSION=22.10.0

# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
Expand Down Expand Up @@ -214,7 +214,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app/
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app/

WORKDIR /app

Expand All @@ -225,7 +225,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app/
COPY --from=build /app/full/llama /app/full/llama-server /app/

WORKDIR /app

Expand Down
4 changes: 2 additions & 2 deletions .devops/rocm.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -138,7 +138,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
4 changes: 2 additions & 2 deletions .devops/s390x.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,7 @@ WORKDIR /llama.cpp/bin

# Copy llama.cpp binaries and libraries
COPY --from=collector /llama.cpp/bin/*.so /llama.cpp/bin
COPY --from=collector /llama.cpp/bin/llama-cli /llama.cpp/bin/llama-completion /llama.cpp/bin
COPY --from=collector /llama.cpp/bin/llama /llama.cpp/bin/llama-cli /llama.cpp/bin/llama-completion /llama.cpp/bin

ENTRYPOINT [ "/llama.cpp/bin/llama-cli" ]

Expand All @@ -138,7 +138,7 @@ WORKDIR /llama.cpp/bin

# Copy llama.cpp binaries and libraries
COPY --from=collector /llama.cpp/bin/*.so /llama.cpp/bin
COPY --from=collector /llama.cpp/bin/llama-server /llama.cpp/bin
COPY --from=collector /llama.cpp/bin/llama /llama.cpp/bin/llama-server /llama.cpp/bin

EXPOSE 8080

Expand Down
4 changes: 2 additions & 2 deletions .devops/vulkan.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -107,7 +107,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -118,7 +118,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
4 changes: 2 additions & 2 deletions .devops/zendnn.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,7 @@ ENTRYPOINT ["/app/tools.sh"]
### Light, CLI only
FROM base AS light

COPY --from=build /app/full/llama-cli /app/full/llama-completion /app
COPY --from=build /app/full/llama /app/full/llama-cli /app/full/llama-completion /app

WORKDIR /app

Expand All @@ -108,7 +108,7 @@ FROM base AS server

ENV LLAMA_ARG_HOST=0.0.0.0

COPY --from=build /app/full/llama-server /app
COPY --from=build /app/full/llama /app/full/llama-server /app

WORKDIR /app

Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/build-cache.yml
Original file line number Diff line number Diff line change
Expand Up @@ -68,8 +68,8 @@ jobs:

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2"
OPENVINO_VERSION_FULL: "2026.2.0.21903.52ddc073857"
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"

steps:
- name: Clone
Expand All @@ -96,8 +96,8 @@ jobs:

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2"
OPENVINO_VERSION_FULL: "2026.2.0.21903.52ddc073857"
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"

steps:
- name: Clone
Expand Down
5 changes: 4 additions & 1 deletion .github/workflows/build-gfx11-rocm.yml
Original file line number Diff line number Diff line change
Expand Up @@ -444,7 +444,10 @@ jobs:
# The parsed assistant answer ("content" field, printed with -v) must
# contain 4 and no other digit, so "4" / "The answer is 4." pass while
# "5", "14", "22" fail. This is the deterministic correctness check.
answer_content=$(grep -oE '"content":"[^"]*"' "$output_file" | tail -1)
# llama-cli now runs via an in-process server (upstream #24948); its
# stream ends with a verbose stop chunk carrying an empty "content":"",
# so drop empty matches before taking the last real answer chunk.
answer_content=$(grep -oE '"content":"[^"]*"' "$output_file" | grep -v '"content":""' | tail -1)
if echo "$answer_content" | grep -qE '4' && ! echo "$answer_content" | grep -qE '[0-35-9]'; then
found_answer=true
fi
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/build-openvino.yml
Original file line number Diff line number Diff line change
Expand Up @@ -39,8 +39,8 @@ jobs:

env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2"
OPENVINO_VERSION_FULL: "2026.2.0.21903.52ddc073857"
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"

steps:
- name: Clone
Expand Down Expand Up @@ -96,8 +96,8 @@ jobs:

env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2"
OPENVINO_VERSION_FULL: "2026.2.0.21903.52ddc073857"
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"

steps:
- name: Clone
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/build-self-hosted.yml
Original file line number Diff line number Diff line change
Expand Up @@ -266,8 +266,8 @@ jobs:

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2"
OPENVINO_VERSION_FULL: "2026.2.0.21903.52ddc073857"
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"

steps:
- name: Clone
Expand Down
4 changes: 4 additions & 0 deletions .github/workflows/hip-quality-check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,8 @@ on:
'.github/workflows/hip-quality-check.yml',
'**/*.cu',
'**/*.cuh',
'ggml/src/ggml-hip/CMakeLists.txt',
'ggml/src/ggml-cuda/vendors/hip.h',
'scripts/hip/gcn-cdna-vgpr-check.py'
]

Expand All @@ -18,6 +20,8 @@ on:
'.github/workflows/hip-quality-check.yml',
'**/*.cu',
'**/*.cuh',
'ggml/src/ggml-hip/CMakeLists.txt',
'ggml/src/ggml-cuda/vendors/hip.h',
'scripts/hip/gcn-cdna-vgpr-check.py'
]

Expand Down
Loading
Loading