Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
108 commits
Select commit Hold shift + click to select a range
933f46f
ggml : bump version to 0.19.0 (ggml/1581)
ggerganov Aug 7, 2026
4cf5cab
sync : ggml
ggerganov Aug 7, 2026
4cb22cd
mtmd: fix longest_edge ignoring min/max pixels (#26638)
ngxson Aug 7, 2026
2363478
ui: Filesystem `@mentions` for Chat Form (#26715)
ServeurpersoCom Aug 7, 2026
a194a75
metal : fix NORM/RMS_NORM for row lengths that leave a partial simdgr…
robertomeroni Aug 7, 2026
f8e3026
sycl: coalesce the ssm_conv window loads (#26612)
Titaniumtown Aug 7, 2026
6de1b63
allozaur/feat/chat slash commands (#26716)
ServeurpersoCom Aug 7, 2026
1621a3d
tests : speed-up server test suite 3x (#26734)
ggerganov Aug 7, 2026
fc6545d
allozaur/feat/chat form contenteditable (#26717)
ServeurpersoCom Aug 7, 2026
3653e6d
tts: account for the vocoder pass in the timings line (#26733)
ServeurpersoCom Aug 7, 2026
69bf643
CUDA: fix thread/block count in quantized cpy kernel launches (#26731)
grafail Aug 8, 2026
dd2c7c4
server: add initial tool isolation support (via docker) (#26507)
ngxson Aug 8, 2026
18f7ad7
server, ui: only offer a working directory when a tool reads it (#26762)
ServeurpersoCom Aug 8, 2026
687e778
CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767)
grafail Aug 8, 2026
7ba604f
server: report the isolate working directory from get_info (#26773)
ServeurpersoCom Aug 8, 2026
61141f1
ci: rm `GGML_HIP_ROCWMMA_FATTN` (#26760)
taronaeo Aug 9, 2026
0865990
ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792)
Hao-Chen2337 Aug 9, 2026
9369185
ci: add pr-draft-label (#26801)
ngxson Aug 9, 2026
74ce157
ui: degrade the working directory picker when file search is off (#26…
ServeurpersoCom Aug 9, 2026
f401bb1
ggml-webgpu : refactor several wgsl files and simplify flash_attn wgs…
yomaytk Aug 10, 2026
aea252f
ci: fix the ctest sanitize runs (#26593)
netrunnereve Aug 10, 2026
0377426
model-saver : fix expert shared/chunk FFN length key clobber (#26693)
SolshineCode Aug 10, 2026
1e396e7
server: gate the docker tools runtime tests on a real container run (…
ServeurpersoCom Aug 10, 2026
92d1bb0
ui: Linting & Formatting scripts (#26819)
allozaur Aug 10, 2026
6ad4ab0
readme : remove dev branches (#26832)
ggerganov Aug 10, 2026
157b81f
model : Granite-Switch Architecture (#25107)
barvhaim Aug 10, 2026
e23e944
vendor : update cpp-httplib to 0.53.0 (#26821)
cabelo Aug 10, 2026
7a20b41
model: add MTP support for Nemotron model (#26725)
ruixiang63 Aug 10, 2026
2e2d99c
ci: Add support for CUDA 13.4 ARM64 builds for Windows (#26650)
shivamkumard-ctrl Aug 10, 2026
86c298f
llama: Restore quantization of mmprojs (#26818)
pcuenca Aug 10, 2026
4c6766f
vendor: sync subprocess.h and drop local patches (#26808)
ServeurpersoCom Aug 10, 2026
a52077c
chat : Align Laguna-S-2.1 chat template to huggingface (#26232)
crusaderky Aug 10, 2026
62bf73d
model: Muse Glimmer Support (#26841)
pcuenca Aug 10, 2026
4ae84de
server: add more tool isolation support (ssh remote + podman rootless…
ServeurpersoCom Aug 10, 2026
e5275f6
ci : don't specify python version in server-sanitize for broader runn…
CISC Aug 10, 2026
4dee52f
ui: UI/chat form follow ups (#26743)
ServeurpersoCom Aug 10, 2026
f8def7f
ggml : require contiguous src for ROLL on CUDA and Metal (#25928)
devYRPauli Aug 10, 2026
d2f8305
ggml-cpu : fix CPU affinity mask being ignored on Android (#26838)
hiteshchopra11 Aug 10, 2026
dd1ea52
llama : support multi-output backend sampling (#25532)
gaugarg-nv Aug 10, 2026
0666ad2
ci : target ROCm 7.14 for build and release (#25775)
superm1 Aug 10, 2026
689e227
opencl: transpose the K tile in local memory for FA prefill kernels (…
wanghqc Aug 10, 2026
030ebb5
Address review comment of PR 25532 (#26852)
gaugarg-nv Aug 10, 2026
84f7129
ggml-webgpu: fix CI errors from #25025 and #25262 (#26566)
yomaytk Aug 11, 2026
48d22e2
common/peg : suppress incomplete escape sequences (#26780)
aldehir Aug 11, 2026
14e78dd
model : fix SWA not being enabled for EXAONE 4.5 (#26848)
junmo-kim Aug 11, 2026
4801e3c
tests : disable backend sampler hip multi output (#26878)
jimw567 Aug 11, 2026
b3df572
tests : clean-up server test, use `tests.sh` in ci (#26886)
ggerganov Aug 11, 2026
153d324
llama: add default load-mode auto, which avoids mmap on iGPUs (#26081)
0cc4m Aug 11, 2026
9afff1b
tests : fix running server tests on windows (#26889)
ggerganov Aug 11, 2026
1138b85
model-conversion : use save_output_data for causual embeddings [no ci…
danbev Aug 11, 2026
7044859
ci: hip-quality-check: update vgpr spill ignore list (#26859)
IMbackK Aug 11, 2026
8d274dd
ui: fix context gauge for single-model usage (#25738)
intel00000 Aug 11, 2026
6e62ba5
mtmd: support pocket-tts (#26871)
ngxson Aug 11, 2026
cc078b4
Dflash support for nemotron-3.5 (#26905)
lnigam Aug 11, 2026
5d16e81
convert : keep quantization scales for nemotron --mtp export (#26903)
ynankani Aug 11, 2026
2468576
requirements: use stable torch packages on s390x (#26864)
nikwen Aug 11, 2026
38406d5
imatrix.cpp: Move finite check and only check touched experts (#26861)
bartowski1182 Aug 11, 2026
70dfba5
ci : add windows-rocm to check-release (#26897)
CISC Aug 11, 2026
ba360ef
chat : tighten bare function parsing for Qwen models (#26793)
aldehir Aug 11, 2026
f785fc9
spec : update speculative-simple (#26904)
ggerganov Aug 11, 2026
5988633
cuda : add warp-per-row wkv7 kernel for single-token decode (#26111)
123123213weqw Aug 11, 2026
ebb546b
CUDA: only disable CUDA graphs when mul_mat_id actually needs a strea…
grafail Aug 11, 2026
7b13a84
ci : add missing release check (#26923)
CISC Aug 11, 2026
0b1bad1
chat : fix muse-glimmer detection of tool calls after EOM (#26879)
ruanslv Aug 11, 2026
cb27fe9
opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (…
lhez Aug 12, 2026
6eff593
convert : handle per_layer_config in Gemma4 (transformers 5.15) (#26882)
pluvium27 Aug 12, 2026
55f453b
wavtokenizer-dec : bound posnet/convnext block_count against n_layer_…
oakkaya Aug 12, 2026
a7cd2f0
vulkan: add TQ2_0 (ternary) support (#25850)
michaeltrabalka-tech Aug 12, 2026
a4a4c51
tests : update speculative params (#26925)
ggerganov Aug 12, 2026
89e0aa6
opencl: default FA c8 cluster width to 16 on X1E (#26433)
wanghqc Aug 12, 2026
4dd1275
ui: add read_media tool (#25877)
parabelboi Aug 12, 2026
5d9e5ac
server : support slot save/restore with media inputs (#26640)
CHIPMUNK-T0T Aug 12, 2026
13fd0bb
cmake : add config version support (ggml/1582)
danbev Aug 12, 2026
af05a42
sync : ggml
ggerganov Aug 12, 2026
ece98b8
model : disallow integer dflash sliding_window_pattern (#26900)
CISC Aug 12, 2026
132753b
kleidiai: Add runtime feature detection mechanism for aarch64/kleidia…
JonathanC-ARM Aug 12, 2026
d8a8bea
gguf : harden loader against malformed tensor dims and metadata types…
harrison001 Aug 12, 2026
680a9ae
cmake : introduce semantic versioning (#26839)
danbev Aug 12, 2026
7a9ff95
disable rocm cache (#26962)
ggerganov Aug 12, 2026
9558fa4
ci : disable ubuntu-rocm (#26969)
ggerganov Aug 12, 2026
84e908c
ci: fix thread sanitizer + remove ccache (#26927)
netrunnereve Aug 12, 2026
8e7f22b
common: add system-level config file (#26118)
jcmdln Aug 12, 2026
e21152d
ui: Constants refactor (#26908)
ServeurpersoCom Aug 13, 2026
1f368f3
ggml : fix arm builds, unused var (#26991)
ggerganov Aug 13, 2026
a6040c9
refactor: Clean up UI types (#26909)
allozaur Aug 13, 2026
094e53d
ui: Stores architecture improvements (#26910)
allozaur Aug 13, 2026
f2efd64
ui: Move `styles/` to `$lib` scope (#26950)
allozaur Aug 13, 2026
d86c7d6
ui: Clean up contexts, remove prop drilling from Chat Form Actions (#…
allozaur Aug 13, 2026
e79e4bf
ggml-hip : remove -funsafe-math-optimizations (#26696)
jimw567 Aug 13, 2026
decaf50
server: refactor + correctness fixes for metrics (#26920)
ngxson Aug 13, 2026
d415e65
sycl : enhance concat to support Q4_0, Q4_1, Q5_0, Q5_1, Q8_0 (#26800)
arthw Aug 13, 2026
8efbf65
sycl : Add DMMV ESIMD Q3_K kernel (#26251)
malsbat Aug 13, 2026
1ee1cd9
sycl: fuse UNARY(silu|sigmoid|softplus) + MUL (#26411)
Titaniumtown Aug 13, 2026
154d57a
sycl: remove separate fp32 type promotion in gemm non-oneDNN path (#2…
icfaust Aug 13, 2026
eeae28b
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion…
jinzihao Aug 13, 2026
0d0bfcd
spec: enable backend sampling for both dflash & dspark (#26958)
ruixiang63 Aug 13, 2026
f65e568
common : auto-detect spec type from draft GGUF metadata (#26814)
aic0d3r Aug 13, 2026
4a84b0a
metal : add TQ2_0 support (#26980)
ggerganov Aug 13, 2026
1d2869c
spec : auto-detect mtp draft model type (#27005)
CISC Aug 13, 2026
981184e
server : serve index.html with no-cache (#27006)
erusev Aug 13, 2026
2606220
chat : fix LFM2 tool call arg name prefix ambiguity (#26960)
aldehir Aug 13, 2026
a97123e
[SYCL] Support host pinned mem to improve SYCL Host-to-Device Memory …
arthw Aug 13, 2026
aee56b3
OpenVINO: Qwen3.5, memory optimization, and test-recurrent-state-roll…
wine99 Aug 13, 2026
9c5531e
ui: fix VITE_PUBLIC_SERVER variable reading (#24845)
brainrom Aug 13, 2026
fa4ec45
refactor: Naming (#27001)
allozaur Aug 13, 2026
bdffafa
ui: Refactor data-attrs constants, enum for bool strings (#27002)
ServeurpersoCom Aug 13, 2026
a94d563
common: apply CPU parameters across tools (#27026)
nikwen Aug 13, 2026
de8775a
Merge remote-tracking branch 'upstream/master' into jimwu.sync-upstre…
Aug 13, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
1 change: 0 additions & 1 deletion .devops/rocm.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,6 @@ COPY --from=web /app/tools/ui/dist tools/ui/dist
RUN HIPCXX="$(hipconfig -l)/clang" HIP_PATH="$(hipconfig -R)" \
cmake -S . -B build \
-DGGML_HIP=ON \
-DGGML_HIP_ROCWMMA_FATTN=ON \
-DAMDGPU_TARGETS="$ROCM_DOCKER_ARCH" \
-DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON \
-DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_TESTS=OFF \
Expand Down
27 changes: 27 additions & 0 deletions .github/actions/windows-setup-cuda/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,10 @@ inputs:
cuda_version:
description: "CUDA toolkit version"
required: true
cuda_arch:
description: "CUDA target architecture"
required: false
default: "x64"

runs:
using: "composite"
Expand Down Expand Up @@ -127,3 +131,26 @@ runs:
echo "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.3\bin" | Out-File -FilePath $env:GITHUB_PATH -Encoding utf8 -Append
echo "CUDA_PATH=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.3" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8
echo "CUDA_PATH_V13_3=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.3" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8

- name: Install Cuda Toolkit 13.4 for ARM64
if: ${{ inputs.cuda_version == '13.4' && inputs.cuda_arch == 'arm64' }}
shell: pwsh
run: |
mkdir -p "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4"
choco install unzip -y
curl -O "https://packages.nvidia.com/bin-archive/pool/windows-x86_64/5B515474-7E78-11F1-8656-C51E4F4B317F/cccl-windows-x86_64-13.3.4.1.2-archive.zip"
curl -O "https://packages.nvidia.com/bin-archive/pool/windows-x86_64/5B515474-7E78-11F1-8656-C51E4F4B317F/cuda_crt-windows-x86_64-13.4.46-archive.zip"
curl -O "https://packages.nvidia.com/bin-archive/pool/windows-x86_64/5B515474-7E78-11F1-8656-C51E4F4B317F/cuda_nvcc-windows-x86_64-13.4.46-archive.zip"
curl -O "https://packages.nvidia.com/bin-archive/pool/windows-x86_64/5B515474-7E78-11F1-8656-C51E4F4B317F/libnvvm-windows-x86_64-13.4.46-archive.zip"
curl -O "https://packages.nvidia.com/bin-archive/pool/windows-arm64/5B515474-7E78-11F1-8656-C51E4F4B317F/cuda_cudart-windows-arm64-13.4.46-archive.zip"
curl -O "https://packages.nvidia.com/bin-archive/pool/windows-arm64/5B515474-7E78-11F1-8656-C51E4F4B317F/libcublas-windows-arm64-13.7.0.10-archive.zip"
unzip '*.zip' -d "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4"
xcopy "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\cccl-windows-x86_64-13.3.4.1.2-archive\*" "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" /E /I /H /Y
xcopy "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\cuda_crt-windows-x86_64-13.4.46-archive\*" "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" /E /I /H /Y
xcopy "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\cuda_nvcc-windows-x86_64-13.4.46-archive\*" "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" /E /I /H /Y
xcopy "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\libnvvm-windows-x86_64-13.4.46-archive\*" "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" /E /I /H /Y
xcopy "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\cuda_cudart-windows-arm64-13.4.46-archive\*" "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" /E /I /H /Y
xcopy "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\libcublas-windows-arm64-13.7.0.10-archive\*" "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" /E /I /H /Y
echo "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4\bin" | Out-File -FilePath $env:GITHUB_PATH -Encoding utf8 -Append
echo "CUDA_PATH=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8
echo "CUDA_PATH_V13_4=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8
28 changes: 23 additions & 5 deletions .github/actions/windows-setup-rocm/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,26 @@ inputs:
runs:
using: "composite"
steps:
- name: Setup ROCm
uses: ./.github/actions/install-exe
with:
url: https://download.amd.com/developer/eula/rocm-hub/AMD-Software-PRO-Edition-${{ inputs.version }}-Win11-For-HIP.exe
args: -install
- name: Install ROCm with Wheels
shell: pwsh
run: |
$ErrorActionPreference = "Stop"
write-host "Setting up Python virtual environment"

# Create the venv directly at the cache location to avoid relocation issues
New-Item -Path "C:\TheRock\build" -ItemType Directory -Force | Out-Null
python -m venv C:\TheRock\build\.venv
& C:\TheRock\build\.venv\Scripts\Activate.ps1

write-host "Upgrading pip"
python -m pip install --upgrade pip

write-host "Installing ROCm wheels for multi-arch support"
# Install ROCm wheels for multi-arch support (this may take several minutes)
python -m pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "rocm[libraries,devel]==${{ inputs.version }}"

# Pre-expand the devel tree so it is included in the cache
write-host "Initializing ROCm devel tree"
rocm-sdk init
if ($LASTEXITCODE -ne 0) { throw "rocm-sdk init failed with exit code $LASTEXITCODE" }
write-host "Completed ROCm wheel installation to C:\TheRock\build"
48 changes: 24 additions & 24 deletions .github/workflows/build-cache.yml
Original file line number Diff line number Diff line change
Expand Up @@ -119,27 +119,27 @@ jobs:
version_major: ${{ env.OPENVINO_VERSION_MAJOR }}
version_full: ${{ env.OPENVINO_VERSION_FULL }}

windows-2022-rocm-cache:
runs-on: windows-2022

env:
# Make sure this is in sync with build.yml
HIPSDK_INSTALLER_VERSION: "26.Q1"

steps:
- name: Clone
id: checkout
uses: actions/checkout@v6

- name: Setup Cache
uses: actions/cache@v5
id: cache-rocm
with:
path: C:\Program Files\AMD\ROCm
key: cache-gha-rocm-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ runner.os }}

- name: Setup ROCm
if: steps.cache-rocm.outputs.cache-hit != 'true'
uses: ./.github/actions/windows-setup-rocm
with:
version: ${{ env.HIPSDK_INSTALLER_VERSION }}
# windows-2022-rocm-cache:
# runs-on: windows-2022

# env:
# # Make sure this is in sync with release.yml and build-cuda-windows.yml
# ROCM_VERSION: "7.14.0"

# steps:
# - name: Clone
# id: checkout
# uses: actions/checkout@v6

# - name: Setup Cache
# uses: actions/cache@v5
# id: cache-rocm
# with:
# path: C:\TheRock\build
# key: rocm-wheels-${{ env.ROCM_VERSION }}-multi-arch-${{ runner.os }}

# - name: Setup ROCm
# if: steps.cache-rocm.outputs.cache-hit != 'true'
# uses: ./.github/actions/windows-setup-rocm
# with:
# version: ${{ env.ROCM_VERSION }}
14 changes: 10 additions & 4 deletions .github/workflows/build-cmake-pkg.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ on:

jobs:
linux:
runs-on: [self-hosted, Linux, CPU]
runs-on: [self-hosted, Linux]
steps:
- uses: actions/checkout@v6
with:
Expand All @@ -21,15 +21,21 @@ jobs:
-DLLAMA_BUILD_TOOLS=OFF \
-DLLAMA_BUILD_EXAMPLES=OFF \
-DLLAMA_BUILD_APP=OFF \
-DLLAMA_BUILD_IS_DEV=OFF \
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release
cmake --build build --config Release -j $(nproc)
cmake --install build --prefix "$PREFIX" --config Release

export LLAMA_CONFIG="$PREFIX"/lib/cmake/llama/llama-config.cmake
tclsh <<'EOF'
set build(commit) [string trim [exec git rev-parse --short HEAD]]
set build(number) [string trim [exec git rev-list --count HEAD]]
set build(version) "0.0.$build(number)"

set cmakelists [read [open "CMakeLists.txt" r]]
regexp {set\(LLAMA_VERSION_MAJOR\s+(\d+)\)} $cmakelists -> major
regexp {set\(LLAMA_VERSION_MINOR\s+(\d+)\)} $cmakelists -> minor
regexp {set\(LLAMA_VERSION_PATCH\s+(\d+)\)} $cmakelists -> patch
set build(version) "$major.$minor.$patch"

set llamaconfig [read [open "$env(LLAMA_CONFIG)" r]]
set checks [list "set\\(LLAMA_VERSION \\s+$build(version)\\)" \
Expand All @@ -48,4 +54,4 @@ jobs:

cd examples/simple-cmake-pkg
cmake -S . -B build -DCMAKE_PREFIX_PATH="$PREFIX"/lib/cmake
cmake --build build
cmake --build build -j $(nproc)
4 changes: 3 additions & 1 deletion .github/workflows/build-cpu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -94,8 +94,10 @@ jobs:
id: cmake_build
run: |
cmake -B build \
-DGGML_NATIVE=OFF \
-DLLAMA_FATAL_WARNINGS=ON \
-DGGML_RPC=ON
-DGGML_RPC=ON \
-DGGML_NATIVE=OFF
time cmake --build build --config Release -j $(nproc)
- name: Test
Expand Down
1 change: 0 additions & 1 deletion .github/workflows/build-cuda-ubuntu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -99,7 +99,6 @@ jobs:
run: |
cmake -B build -S . \
-DCMAKE_HIP_COMPILER="$(hipconfig -l)/clang" \
-DGGML_HIP_ROCWMMA_FATTN=ON \
-DGPU_TARGETS="gfx1030" \
-DGGML_HIP=ON
cmake --build build --config Release -j $(nproc)
Expand Down
85 changes: 50 additions & 35 deletions .github/workflows/build-cuda-windows.yml
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,7 @@ jobs:

env:
# Make sure this is in sync with build-cache.yml
HIPSDK_INSTALLER_VERSION: "26.Q1"
ROCM_VERSION: "7.14.0"

strategy:
matrix:
Expand All @@ -97,66 +97,81 @@ jobs:
id: checkout
uses: actions/checkout@v6

- name: Grab rocWMMA package
id: grab_rocwmma
run: |
curl -o rocwmma.deb "https://repo.radeon.com/rocm/apt/7.2.1/pool/main/r/rocwmma-dev/rocwmma-dev_2.2.0.70201-81~24.04_amd64.deb"
7z x rocwmma.deb
7z x data.tar

- name: Use ROCm Installation Cache
uses: actions/cache@v5
id: cache-rocm
with:
path: C:\Program Files\AMD\ROCm
key: cache-gha-rocm-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ runner.os }}
# - name: Cache ROCm Installation
# uses: actions/cache@v5
# id: cache-rocm
# with:
# path: C:\TheRock\build
# key: rocm-wheels-${{ env.ROCM_VERSION }}-multi-arch-${{ runner.os }}

- name: Setup ROCm
if: steps.cache-rocm.outputs.cache-hit != 'true'
# if: steps.cache-rocm.outputs.cache-hit != 'true'
uses: ./.github/actions/windows-setup-rocm
with:
version: ${{ env.HIPSDK_INSTALLER_VERSION }}
version: ${{ env.ROCM_VERSION }}

- name: Setup ROCm Environment
run: |
$ErrorActionPreference = "Stop"

# Activate venv from cache or fresh install
& C:\TheRock\build\.venv\Scripts\Activate.ps1

# Expand the devel tree (idempotent; no-op if already done during install)
rocm-sdk init
if ($LASTEXITCODE -ne 0) { throw "rocm-sdk init failed with exit code $LASTEXITCODE" }

# Get ROCm installation paths using the rocm-sdk CLI tool
$rocmPath = (rocm-sdk path --root)
if (-not $rocmPath) { throw "rocm-sdk path --root returned empty - devel package may not be installed" }
$rocmPath = $rocmPath.Trim()
$cmakePath = (rocm-sdk path --cmake).Trim()
$binPath = (rocm-sdk path --bin).Trim()
write-host "ROCm root: $rocmPath"

echo "HIP_PATH=$rocmPath" >> $env:GITHUB_ENV
echo "CMAKE_PREFIX_PATH=$cmakePath" >> $env:GITHUB_ENV
echo "HIP_DEVICE_LIB_PATH=$rocmPath\lib\llvm\amdgcn\bitcode" >> $env:GITHUB_ENV
echo "HIP_PLATFORM=amd" >> $env:GITHUB_ENV
echo "LLVM_PATH=$rocmPath\lib\llvm" >> $env:GITHUB_ENV
echo "$binPath" >> $env:GITHUB_PATH

# Keep venv in PATH for subsequent steps
echo "C:\TheRock\build\.venv\Scripts" >> $env:GITHUB_PATH

- name: Verify ROCm
id: verify
run: |
# Find and test ROCm installation
$clangPath = Get-ChildItem 'C:\Program Files\AMD\ROCm\*\bin\clang.exe' | Select-Object -First 1
if (-not $clangPath) {
Write-Error "ROCm installation not found"
exit 1
}
& $clangPath.FullName --version
# Test the ROCm clang shipped in the installed wheel
& "${env:HIP_PATH}\lib\llvm\bin\clang.exe" --version

- name: ccache
uses: ggml-org/ccache-action@v1.2.21
with:
# TODO: this build does not match the build in release.yml, so we use a different cache key
# ideally, the builds should match, similar to the CUDA build above so that we would be able
# to populate the ccache for the release with manual runs of this workflow
#key: release-windows-2022-x64-hip-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ matrix.name }}
key: cuda-windows-2022-x64-hip-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ matrix.name }}
#key: release-windows-2022-x64-hip-${{ env.ROCM_VERSION }}-${{ matrix.name }}
key: cuda-windows-2022-x64-hip-${{ env.ROCM_VERSION }}-${{ matrix.name }}

- name: Build
id: cmake_build
run: |
$env:HIP_PATH=$(Resolve-Path 'C:\Program Files\AMD\ROCm\*\bin\clang.exe' | split-path | split-path)
$env:CMAKE_PREFIX_PATH="${env:HIP_PATH}"
cmake -G "Unix Makefiles" -B build -S . `
-DCMAKE_C_COMPILER="${env:HIP_PATH}\bin\clang.exe" `
-DCMAKE_CXX_COMPILER="${env:HIP_PATH}\bin\clang++.exe" `
-DCMAKE_CXX_FLAGS="-I$($PWD.Path.Replace('\', '/'))/opt/rocm-7.2.1/include/" `
-DCMAKE_PREFIX_PATH="${env:HIP_PATH}" `
-DCMAKE_C_COMPILER="${env:HIP_PATH}\lib\llvm\bin\clang.exe" `
-DCMAKE_CXX_COMPILER="${env:HIP_PATH}\lib\llvm\bin\clang++.exe" `
-DCMAKE_HIP_COMPILER="${env:HIP_PATH}\lib\llvm\bin\clang.exe" `
-DCMAKE_BUILD_TYPE=Release `
-DLLAMA_BUILD_BORINGSSL=ON `
-DROCM_DIR="${env:HIP_PATH}" `
-DHIP_PATH="${env:HIP_PATH}" `
-DGGML_HIP=ON `
-DGGML_HIP_ROCWMMA_FATTN=ON `
-DGPU_TARGETS="gfx1100" `
-DGPU_TARGETS="gfx1100" `
-DGGML_RPC=ON
cmake --build build -j ${env:NUMBER_OF_PROCESSORS}

- name: ccache-clear
uses: ./.github/actions/ccache-clear
with:
#key: release-windows-2022-x64-hip-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ matrix.name }}
key: cuda-windows-2022-x64-hip-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ matrix.name }}
#key: release-windows-2022-x64-hip-${{ env.ROCM_VERSION }}-${{ matrix.name }}
key: cuda-windows-2022-x64-hip-${{ env.ROCM_VERSION }}-${{ matrix.name }}
28 changes: 25 additions & 3 deletions .github/workflows/build-sanitize.yml
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,12 @@ on:
'**/*.cpp'
]

pull_request:
types: [opened, synchronize, reopened]
paths: [
'.github/workflows/build-sanitize.yml'
]

concurrency:
group: ${{ github.workflow }}-${{ github.head_ref && github.ref || github.run_id }}
cancel-in-progress: true
Expand All @@ -28,19 +34,35 @@ env:

jobs:
ctest:
runs-on: [self-hosted, X64, CPU, Linux]

continue-on-error: true

strategy:
matrix:
sanitizer: [ADDRESS, THREAD, UNDEFINED]
include:
# thread and address doesn't run properly on some self hosted machines, so run it on Github instead
- sanitizer: ADDRESS
machine: ubuntu-24.04
- sanitizer: THREAD
machine: ubuntu-24.04
- sanitizer: UNDEFINED
machine: [self-hosted, X64, Linux]

runs-on: ${{ matrix.machine }}

steps:
- name: Clone
id: checkout
uses: actions/checkout@v6

# - name: ccache
# uses: ggml-org/ccache-action@v1.2.21
# if: ${{ matrix.sanitizer != 'UNDEFINED' }}
# with:
# key: ctest-${{ matrix.sanitizer }}-ubuntu-24.04
# variant: ccache
# evict-old-files: 1d
# save: ${{ github.event_name == 'push' && github.ref == 'refs/heads/master' }}

# with UNDEFINED sanitizer, we have to build in Debug to avoid GCC 13 false-positive warnings
- name: Build (undefined)
id: cmake_build_undefined
Expand Down
Loading
Loading