Skip to content

Merge latest upstream changes in master to gfx11. - #72

Merged
jimw567 merged 192 commits into
gfx11from
lichang.gfx11-master-sync
Jul 30, 2026
Merged

Merge latest upstream changes in master to gfx11.#72
jimw567 merged 192 commits into
gfx11from
lichang.gfx11-master-sync

Conversation

@liangliangchang

@liangliangchang liangliangchang commented Jul 27, 2026

Copy link
Copy Markdown

Overview

Merge latest upstream changes in master to gfx11. Several fixes:

  1. Fixed bitwise Q1_0 unpacking on HIP to recover AMD MMQ and MMVQ performance while retaining the optimized byte-permute path elsewhere.
  2. Use two WMMA row tiles only when the column tile is aligned, preventing partial tiles from overflowing accumulators. (Crashes Gemma 8K and SmolLM2 8K)
  3. Fixed VGPR pressure check because of mmq register renaming.
  4. Recover J=16 prefill performance lost in the MMQ refactor and cover the affected Q4_K shapes in backend operation tests.
  5. Keep integrated GPU detection while staging gfx1151 buffers in device memory to avoid UMA prefill latency regressions.

Full 21-type production matrix

This table compares one latest-sync pass against one adjacent premerge pass. Each row summarizes 13 production shapes.

Quantization Geomean speedup Minimum Maximum
Q1_0 0.9978x 0.9938x 1.0032x
Q4_0 1.0020x 0.9675x 1.0285x
Q4_1 1.0410x 0.9623x 1.1282x
Q5_0 1.0096x 0.9971x 1.0429x
Q5_1 1.0686x 0.9939x 1.3270x
Q8_0 0.9991x 0.9949x 1.0014x
Q2_K 0.9992x 0.9956x 1.0094x
Q3_K 0.9991x 0.9930x 1.0068x
Q4_K 0.9898x 0.9567x 1.0057x
Q5_K 0.9977x 0.9905x 1.0048x
Q6_K 1.0001x 0.9953x 1.0037x
IQ1_S 0.9998x 0.9912x 1.0112x
IQ2_XXS 1.0038x 0.9808x 1.0252x
IQ2_XS 0.9943x 0.9745x 1.0124x
IQ2_S 0.9904x 0.9734x 1.0088x
IQ3_XXS 1.0012x 0.9837x 1.0202x
IQ3_S 1.0019x 0.9804x 1.0195x
IQ4_XS 0.9948x 0.9837x 1.0022x
IQ4_NL 0.9874x 0.9432x 1.0047x
MXFP4 0.9961x 0.9897x 1.0038x
NVFP4 0.9865x 0.9640x 1.0034x
All 273 cases 1.0027x

Standard model comparison

Test Pre decode Fixed decode Decode delta Pre TTFT ms Fixed TTFT ms TTFT delta
GLM-4.7-Flash_Q4_K_M_GGUF_128 62.06 62.18 +0.19% 127.22 125.42 -1.42%
Gemma-2B_Q4_K_M_GGUF_128 112.87 111.12 -1.55% 32.24 32.44 +0.63%
Gemma-2B_Q4_K_M_GGUF_8000 106.38 106.87 +0.46% 3700.97 3682.55 -0.50%
Gemma-3-12B-IT_Q4_K_M_GGUF_4096 24.55 24.52 -0.09% 5340.52 5281.17 -1.11%
Gemma-3-12B-IT_Q4_K_M_GGUF_VLM 26.63 26.58 -0.19% 1283.83 1282.82 -0.08%
Gemma-3-4B-IT_VLM 71.55 71.34 -0.30% 961.95 960.69 -0.13%
Gemma-4-12B-it_Q4_K_M_GGUF_128 27.47 27.42 -0.16% 154.16 153.17 -0.64%
Gemma-4-26B-A4B-IT_VLM 48.54 48.54 +0.01% 661.00 671.90 +1.65%
Gemma-4-31B-IT_VLM 11.01 10.98 -0.24% 1732.22 1760.80 +1.65%
Gemma-4-E2B-IT_Q4_K_M_GGUF_128 100.02 100.03 +0.01% 42.03 40.66 -3.28%
Gemma-4-E2B-IT_Q4_K_M_GGUF_3968 94.74 94.55 -0.20% 1679.89 1691.74 +0.71%
Janus-Pro-1B_VLM 159.47 159.35 -0.08% 164.04 162.99 -0.64%
Janus-Pro-7B_VLM 45.05 45.00 -0.11% 491.65 489.29 -0.48%
Llama-2-7B_Q4_K_M_GGUF_128 46.78 48.19 +3.00% 95.01 94.19 -0.86%
Llama-2-7B_Q4_K_M_GGUF_1920 35.12 35.16 +0.09% 1684.97 1675.74 -0.55%
Llama-3.1-8B-Instruct_Q4_K_M_GGUF_128 42.00 42.04 +0.10% 98.36 97.75 -0.62%
MiniCPM-V-2_6_VLM 45.72 45.65 -0.14% 1405.58 1404.34 -0.09%
MiniCPM-o-2_6_VLM 44.87 45.66 +1.75% 1429.03 1407.97 -1.47%
Qwen2.5-0.5B-Instruct_Q4_K_M_GGUF_128 291.67 291.75 +0.03% 12.33 12.12 -1.65%
Qwen2.5-0.5B-Instruct_Q4_K_M_GGUF_3968 285.06 285.61 +0.19% 389.54 393.73 +1.08%
Qwen2.5-3B-Instruct_Q4_K_M_GGUF_128 87.18 87.16 -0.02% 47.47 46.19 -2.69%
Qwen2.5-3B-Instruct_Q4_K_M_GGUF_3968 84.75 85.57 +0.97% 1555.32 1568.56 +0.85%
Qwen2.5-7B-Instruct_Q4_K_M_GGUF_128 46.08 46.68 +1.31% 90.45 89.97 -0.53%
Qwen2.5-7B-Instruct_Q4_K_M_GGUF_3968 43.71 43.67 -0.09% 3064.18 3055.75 -0.27%
Qwen2.5-VL-3B-Instruct_VLM 87.25 86.86 -0.44% 2083.14 2078.42 -0.23%
Qwen2.5-VL-7B-Instruct_VLM 45.19 45.09 -0.22% 2472.44 2463.99 -0.34%
Qwen3-1.7B_Q4_K_M_GGUF_128 141.45 138.83 -1.85% 27.67 27.84 +0.62%
Qwen3-1.7B_Q4_K_M_GGUF_3968 113.66 112.38 -1.13% 1144.27 1143.49 -0.07%
Qwen3-30B-A3B-Instruct-2507_Q4_K_M_GGUF_128 75.99 77.65 +2.19% 104.12 102.74 -1.33%
Qwen3-4B_Q4_K_M_GGUF_128 74.23 73.74 -0.66% 60.41 59.86 -0.91%
Qwen3-4B_Q4_K_M_GGUF_3968 61.11 62.65 +2.51% 2372.51 2363.21 -0.39%
Qwen3-8B_Q4_K_M_GGUF_128 42.31 41.79 -1.23% 99.26 99.22 -0.04%
Qwen3-8B_Q4_K_M_GGUF_3968 37.89 37.89 +0.02% 3776.69 3771.22 -0.14%
Qwen3-VL-4B-Instruct_VLM 69.98 69.75 -0.32% 872.82 868.45 -0.50%
Qwen3.5-35B-A3B_Q4_K_M_GGUF_128 53.89 54.13 +0.45% 104.20 105.53 +1.28%
Qwen3.5-4B_Q4_K_M_GGUF_128 61.66 62.07 +0.67% 68.14 67.42 -1.06%
Qwen3.5-9B_Q4_K_M_GGUF_128 36.56 37.34 +2.14% 111.88 110.47 -1.26%
Qwen3.6-27B_Q4_K_M_GGUF_4096 12.23 12.14 -0.77% 12087.62 12007.14 -0.67%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_128 59.06 59.27 +0.36% 109.80 108.57 -1.12%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_4096 57.77 58.01 +0.42% 2891.05 2885.76 -0.18%
SmolLM2-1.7B-Instruct_Q4_K_M_GGUF_128 146.89 148.97 +1.41% 30.81 30.06 -2.42%
SmolLM2-1.7B-Instruct_Q4_K_M_GGUF_8000 37.04 36.46 -1.56% 6820.09 6826.11 +0.09%

MTP comparison

Test Pre decode Fixed decode Decode delta Pre TTFT ms Fixed TTFT ms TTFT delta
Qwen3.5-4B_Q4_0_GGUF_baseline_code 66.28 66.58 +0.46% 68.76 68.01 -1.09%
Qwen3.5-4B_Q4_0_GGUF_baseline_list 66.80 67.16 +0.54% 66.56 66.87 +0.46%
Qwen3.5-4B_Q4_0_GGUF_baseline_prose 66.34 66.68 +0.51% 66.18 67.43 +1.89%
Qwen3.5-4B_Q4_0_GGUF_mtp_code 116.73 117.15 +0.37% 78.20 78.66 +0.59%
Qwen3.5-4B_Q4_0_GGUF_mtp_list 133.18 133.59 +0.31% 77.31 77.44 +0.16%
Qwen3.5-4B_Q4_0_GGUF_mtp_prose 85.10 85.32 +0.26% 76.03 76.52 +0.64%
Qwen3.5-4B_Q4_K_M_GGUF_baseline_code 62.79 63.36 +0.91% 61.97 62.29 +0.52%
Qwen3.5-4B_Q4_K_M_GGUF_baseline_list 63.48 63.66 +0.28% 61.88 61.41 -0.75%
Qwen3.5-4B_Q4_K_M_GGUF_baseline_prose 62.89 63.43 +0.86% 60.87 61.40 +0.87%
Qwen3.5-4B_Q4_K_M_GGUF_mtp_code 102.40 102.79 +0.38% 72.54 73.16 +0.85%
Qwen3.5-4B_Q4_K_M_GGUF_mtp_list 110.20 109.94 -0.23% 71.95 72.43 +0.66%
Qwen3.5-4B_Q4_K_M_GGUF_mtp_prose 68.54 67.84 -1.02% 71.24 71.61 +0.53%
Qwen3.5-9B_Q4_0_GGUF_baseline_code 39.96 40.05 +0.22% 103.47 102.85 -0.59%
Qwen3.5-9B_Q4_0_GGUF_baseline_list 40.38 40.50 +0.29% 102.29 100.47 -1.78%
Qwen3.5-9B_Q4_0_GGUF_baseline_prose 40.04 40.13 +0.20% 100.94 99.99 -0.94%
Qwen3.5-9B_Q4_0_GGUF_mtp_code 80.01 79.18 -1.03% 115.14 114.82 -0.28%
Qwen3.5-9B_Q4_0_GGUF_mtp_list 86.19 86.25 +0.07% 113.97 114.09 +0.11%
Qwen3.5-9B_Q4_0_GGUF_mtp_prose 60.74 60.83 +0.16% 113.49 113.35 -0.12%
Qwen3.5-9B_Q4_K_M_GGUF_baseline_code 37.18 37.28 +0.27% 94.71 94.98 +0.28%
Qwen3.5-9B_Q4_K_M_GGUF_baseline_list 37.57 37.69 +0.34% 93.56 93.39 -0.19%
Qwen3.5-9B_Q4_K_M_GGUF_baseline_prose 37.28 37.40 +0.30% 93.27 93.97 +0.76%
Qwen3.5-9B_Q4_K_M_GGUF_mtp_code 61.94 62.22 +0.45% 105.84 104.93 -0.86%
Qwen3.5-9B_Q4_K_M_GGUF_mtp_list 66.90 67.73 +1.24% 105.63 104.46 -1.11%
Qwen3.5-9B_Q4_K_M_GGUF_mtp_prose 45.25 44.45 -1.76% 104.99 103.74 -1.19%
Qwen3.6-27B_Q4_0_GGUF_baseline_code 13.39 13.46 +0.49% 313.00 304.29 -2.78%
Qwen3.6-27B_Q4_0_GGUF_baseline_list 13.50 13.56 +0.43% 310.61 302.53 -2.60%
Qwen3.6-27B_Q4_0_GGUF_baseline_prose 13.34 13.46 +0.85% 304.16 297.76 -2.10%
Qwen3.6-27B_Q4_0_GGUF_mtp_code 36.03 36.14 +0.30% 342.49 334.06 -2.46%
Qwen3.6-27B_Q4_0_GGUF_mtp_list 36.59 36.73 +0.37% 338.30 332.68 -1.66%
Qwen3.6-27B_Q4_0_GGUF_mtp_prose 26.14 27.14 +3.81% 333.30 329.98 -1.00%
Qwen3.6-27B_Q4_K_M_GGUF_baseline_code 12.50 12.47 -0.23% 280.45 281.36 +0.33%
Qwen3.6-27B_Q4_K_M_GGUF_baseline_list 12.61 12.66 +0.32% 280.05 282.28 +0.80%
Qwen3.6-27B_Q4_K_M_GGUF_baseline_prose 12.51 12.55 +0.34% 274.44 275.53 +0.40%
Qwen3.6-27B_Q4_K_M_GGUF_mtp_code 25.41 25.51 +0.39% 311.29 311.11 -0.06%
Qwen3.6-27B_Q4_K_M_GGUF_mtp_list 26.16 26.22 +0.25% 309.81 307.86 -0.63%
Qwen3.6-27B_Q4_K_M_GGUF_mtp_prose 19.02 19.08 +0.30% 306.73 306.07 -0.21%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_baseline_code 58.16 58.17 +0.02% 105.75 105.45 -0.29%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_baseline_list 58.20 58.51 +0.54% 102.75 99.74 -2.94%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_baseline_prose 58.19 58.49 +0.51% 100.27 100.41 +0.13%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_mtp_code 98.08 98.13 +0.05% 121.65 122.88 +1.01%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_mtp_list 98.19 98.26 +0.07% 116.51 118.11 +1.37%
Qwen3.6-35B-A3B_Q4_K_M_GGUF_mtp_prose 72.12 73.56 +1.99% 115.94 115.92 -0.02%

Jim Wu and others added 30 commits July 9, 2026 13:33
* llama-batch: add unit test

* fix win32 builds

* add not implemented assertion in unused methods

* remove unreachable code
* ggml-et: Add performance logging

* ggml-et: Quants helpers

* ggml-et: Add MUL_MAT kernel

* ggml-et: Add ROPE kernel

* ggml-et: Add RMS_NORM kernel

* ggml-et: Add GLU kernel

* ggml-et: Add SOFT_MAX kernel

* ggml-et: Add GET_ROWS kernel

* ggml-et: Add CONT kernel

* ggml-et: Add SET_ROWS kernel

* ggml-et: Add MUL_MAT_ID kernel

* ggml-et: Build et kernels as part of ggml

* ggml-et: Embed kernels with fs fallback

* ggml-et: Build fixes

* ggml-et: Add MUL_MAT F32xF32 op

* ggml_et: Add MUL_MAT_ID op

* ggml-et: Disable offloading for debug

* ggml-et: Refactor out block ops

* ggml-et: ggml backend API changes

* ggml-et: Add RESHAPE/TRANSPOSE to supported

* ggml-et: Add CONT_F16

* ggml-et: Add supported ops doc

* gglm-et: Initial doc

* ggml-et: Remove  runtime import hacks

We can now import the runtime by a simple find_package(), so we
can cleanup the CMakeLists.txt.

* ggml-et: Fix GET_ROWS kernel

Fix lost batch dimension.

Also clean vibe-comments.

* ggml-et: Fix SET_ROWS kernel

Remove incorrect broadcasting guard.

* ggml-et: Use custom instruction for fp32->fp16

* ggml-et: Vectorize set_rows fp32->fp16

* ggml-et: Fix ROPE kernel (yarn)

ggml-et: fix et_logf

WIP: Fix ramp

WIP: fix ROPE!

* ggml-et: Better sinf

* ggml-et: Fix SOFT_MAX

Add `max_bias` and `sink` support.

* ggml-et: Fix CONT

Reorder from contiguous write to read with atomic stores.

* ggml-et: Fix elmap kernel

Remainder handlin

* ggml-et: Fix MUL_MAT MUL_MAT_ID remainders

* ggml-et: Fix ET-SOC reference

* ggml-et: Fix embed kernels scripts for old python

This allows GGML-ET to build on pre-3.8 python.

* Add sysemu support with compile time flag `-DGGML_ET_SYSEMU=ON` (#6)

* Example using ET-Soc-1 emulator configuration

Example usage:
```bash
cmake -B build -DGGML_CUDA=OFF -DGGML_ET=ON -DLLAMA_CURL=OFF -DGGML_CCACHE=ON
cmake --build build --config Release -j $(nproc)

time ./build/bin/test-backend-ops

./build/bin/llama-server \
    --model Qwen3-0.6B-Q8_0.gguf \
    --alias Qwen3-0.6B-Q8_0 \
    -fa 0 \
    --ctx-size 1024 \
    --no-warmup \
    --host 127.0.0.1 \
    --port 8080
```

* build: proper dep tracking for kernels

* support host using MOLD linker

* initial multi core GET_ROW F32 implementation

* vectorized q8 dequant

* wip: cland warning clenaups and initial logging refactor

* wip: message default message cleanup

* chore: message cleanups

* cmake cleanup

* migrate to use platform provided functions

* cmake back into subdir

* support et_print() in kernels

* fix: repair kernel building

* perf: operations run async by default

* debug: proper kernel dep tracking and error detection on kenrel launch

* fix: kernel binary dep tracking and fixing get_rows_f32 erroring

* perf: back to doing async kernel runs by default

* perf: vectorize and parallel device memset

* merge matmul work

* misc: align allocation and enable all offload

* misc: delete deadcode and respect memory limits

* fix: repair tensor debug print

* fix: loosen RMS_NORM op percision

* feat: Q4_0 GET_ROWS

* perf: FP32 MUL_MAT using TensorFMA

* update limitations

* perf: redue L1 load in compute_block_dot_product_q8_0

* feat: save kernel mapping (name to id) when profiling is enabled

* chore: memops cleanup

* perf: parallelize softmax by rows

* perf: vectorize 2nd phase of softmax

* perf: ban GET_ROWS from offloaded

* perf: vectorize and non-atomic for eltwise ops and sub support

* perf: vectorize normal rope

* perf: glu runs in parallel

* merge: manually merge saqib's work on kernel fixes

* perf: more vectorized RoPE

* perf: parallelize mul_mat_id

* perf: parallelize set_rows_f32

* perf: vectorize softmax

* feat: support kernel fusion and fuse RMS_NORM + MUL

* fix: mostly resolve test-backend-ops failure in SOFT_MAX and ROPE

* fix: bump max rope dims for gemma

* feat: GeGLU and SCALE support to fully offload Gemma

* perf: faster device memset

* feat: get_rows supporting Q4_K and avoid cont cache coherent issues

* better F32 MM

* feat: NORM for ET backend

* feat: SQR for ET backend

* feat: UNARY on ET

* feat: el_map support broadcasting for ET

* feat: SUM_ROWS in ET backend

* feat: more ops in ET backend

* feat: WKV* operators in ET backend

* perf: parallelize operators across cacheline instead of row

* perf: parallelize get_rows on cacheline

* wip: baseline FlashAttention for ET backend

* wip: enough FA and CPY f32->f16 to run llama 3.1 fully offloaded with FA on

* feat: f16 x f16 -> f32 MM using matrix engine

* wip: f16 FlashAttention using matrix engine

* wip: clean up

* feat: barriers

* perf: optimize FA_F16 in ET

* perf: vectorize pack_k_for_transpose16

* perf: prefetch next loop matrix tile

* perf: FlashAttention 2nd MM uses TensorFMA and optimizations

* cleanup: flashattention reorg

* perf: optimizations and fixes

* feat: L2SCP API and make FlashAttention support DV = 256 for gemma

* perf: parallelize norms beyond single row

* feat: GATED_DELTA_NET support and relaxed L2_NORM requirment

* feat: loosen RMS_NORM, NORM, ROPE contingous req too

* feat: repeat supports brocasting on dim 0 and loosen cont check

* feat: FILL and DIAG operator

* feat: loosen UNARY support chcek

* feat: TRI support

* feat: SOLVE_TRI support

* feat: basic SET support

* feat: loosen CONT req

* perf: fp16_to_fp32 use ASM

* feat: IMROPE support

* feat: PAD support

* feat: global barrier

* fix: view must live on the same backend as backing tensor

* feat: relax CONCAT in ET backend

* feat: dead simple CUMSUM implementation

* feat: basic SSM_CONV support

* feat: loosen CONCAT req

* feat: relax GATED_DELTA_NET and add SET support proper

* cleanup: cleanup LCM math

* feat: SWIGLU single input

* feat: SSM_SCAN support

* feat: el_map supports non aligned tensors in best effort

* feat: basic GROUP_NORM support

* feat: loosen MUL_MAT capablities slightly

* feat: loosen MUL_MAT and GET_ROWS and add IM2COL

* feat: special case for softmax 1x1x1x1

* feat: loosen SOFT_MAX req in ET backend

* fix: el_map unaligned acse fixes

* perf: optimize zero_acc_vec in flash_attn_ext_f16_me

* perf: use hart 1 for packing in MM and FA for FP16

* feat: kernel semaphore

* perf: better instruction sequence in FlashAttention

* fix: gated_delta_net with proper masking

* perf: better parallelization for GATED_DELTA_NET

* perf: parallelize SSM_CONV over nr

* perf: vectorize SSM_CONV

* perf: optimize MUL_MAT for q8

* feat: support Gemma 4

* fix: support multi-device

* feat: broader GLU support

* feat: unary ops supports view

* fix: repair fp16 MM using matrix engine

* perf: handle large N GEMV better

* perf: better q8_0 MM

* perf: better set_rows

* add back deleted files

* fix: repair after merge

* feat: POC version of uberkernel

* feat: RMS_NORM in uberkernel

* feat: add more kernels into usage

* chore: clean up uberkernel compilation

* perf: faster flash attention

* perf: opt flash attention for large seq length

* feat: loosen op bounds. clamp and mean support

* perf: vectorize ssm_scan

* perf: slightly faster FA

* perf: FlashAttention parallel MM and load

* perf: fuse Q8 MM and ADD

* feat: basic conv kernel for ET

* softMAx_test

* set_rows_f32

* get_rows and cont

* testing

* set_rows_exp

* Junk addition

* Narrowing the issue

* Update flash_attn_ext_f16_me.c

Focusing FA_ext_f16_me

* test

* Eviction updated

* Detailed cache eviction debug

* mulmat

* removeal of `BUILD_FOR_UBERKERNEL` flag

* cleaning...

* fix: balance FCC0 count

* feat: implement mul_mat and mul_mat_id for Q4_0 type

* optimize uberkernel plan upload

* add mul_mat q4 into uberkernel

* enable gating flush to just uberkernel

* update docs for ET

* update op support for ET

* et-backend: optimize Q4_0 and Q8_0 mul_mat_id row accumulations

* et-backend: specialize mul_mat_id kernels for Q4_0 and Q8_0

* et-backend: fix RoPE YaRN corr_dim formula and handle degenerate inputs

* test-backend-ops: add DeepSeek-V2-Lite RoPE test coverage

* et-backend: add Q4_0 mul_mat matrix-engine kernel using TensorFMA32

* et-backend: vectorize Q4_0 matrix-engine dequantization

* et-backend: support hybrid matrix/vector engine execution for Q4_0 mul_mat tail

* et-backend: run partial-N tiles on matrix engine for Q4_0 mul_mat

* et-backend: route Q4_0 mul_mat N < 53 to vecdot for better prefill latency

* Update uberkernel.c

* Update unary_f32.c

* gemma 4

* bisect gemma4: enable scale_f32 only

* bisect gemma4: +rms_norm_f32

* bisect gemma4: +rms_norm_mul_f32

* bisect gemma4: disable rms_norm_mul_f32 -- BREAKS OUTPUT

* bisect gemma4: +rope_f32 (skip rms_norm_mul)

* bisect gemma4: +el_map_f32

* bisect gemma4: +softmax_f32

* bisect gemma4: +get_rows_f32

* bisect gemma4: +glu_f32

* bisect gemma4: +mul_mat_f32 +mul_mat_f32_matrix_engine

* bisect gemma4: +mul_mat_f16 +mul_mat_f16_matrix_engine

* bisect gemma4: +mul_mat_Q8_0 +mul_mat_Q4_0

* bisect gemma4: +flash_attn_ext_f32 +flash_attn_ext_f16_me

* bisect gemma4: +mul_mat_id_f32

* bisect gemma4: +sum_rows_f32

* bisect gemma4: +cont_f16

* bisect gemma4: +fill_f32

* bisect gemma4: +unary_f32 (all ops re-enabled except rms_norm_mul)

* Update rms_norm_mul_f32.c

* bisect2 gemma4 n64: +scale_f32 only

* bisect2 gemma4 n64: +rms_norm_f32 +rope_f32

* bisect2 gemma4 n64: +rms_norm_mul_f32 (with ET_UBERKERNEL eviction fix)

* bisect2 gemma4 n64: +el_map +get_rows +glu +softmax (skip rms_norm_mul)

* bisect2 gemma4 n64: all ops enabled except rms_norm_mul

* bisect2 n64: test unary+cont+fill+sum_rows (no mul_mat/flash_attn)

* bisect2 n64: +mul_mat_f32 +mul_mat_f32_matrix_engine

* bisect2 n64: +mul_mat_f16 +mul_mat_f16_matrix_engine

* bisect2 n64: +mul_mat_Q8_0 +mul_mat_Q4_0

* bisect2 n64: +mul_mat_Q8_0 only (disable Q4_0)

* bisect2 n64: +mul_mat_Q4_0 only (Q8_0 breaks)

* bisect2 n64: +mul_mat_id +flash_attn_ext (skip Q8_0)

* run-3: matmul + rms_norm_mul

* run-4

* Revert "run-4"

* run5

* changes after cleanup

* cleanup before upstream

* restrict changes into ET backend

* move kernel embedding from Python to CMake

* move uberkernel gen into CMake

* apply clang format

* update CMake style

* update to match C and C++ style

* use source ggml and quant headers instead of ET's

* MROPE support

* absorb view ops into same branch as none

* fix bad rebase

* add marty1885 to codeowners

* oops

* remove redundant newline

* fix CI editor warnings

---------

Co-authored-by: Vidas <vidas@nuolat.lt>
Co-authored-by: Gianluca Guida <glguida@tlbflush.org>
Co-authored-by: Gianluca Guida <gianluca@nekko.ai>
Co-authored-by: ubergarm <leimgrub@gmail.com>
Co-authored-by: SaqibAkram-10xE <saqib.akram@10xengineers.ai>
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
…as, remove raw_k repeats in DeepSeek V4 (ggml-org#25370)

* llama : make all KQ masks (except the lightning indexer one) f16 if FA is used and remove zero attention bias in DeepSeek V4

* llama : remove dead code that repeats unified raw_k cache for each stream in DeepSeek V4 - no longer needed as raw_k is always non-unified.

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
)

downloadConversation serialized activeMessages, the root -> currNode
path, so exporting a conversation with edited or regenerated messages
dropped every alternate version and kept only the selected one.

Fetch the whole message tree via getConversationMessages so the export
carries all message versions, matching the multi-conversation export
path which already did this. Keep the active conversation as the header
source to preserve an up-to-date currNode.

Forks are separate conversations, each with its own convId, and are
exported on their own.
* ggml : bump version to 0.16.0 (ggml/1559)

* sync : ggml
* llama-cli: fix crash on wrong server base url by catching exceptions and graceful exit

* review: leaner catch group: json error and standard exception
* server: improve tools, remove apply_diff

* improve edit tool

* add tools_io abstraction

* add tools_io_basic

* fix build

* move utils to class member

* add const
* server: remove loading.html

* apply ui changes
Co-authored-by: example name <example@example.org>
* mtmd: deepseek-ocr v1 multi-tile dynamic resolution + unified image-preprocessors for both versions (ds-ocr v1 and v2)

* remove hacky API

* fuse row into a long image

* almost working

* adapt to new preprocessor api

* rm debugging printf

* improve

* mtmd: dsocr-tiles fixes (ggml-org#25481)

* ds-ocr img-preproc fuse_row tile-drop fix for multi rows and columns images

* mtmd drop the duplicate redundant img_end

* deepseekocr graph simplify CLS broadcast cleanup

* test-deepseek-ocr: relax v1 single-view tolerance; drop trailing prompt space; make DRY opt-in and n_predict model-specific (ggml-org#25486)

---------

Co-authored-by: Saba Fallah <10401143+sfallah@users.noreply.github.com>
Co-authored-by: Saba Fallah <sabafallah@gmail.com>
* hex-sort: add efficient bitomic sort in hvx regs up to 1024 elements

* hex-sort: fix inverted vrors

* hex-sort: specialize sort functions for the common cases

* hex-sort: add tracing and local context
llama_meta_device_get_split_state() recompiled 29 std::regex on every call.
In -sm tensor mode the callback runs once per tensor per token, so this
dominated the decode thread in profiling. Mark them static const so they are
compiled once. Kept inside the function (local statics are thread-safe since
C++11). Patterns are literal and stateless, so behavior is unchanged.
* server: accept null sampling params

Extend the schema validation to treat a null value as absent, so
clients can send null on nullable params (temperature, top_p, ...)
to request the server default. This matches the OpenAI spec and the
json_value convention used elsewhere.

Add has_field() to skip null in the field eval guards.

* has_field -> has_value​
…Us (ggml-org#25537)

* opencl: add int8 dp4 dense and moe GEMM

* opencl: refactor

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
* [Vulkan] Fixes llama-cli breaking over longer promts sizes

The llama-cli was breaking for longer promts sizes for q4_0 quantized networks. Causing due to insufficient shared memory.

* Removed the un-used Adreno device

* Updated matmul for small pipeline.
… lightning indexer (ggml-org#24231)

* ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer

* ggml : remove scale parameters from lightning indexer OP, add f16 mask parameter

* tests : add GGML_OP_LIGHTNING_INDEXER tests

* ggml : bump RPC version

* chore : check if lightning indexer input tensors are not transposed

* tests : count flops instead of bandwidth in lightning indexer test

* chore : add missing const

* chore : whitespace

* ggml : renamed variables in CPU lightning indexer implementation

* ggml : fix lightning indexer mask broadcasting

* tests : tests for lightning indexer mask broadcasting

* chore : whitespace

* llama : use GGML_OP_LIGHTNING_INDEXER in DeepSeek V3.2 and DeepSeek V4 models

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
* server: refactoring, remove spipe from server_http_res

* wip

* remove non-thread-safe rd.stop() call

* move server_res_spipe

* nits

* improve server_stream_create_spipe

* server-stream: update dev docs for the improved API

---------

Co-authored-by: Pascal <admin@serveurperso.com>
* init stream

* add stream for shell tool

* add test

* nits

* update docs
…ggml-org#25157)

If a Cuda device has no or limited available memory, the actual call
to cudaMemGetInfo() itself can cause a fatal crash due to a cuda out
of memory error (there is not enough memory to actually query memory)

This causes an issue because we query memory for all devices at
startup even if the user isn't trying to use the device for inference.

Fix this by making the error non-fatal and assigning zero total/free
memory to the device. This will have the downstream effect of the fit
algorithm not trying to put any layers on it, which is desired outcome
vs hard crashing.

this also prevents crashes in cuda enabled builds when user explicitly
passes '-dev none'
…ic OpenAI conversion (ggml-org#22536)

* server : fix image blocks in tool_result being dropped during Anthropic→OpenAI conversion

server_chat_convert_anthropic_to_oai() silently discarded image blocks

inside Anthropic tool_result content. This broke multimodal tool outputs

(e.g. a tool that returns an image) because the model never received the

image.

When tool_result contains image blocks, convert them to OpenAI

multimodal content parts (text + image_url array). Plain-text results

remain simple strings for backwards compatibility.

* server : add test for image blocks in Anthropic tool_result conversion
helanfxz and others added 13 commits July 25, 2026 10:23
…et sampler (ggml-org#25544)

* common : extract trie/ac to a separate file

* common : support multiple token sequences in the reasoning budget sampler

* common/trie : return matched word index

* common/trie : rename "word" to "pattern"

* common/reasoning-budget : expose matched end sequence

* common/sampling : replay end sequence when reasoning budget is done

* cont : update to use multiple end sequences

* cont : clean up
…gguf_writer_base (ggml-org#25867)

Without a virtual destructor, deleting a derived object through a
base-class pointer only invokes the base destructor, skipping the
derived one.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
… in generation_settings (ggml-org#25830)

These two parameters were overlooked when task parameters were being JSONized
within `generation_settings` and have been added. A regression test has been
added to prevent the problem from recurring and it passes.

Fixes ggml-org#25803
The INI parser creates an implicit default section for top-level metadata.
After reserved keys such as `version` are skipped, that section can have no
model options but was still added and exposed in router mode.

Skip only the empty implicit default while preserving real default presets,
named presets, and the global `[*]` settings.

Signed-off-by: JS van Dijk <267467744+hogeheer499-commits@users.noreply.github.com>
Co-authored-by: JS van Dijk <267467744+hogeheer499-commits@users.noreply.github.com>
Sync with upstream ggml-org/master (2026-07-25)
Preserve the gfx11 RDNA3.5 MMQ paths while adopting the latest master refactors and quantization support.

Co-authored-by: Cursor <cursoragent@cursor.com>
Run the RDNA3.5 shape matrix through every quantized mul_mat_q specialization so master merges expose type-specific correctness regressions.

Co-authored-by: Cursor <cursoragent@cursor.com>
Cover the remaining mul_mat_q specialization with the production-shape correctness sweep.

Co-authored-by: Cursor <cursoragent@cursor.com>
Restore bitwise Q1_0 unpacking on HIP to recover AMD MMQ and MMVQ performance while retaining the optimized byte-permute path elsewhere.

Co-authored-by: Cursor <cursoragent@cursor.com>
@liangliangchang liangliangchang changed the title Lichang.gfx11 master sync Merge latest upstream changes in master to gfx11. Jul 27, 2026
liangliangchang and others added 2 commits July 27, 2026 14:28
Match the post-merge mul_mat_q signature so the quality check continues to ignore the documented pre-existing gfx908 spills.

Co-authored-by: Cursor <cursoragent@cursor.com>
Use two WMMA row tiles only when the column tile is aligned, preventing partial tiles from overflowing accumulators.

Co-authored-by: Cursor <cursoragent@cursor.com>
@liangliangchang
liangliangchang marked this pull request as ready for review July 28, 2026 17:00
Recover J=16 prefill performance lost in the MMQ refactor and cover the affected Q4_K shapes in backend operation tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
@liangliangchang
liangliangchang marked this pull request as draft July 28, 2026 23:15
@liangliangchang

Copy link
Copy Markdown
Author

Moving it to draft because I found there are still small changes causing 2-3% regressions on different models. I will fix them all before reopening this PR.

liangliangchang and others added 2 commits July 28, 2026 17:21
Keep integrated GPU detection while staging gfx1151 buffers in device memory to avoid UMA prefill latency regressions.

Co-authored-by: Cursor <cursoragent@cursor.com>
@liangliangchang
liangliangchang marked this pull request as ready for review July 29, 2026 17:05
Adapt the tile-width tuning to the refactored J-based MMQ configuration so the sync branch builds again.

Co-authored-by: Cursor <cursoragent@cursor.com>

@jimw567 jimw567 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good

@jimw567
jimw567 merged commit 3c53e42 into gfx11 Jul 30, 2026
8 checks passed
@jimw567
jimw567 deleted the lichang.gfx11-master-sync branch July 30, 2026 15:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.