-
Notifications
You must be signed in to change notification settings - Fork 21.3k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
tools : split Metal FA-vec tuning into a standalone tuner
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
documentation
Improvements or additions to documentation
examples
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#26498
opened Aug 3, 2026 by
forforever73
Contributor
Loading…
Deepseek4: concat inputs to stay under limit
#26496
opened Aug 3, 2026 by
am17an
Contributor
Loading…
CUDA: support ncols > 1024 in bitonic argsort/top-k The original bito…
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26493
opened Aug 3, 2026 by
Geramy
Contributor
Loading…
Deepseek 4: changes relating to the ggml tensor library for machine learning
testing
Everything test related
-sm tensor
ggml
#26490
opened Aug 3, 2026 by
am17an
Contributor
Loading…
docs: fix inaccurate non-terminal naming rules in GBNF README
documentation
Improvements or additions to documentation
#26489
opened Aug 3, 2026 by
JingliangGao
Loading…
cuda: add GGML_CUDA_BLOCKING_SYNC to eliminate 100% CPU busy-wait on full GPU offload
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26487
opened Aug 3, 2026 by
hksk
Loading…
2 tasks done
vendor : update cpp-httplib to 0.52.0
vendor
#26485
opened Aug 3, 2026 by
cabelo
Contributor
Loading…
docs(server): note OpenAI client base_url for multi-model gateways
documentation
Improvements or additions to documentation
server
#26483
opened Aug 3, 2026 by
seven7763
Loading…
3 tasks
Add Vulkan SDK setup and caching to ubuntu release workflow
devops
improvements to build systems and github actions
#26479
opened Aug 2, 2026 by
oscarbg
Contributor
Loading…
opencl: quant lm_head / decode GEMV and medium-batch GEMM optimizations (speculative decoding/MTP)
ggml
changes relating to the ggml tensor library for machine learning
OpenCL
Issues specific to the OpenCL backend
opencl: fix q6_K flat mul_mat for Adreno A6x/A7x GPUs with older E031 compilers
ggml
changes relating to the ggml tensor library for machine learning
OpenCL
Issues specific to the OpenCL backend
llama : allocate GLM_DSA indexer cache only in "full" indexer layers
#26474
opened Aug 2, 2026 by
fairydreaming
Collaborator
•
Draft
Preserve omitted assistant content during OpenAI-compatible message round-tripping
testing
Everything test related
#26473
opened Aug 2, 2026 by
i386
Loading…
Fix issue where tool arguments can appear in any JSON-valid order
testing
Everything test related
#26472
opened Aug 2, 2026 by
i386
Loading…
ggml : fuse soft_max sweeps into fewer passes
ggml
changes relating to the ggml tensor library for machine learning
#26468
opened Aug 2, 2026 by
cschanaj
Loading…
Instella moe
conversion
model
Model specific
#26467
opened Aug 2, 2026 by
csabakecskemeti
Contributor
Loading…
ggml-cuda: HIP replace __shfl_xor_sync with dpp instructions
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26466
opened Aug 2, 2026 by
thelittlefireman
•
Draft
opencl: make the MoE expert scatter deterministic
ggml
changes relating to the ggml tensor library for machine learning
OpenCL
Issues specific to the OpenCL backend
ci: prepare to onboard AMD ROCm CI
devops
improvements to build systems and github actions
#26461
opened Aug 2, 2026 by
taronaeo
Member
Loading…
llama: re-create the KV cache when flash attention resolves to disabled (performance and layout bug)
#26460
opened Aug 2, 2026 by
wanghqc
Contributor
Loading…
ggml: add gfx90c HIP support
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26454
opened Aug 2, 2026 by
AuroraRAS
Loading…
docs : clarify tensor split context sizing
documentation
Improvements or additions to documentation
#26453
opened Aug 2, 2026 by
zcxGGmu
Loading…
1 task done
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.