Name and Version
version: 10216 (876a432)
built with Clang 22.1.8 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
Models
GLM-5.2
Problem description & steps to reproduce
When loading GLM-5.2 with -ctk q5_1 and no -ctv:
llama_init_from_model: model does not support different K (q5_1) and V (f16) cache types
First Bad Commit
No response
Relevant log output
Name and Version
version: 10216 (876a432)
built with Clang 22.1.8 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
Models
GLM-5.2
Problem description & steps to reproduce
When loading GLM-5.2 with
-ctk q5_1and no-ctv:First Bad Commit
No response
Relevant log output