Skip to content

Eval bug: same K and V cache type enforced for models with no V cache #26382

Description

@wolfpld

Name and Version

version: 10216 (876a432)
built with Clang 22.1.8 for Linux x86_64

Operating systems

Linux

GGML backends

CUDA

Hardware

Models

GLM-5.2

Problem description & steps to reproduce

When loading GLM-5.2 with -ctk q5_1 and no -ctv:

llama_init_from_model: model does not support different K (q5_1) and V (f16) cache types

First Bad Commit

No response

Relevant log output

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions