Skip to content

llama: re-create the KV cache when flash attention resolves to disabled (performance and layout bug) - #26460

Open
wanghqc wants to merge 1 commit into
ggml-org:masterfrom
qualcomm:hq/llama-fa-auto-recreate-kv-cache
Open

llama: re-create the KV cache when flash attention resolves to disabled (performance and layout bug)#26460
wanghqc wants to merge 1 commit into
ggml-org:masterfrom
qualcomm:hq/llama-fa-auto-recreate-kv-cache

llama: re-create the KV cache when flash attention resolves to disabled

2ffd0d7
Select commit
Loading
Failed to load commit list.
Sign in for the full log view