Skip to content

fix: allow mixed turbo/q8_0 KV cache types in CUDA flash-attention selection #12

fix: allow mixed turbo/q8_0 KV cache types in CUDA flash-attention selection

fix: allow mixed turbo/q8_0 KV cache types in CUDA flash-attention selection #12

Job Run time
20m 18s
17m 9s
37m 27s