It's not entirely clear:
A novel MoE-aware mixed-precision quantization technique for llama.cpp
Can llama.cpp run these already? It this just a quantization tool that applies existing GGUF capabilities more optimally, or is it a new/enhanced format on disk or in memory?
It's not entirely clear:
Can llama.cpp run these already? It this just a quantization tool that applies existing GGUF capabilities more optimally, or is it a new/enhanced format on disk or in memory?