Skip to content

Docs: Clarification request: Which engines can run these? #12

Description

@lee-b

It's not entirely clear:

A novel MoE-aware mixed-precision quantization technique for llama.cpp

Can llama.cpp run these already? It this just a quantization tool that applies existing GGUF capabilities more optimally, or is it a new/enhanced format on disk or in memory?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions