Hi! I'm reaching out from TurboLLM (https://github.com/mohitsoni48/TurboLLM) — a local LLM inference platform built on top of llama.cpp that lets users easily install and run various llama.cpp forks with a single click.
We noticed your llama.cpp-tq3 fork with the TQ3_1S/4S CUDA kernels. It looks like a solid contribution to the low-bit KV-cache quantization space, and we think it could be useful for our users — especially those working with Blackwell GPUs.
We'd love to list your fork in our engine catalog so users can discover and install it directly from the TurboLLM app. If you're open to this, please open an issue here in this repo and we'll take it from there. Or you can reach out to us directly.
TurboLLM currently supports a growing list of llama.cpp forks (official, prism, beellama, etc.) and gives each engine its own card with one-click install + launch. Listing your fork would expose it to a wider audience of local LLM users.
Looking forward to hearing from you!
Hi! I'm reaching out from TurboLLM (https://github.com/mohitsoni48/TurboLLM) — a local LLM inference platform built on top of llama.cpp that lets users easily install and run various llama.cpp forks with a single click.
We noticed your llama.cpp-tq3 fork with the TQ3_1S/4S CUDA kernels. It looks like a solid contribution to the low-bit KV-cache quantization space, and we think it could be useful for our users — especially those working with Blackwell GPUs.
We'd love to list your fork in our engine catalog so users can discover and install it directly from the TurboLLM app. If you're open to this, please open an issue here in this repo and we'll take it from there. Or you can reach out to us directly.
TurboLLM currently supports a growing list of llama.cpp forks (official, prism, beellama, etc.) and gives each engine its own card with one-click install + launch. Listing your fork would expose it to a wider audience of local LLM users.
Looking forward to hearing from you!