Hi! Thank you for this package.
I briefly measured performance of libcrispasr from the windows build with a python ctypes adapter against ONNX. I used a Parakeet model (parakeet-tdt-0.6b-v3-q8_0.gguf), and a 16-bit quant for onnx.
4070 GPU, I saw it was loaded.
Both CUDA and Vulkan builds show 5 times slower inference comparably to ONNX.
Are windows builds not optimized?
Hi! Thank you for this package.
I briefly measured performance of libcrispasr from the windows build with a python ctypes adapter against ONNX. I used a Parakeet model (parakeet-tdt-0.6b-v3-q8_0.gguf), and a 16-bit quant for onnx.
4070 GPU, I saw it was loaded.
Both CUDA and Vulkan builds show 5 times slower inference comparably to ONNX.
Are windows builds not optimized?