Skip to content

Support request: Bielik-11B-v3.0-Instruct-heretic-MPOA on Ryzen AI NPU #704

Description

@Mikikudi

Hello FastFlowLM team,

I would like to request NPU support for the Polish Llama-based model Bielik-11B-v3.0-Instruct-heretic-MPOA.

Model sources:

I converted the Q6_K GGUF to Q4NX_FLL successfully. The resulting model has 50 layers, hidden size 4096, intermediate size 14336, 32 attention heads, and 8 key/value heads.

System:

  • Linux
  • FastFlowLM 1.0.4
  • AMD RyzenAI-npu6 / XDNA2
  • XRT detects the NPU and flm validate --json reports ready: true.

I tested the existing Llama-3.1-8B-NPU2 xclbins as a diagnostic only. The model loads and acquires the NPU, but generation fails with:

runlist failed execution (ERT_CMD_STATE_TIMEOUT)
Kernel Instance: MLIR_AIE

Could you please advise:

  1. Is there an existing compatible xclbin set for this Bielik/Llama architecture?
  2. If not, could support be added, including the required attn.xclbin, dequant.xclbin, layer.xclbin, and mm.xclbin?
  3. Is there a supported public workflow for compiling FLM-compatible xclbins for a custom Q4NX_FLL model?

I can provide the Q4NX tensor metadata, model configuration, and reproducible logs if helpful.

Thank you.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions