Skip to content

[Feature]: Qwen3.8-Flash-Next #692

Description

@damosvil

Suggestion Description

Devs are starting to run this model in CPU @ 6tps in Llama CPP and 12tps in CUDA in a RTX 3060 12Gb. Please, could you bring this model to XDNA2?

More open models are coming and they can even run at decent tps in CPU. By the paper XDNA2 would be able to handle these new models at a good pace.

Operating System

Ubuntu 26.04

GPU

Ryzen AI 9 HX 370

ROCm Component

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions