Skip to content

DeepSeek V4 Flash FP4 Weights #310

Description

@groknt

I am inquiring about whether Heretic, in its current state, is usable, or even effective, against the DeepSeek V4 Flash model.
The first thing that caught me off guard was that the file size is relatively small compared to the parameter size - unlike other models.

There's not much material online, but discovering this: https://outcomeschool.com/blog/decoding-deepseek-v4
It seems DeepSeek used 'FP4 Quantization-Aware Training', name seems to imply the model was trained with FP4 as the base.

This is quite alarming because I am unsure how Heretic would work. Would it even work? If it does, would it be less precise than usual?
Forgive me since I don't know the intricacies regarding this, especially since I've never seen a model like this.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions