I am inquiring about whether Heretic, in its current state, is usable, or even effective, against the DeepSeek V4 Flash model.
The first thing that caught me off guard was that the file size is relatively small compared to the parameter size - unlike other models.
There's not much material online, but discovering this: https://outcomeschool.com/blog/decoding-deepseek-v4
It seems DeepSeek used 'FP4 Quantization-Aware Training', name seems to imply the model was trained with FP4 as the base.
This is quite alarming because I am unsure how Heretic would work. Would it even work? If it does, would it be less precise than usual?
Forgive me since I don't know the intricacies regarding this, especially since I've never seen a model like this.
I am inquiring about whether Heretic, in its current state, is usable, or even effective, against the DeepSeek V4 Flash model.
The first thing that caught me off guard was that the file size is relatively small compared to the parameter size - unlike other models.
There's not much material online, but discovering this: https://outcomeschool.com/blog/decoding-deepseek-v4
It seems DeepSeek used 'FP4 Quantization-Aware Training', name seems to imply the model was trained with FP4 as the base.
This is quite alarming because I am unsure how Heretic would work. Would it even work? If it does, would it be less precise than usual?
Forgive me since I don't know the intricacies regarding this, especially since I've never seen a model like this.