Checkpoint or conversion artifacts for ds4.c GGUF conversion?

#4
by cutiedeng - opened

Hi CyberNeurova team,

Thank you for releasing the abliterated DeepSeek V4 Flash GGUFs.

I am working with the ds4.c DeepSeek V4 Flash inference project, which uses a very specific GGUF tensor layout and quantization mix. The currently released Q8_0 GGUF appears to have routed expert weights quantized as Q8_0, while ds4.c expects a mixed layout where dense paths are mostly Q8_0 but routed experts use formats such as IQ2_XXS / Q2_K / Q4_K.
Would it be possible to share one of the following?

  1. The abliterated checkpoint / safetensors before GGUF quantization
  2. The ablation delta or direction vectors needed to reproduce the abliterated weights
  3. Any intermediate converter output that could be re-quantized into a custom GGUF layout
  4. Guidance on whether the existing Q8_0 GGUF can be safely used as a source for re-quantizing routed experts

My goal is to convert the abliterated model into the GGUF layout currently supported by ds4.c, so I can try running it with that inference engine. I am not asking for support for ds4.c itself, just whether the source/intermediate weights or reproducibility artifacts are available.

Thanks again for the release.

I have made a working ds4 backend for cyberneurova abliterated model with multi-gpu support. Prefill is slow and in general it is not perfect but a good starting point. If anyone is interested I can share the code.

$ ./ds4 -m /workspace/cyberneurova-DeepSeek-V4-Flash-abliterated-Q2_K.gguf
-p "Explain Bitcoin in two sentences."
ds4: context buffers 751.71 MiB (ctx=32768, backend=cuda, prefill_chunk=2048, raw_kv_rows=2304, compressed_kv_rows=8194)
ds4: CUDA backend initialized on NVIDIA RTX A6000 (sm_86)
ds4: CUDA visible devices: 4, primary: 0
ds4: CUDA model cache devices: 0
ds4: CUDA host registration skipped: layer cache requests device-resident weights
ds4: CUDA layer split by bytes: dev0=0-9(25.67GiB) dev1=10-20(23.28GiB) dev2=21-31(23.27GiB) dev3=32-42(23.81GiB)
ds4: CUDA loading model tensors into device cache: 92.02 GiB

ds4: CUDA layer model cache prepared 92.02 GiB of tensors in 38.306s
ds4: cuda backend initialized for graph diagnostics

We need to explain Bitcoin in two sentences. Keep it concise, clear, and accurate. Bitcoin is a decentralized digital currency that allows peer-to-peer transactions without intermediaries, using blockchain technology to secure and verify transactions. It was created in 2009 by an anonymous person or group. That's the core. But need to make it two sentences. Possibly: Bitcoin is a decentralized digital currency that enables peer-to-peer transactions without a central authority, using cryptographic proof and a public ledger called the blockchain. Its limited supply and growing adoption have made it both a medium of exchange and a store of value. That's two sentences.

Bitcoin is a decentralized digital currency that allows peer-to-peer transactions without the need for a central authority, using a public ledger called the blockchain to ensure security and transparency. It is created through a process called mining, with a finite supply capped at 21 million coins, making it both a medium of exchange and a store of value.

ds4: prefill: 12.35 t/s, generation: 27.12 t/s

I have made a working ds4 backend for cyberneurova abliterated model with multi-gpu support. Prefill is slow and in general it is not perfect but a good starting point. If anyone is interested I can share the code.

$ ./ds4 -m /workspace/cyberneurova-DeepSeek-V4-Flash-abliterated-Q2_K.gguf
-p "Explain Bitcoin in two sentences."
ds4: context buffers 751.71 MiB (ctx=32768, backend=cuda, prefill_chunk=2048, raw_kv_rows=2304, compressed_kv_rows=8194)
ds4: CUDA backend initialized on NVIDIA RTX A6000 (sm_86)
ds4: CUDA visible devices: 4, primary: 0
ds4: CUDA model cache devices: 0
ds4: CUDA host registration skipped: layer cache requests device-resident weights
ds4: CUDA layer split by bytes: dev0=0-9(25.67GiB) dev1=10-20(23.28GiB) dev2=21-31(23.27GiB) dev3=32-42(23.81GiB)
ds4: CUDA loading model tensors into device cache: 92.02 GiB

ds4: CUDA layer model cache prepared 92.02 GiB of tensors in 38.306s
ds4: cuda backend initialized for graph diagnostics

We need to explain Bitcoin in two sentences. Keep it concise, clear, and accurate. Bitcoin is a decentralized digital currency that allows peer-to-peer transactions without intermediaries, using blockchain technology to secure and verify transactions. It was created in 2009 by an anonymous person or group. That's the core. But need to make it two sentences. Possibly: Bitcoin is a decentralized digital currency that enables peer-to-peer transactions without a central authority, using cryptographic proof and a public ledger called the blockchain. Its limited supply and growing adoption have made it both a medium of exchange and a store of value. That's two sentences.

Bitcoin is a decentralized digital currency that allows peer-to-peer transactions without the need for a central authority, using a public ledger called the blockchain to ensure security and transparency. It is created through a process called mining, with a finite supply capped at 21 million coins, making it both a medium of exchange and a store of value.

ds4: prefill: 12.35 t/s, generation: 27.12 t/s

Thank you for your implementation. If you're willing to share the code, I'd like to ask: would this be suitable for a MacBook M5 Max? I don't have a CUDA backend. Also, I noticed the generation is only 27 tps on an A6000, and as far as I know that's a pretty expensive high-performance GPU (which I don't have), so I'm a bit concerned about the performance.

Sign up or log in to comment