DeepSeek-V4-Flash-Vision-Exp-Abliterated-NVFP4

This experimental research derivative deliberately changes safety-refusal behavior. It can produce content that the source model would refuse. It is not a safety-tuned release and should not be exposed without an independent policy and moderation layer.

This is an abliterated derivative of s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4. It preserves that checkpoint's NVFP4 routed experts and replaces only the 33 FP8 attention output writers in layers 10 through 42, plus their scales, with exact payloads from the pinned native Vision-Exp abliterated reference. MTP/DSpark, the vision encoder, routers, shared experts, embeddings, and all other tensors remain unchanged.

The edit uses one rank-1 refusal direction at strength 3.5, preserves each edited output row's L2 norm, and uses three FP8 fixed-point requantization passes. The direction and donor revision are pinned in ABLITERATION_BUILD_RECEIPT.json. The local native source and donor source have matching config and weight-index SHA-256 identities.

The term "abliterated" describes a targeted projection edit; it does not guarantee removal of every refusal or preservation of every capability.

Validation status

  • Structural overlay: passed for 48/48 shards. All 66 target payloads match the pinned donor; every non-target byte matches the NVFP4 base.
  • Text generation: passed, exact 4, HTTP 200, finish_reason=stop.
  • Vision generation: passed, exact Corn, HTTP 200, finish_reason=stop.
  • Safe non-operational boundary suite: 8/8 answered, zero refusals.
  • GSM8K first 100: 99/100, +3.0 points against the 96/100 NVFP4 base.
  • Fixed serving workload: 23.600 tok/s, -0.093 tok/s (-0.39%) against the 23.692 tok/s NVFP4 base.
  • GSM8K task generation: 17.939 tok/s, -0.312 tok/s against the 18.251 tok/s NVFP4 base.

GSM8K uses 8-shot prompting, max_tokens=256, temperature 0, seed 42, and concurrency 1. The fixed workload uses random input length 256, output length 128, five prompts, concurrency 1, and the median of three warm runs. See ABLITERATION_RUNTIME_VALIDATION.json for the public aggregate and evidence hashes.

The safe boundary suite contains only safe, non-operational prompts. It checks for excessive refusal and is not an unsafe-compliance test. Abliteration does not guarantee removal of every refusal or preservation of every capability.

Provenance

  • NVFP4 base: s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4, revision 7e17ce5c84b088d08878cbe6708ae84d989fa256
  • Native source: deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, revision 86f746b36186f0e567729a5c06a8c918caba82a9
  • Abliteration donor: apetersson/DeepSeek-V4-Flash-Vision-Exp-Abliterated, revision 71e308afc1120c9688da67e54d1b54fe6ddd5c94, variant Reference-Native-FP8
  • Donor manifest SHA-256: 4247b1e5c39c6892c3e0df60d273faf7604c6f46b53580b0ae6ab27c68717835
  • Refusal-direction SHA-256: 6e4d8a8f3aa9e21795faab2c5b14d29b019acdf2ddbfbd8238430458a5837fe0

Runtime

Use the same pinned two-node vLLM runtime and selected SM121 TP=2 recipe as the NVFP4 base. Point the recipe's local model path at this derivative. Do not infer behavioral or safety properties from structural validation alone.

License

The source and donor repositories declare the MIT License. Their license and provenance notices are retained.

Downloads last month
133
Safetensors
Model size
305B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
I64
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for s-zaizen/DeepSeek-V4-Flash-Vision-Exp-Abliterated-NVFP4

Quantized
(24)
this model