DeepSeek-V3-0.86B-MTP-PR3225-FP8-Dynamic

FP8_DYNAMIC through the model pipeline for LLM Compressor PR #3225. This is a small architecture and checkpoint test fixture, not a production language model. Its backbone was initialized randomly and trained on the repository tiny-model skill's toy text corpus; no pretrained base-model weights were used.

Model and provenance

  • Original architecture/tokenizer: deepseek-ai/DeepSeek-V3-Base, revision afb92e1fa402c2be2a9eb085312bb02e0384d6c7.
  • Transformers class: DeepseekV3ForCausalLM.
  • Total source parameters including synthetic MTP: 860,930,560 (0.861B).
  • Backbone parameters: 849,041,408; MTP parameters: 11,889,152.
  • Estimated active backbone parameters: 643,782,656. This routing estimate includes embeddings/dense components, excludes MTP, and is not a FLOP measurement.
  • Precision: compressed-tensors FP8_DYNAMIC, with exclusions retained in their source dtype.
  • Checkpoint: 3 indexed safetensors shards; 4,197 indexed tensors. Tokenizer assets are included.

Configuration

Field Original Tiny
num_hidden_layers 61 61
hidden_size 7168 768
intermediate_size 18432 3072
moe_intermediate_size 2048 384
n_routed_experts 256 8
num_experts_per_tok 8 4
num_attention_heads 128 8
num_key_value_heads 128 8
q_lora_rank 1536 512
kv_lora_rank 512 256

The saved config.json is authoritative. The original backbone depth is retained to match upstream MTP checkpoint indexing.

Validation and scope

Two-GPU load, FP8_DYNAMIC quantization, normal sharded save, and vLLM generation passed. The FP8 MTP run drafted 60 tokens and accepted 0; its greedy output matched ordinary FP8 generation. These are execution checks, not a speedup measurement.

The reloaded BF16 backbone toy perplexity was 1.768900 on the same small corpus used for training. This demonstrates learning/reload integrity, not generalization or benchmark quality.

MTP projections are synthetic initializations and decoder weights were copied from trained backbone blocks. MTP was not separately trained, so these artifacts do not establish draft acceptance quality or inference acceleration.

The tested model-based environment used Transformers 5.17.0, Torch 2.14.0+cu130, LLM Compressor 2d52420, and compressed-tensors e69c8dc. Serving smoke checks used vLLM 0.30.0. GLM/DeepSeek model-based MTP loading failed on Transformers 5.15.0; 5.16 was not tested. Calibration-dependent MTP quantization remains follow-up work.

The model class/loader was not patched to make tests pass. Construction changes were configuration reductions and explicit synthetic checkpoint fixture creation. Structured results and configuration provenance are in validation.json; artifact-manifest.json records checkpoint file hashes.

Backbone loading

With the tested Transformers environment (and compressed-tensors for FP8), this loads the backbone. MTP is opt-in; ordinary Transformers backbone generation does not execute it.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "inference-optimization/DeepSeek-V3-0.86B-MTP-PR3225-FP8-Dynamic"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, dtype=torch.bfloat16, attn_implementation="eager",
)
inputs = tokenizer("The capital of France is", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=16, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Training

Random seed 3225; AdamW with learning rate 0.0004 and weight decay 0.01; batch size 2; text truncated to 160 tokens. Training stopped after three consecutive checks at toy perplexity <=3. The corpus follows the repository tiny-model workflow, with one additional short saying. These are the trained/reloaded artifacts used in testing.

License and attribution

Upstream architecture/tokenizer provenance and license: deepseek-ai/DeepSeek-V3-Base. The upstream DeepSeek license files are included; see LICENSE-MODEL and its use conditions. This fixture changes configuration dimensions, replaces the original weights with randomly initialized toy-trained weights, adds synthetic MTP fixture weights, and, for the FP8 variant, quantizes the trained checkpoint.

Downloads last month
6
Safetensors
Model size
0.9B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic

Quantized
(1)
this model

Collections including inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic