Instructions to use inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic
- SGLang
How to use inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic with Docker Model Runner:
docker model run hf.co/inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic
DeepSeek-V3-0.86B-MTP-PR3225-FP8-Dynamic
FP8_DYNAMIC through the model pipeline for LLM Compressor PR #3225. This is a small architecture and checkpoint test fixture, not a production language model. Its backbone was initialized randomly and trained on the repository tiny-model skill's toy text corpus; no pretrained base-model weights were used.
Model and provenance
- Original architecture/tokenizer: deepseek-ai/DeepSeek-V3-Base, revision
afb92e1fa402c2be2a9eb085312bb02e0384d6c7. - Transformers class:
DeepseekV3ForCausalLM. - Total source parameters including synthetic MTP: 860,930,560 (0.861B).
- Backbone parameters: 849,041,408; MTP parameters: 11,889,152.
- Estimated active backbone parameters: 643,782,656. This routing estimate includes embeddings/dense components, excludes MTP, and is not a FLOP measurement.
- Precision: compressed-tensors FP8_DYNAMIC, with exclusions retained in their source dtype.
- Checkpoint: 3 indexed safetensors shards; 4,197 indexed tensors. Tokenizer assets are included.
Configuration
| Field | Original | Tiny |
|---|---|---|
num_hidden_layers |
61 | 61 |
hidden_size |
7168 | 768 |
intermediate_size |
18432 | 3072 |
moe_intermediate_size |
2048 | 384 |
n_routed_experts |
256 | 8 |
num_experts_per_tok |
8 | 4 |
num_attention_heads |
128 | 8 |
num_key_value_heads |
128 | 8 |
q_lora_rank |
1536 | 512 |
kv_lora_rank |
512 | 256 |
The saved config.json is authoritative. The original backbone depth is retained to match upstream MTP checkpoint indexing.
Validation and scope
Two-GPU load, FP8_DYNAMIC quantization, normal sharded save, and vLLM generation passed. The FP8 MTP run drafted 60 tokens and accepted 0; its greedy output matched ordinary FP8 generation. These are execution checks, not a speedup measurement.
The reloaded BF16 backbone toy perplexity was 1.768900 on the same small corpus used for training. This demonstrates learning/reload integrity, not generalization or benchmark quality.
MTP projections are synthetic initializations and decoder weights were copied from trained backbone blocks. MTP was not separately trained, so these artifacts do not establish draft acceptance quality or inference acceleration.
The tested model-based environment used Transformers 5.17.0, Torch 2.14.0+cu130, LLM Compressor 2d52420, and compressed-tensors e69c8dc. Serving smoke checks used vLLM 0.30.0. GLM/DeepSeek model-based MTP loading failed on Transformers 5.15.0; 5.16 was not tested. Calibration-dependent MTP quantization remains follow-up work.
The model class/loader was not patched to make tests pass. Construction changes were configuration reductions and explicit synthetic checkpoint fixture creation. Structured results and configuration provenance are in validation.json; artifact-manifest.json records checkpoint file hashes.
Backbone loading
With the tested Transformers environment (and compressed-tensors for FP8), this loads the backbone. MTP is opt-in; ordinary Transformers backbone generation does not execute it.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "inference-optimization/DeepSeek-V3-0.86B-MTP-PR3225-FP8-Dynamic"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype=torch.bfloat16, attn_implementation="eager",
)
inputs = tokenizer("The capital of France is", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=16, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Training
Random seed 3225; AdamW with learning rate 0.0004 and weight decay 0.01; batch size 2; text truncated to 160 tokens. Training stopped after three consecutive checks at toy perplexity <=3. The corpus follows the repository tiny-model workflow, with one additional short saying. These are the trained/reloaded artifacts used in testing.
License and attribution
Upstream architecture/tokenizer provenance and license: deepseek-ai/DeepSeek-V3-Base. The upstream DeepSeek license files are included; see LICENSE-MODEL and its use conditions. This fixture changes configuration dimensions, replaces the original weights with randomly initialized toy-trained weights, adds synthetic MTP fixture weights, and, for the FP8 variant, quantizes the trained checkpoint.
- Downloads last month
- 6
Model tree for inference-optimization/DeepSeek-V3-0.86B-MTP-FP8-Dynamic
Base model
inference-optimization/DeepSeek-V3-0.86B-MTP