Image-Text-to-Text
Transformers
Safetensors
deepseek_v41
text-generation
Eval Results
8-bit precision
fp8
Instructions to use deepseek-ai/DeepSeek-V4.1-Flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use deepseek-ai/DeepSeek-V4.1-Flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="deepseek-ai/DeepSeek-V4.1-Flash")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-V4.1-Flash", device_map="auto") - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use deepseek-ai/DeepSeek-V4.1-Flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "deepseek-ai/DeepSeek-V4.1-Flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4.1-Flash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/deepseek-ai/DeepSeek-V4.1-Flash
- SGLang
How to use deepseek-ai/DeepSeek-V4.1-Flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "deepseek-ai/DeepSeek-V4.1-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4.1-Flash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "deepseek-ai/DeepSeek-V4.1-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4.1-Flash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use deepseek-ai/DeepSeek-V4.1-Flash with Docker Model Runner:
docker model run hf.co/deepseek-ai/DeepSeek-V4.1-Flash
Instantly parse Hugging Face & ModelScope safetensors metadata without downloading weights. View tensor shapes, dtypes, and run side-by-side model diffs.
🔥 1
#33 opened about 3 hours ago
by
alone-wl
smaller model with engram?
🔥 2
8
#32 opened about 7 hours ago
by
ProCreations
More long-context evaluation results?
#31 opened about 16 hours ago
by
ArlenSmith
Hey DeepSeek, could you avoid using such confusing model IDs on the API platform?
4
#30 opened about 19 hours ago
by
RainPPR
Add community evaluation results
#29 opened about 19 hours ago
by
SaylorTwift
Running on 4x RTX PRO 6000 with NVMe offload for ngram
🚀 6
2
#28 opened about 20 hours ago
by
0xSero
Love to see the Harness Benchmark!!! TY!!!
👍 1
#26 opened about 21 hours ago
by
darkmatter2222
哇、DeepSeek!
🔥 1
1
#25 opened about 21 hours ago
by
NILKNARFGonzo
no way w deepseek
#24 opened about 21 hours ago
by
puihl481723
很强,参数量比上个版本翻倍,最强的flash模型,unsloth 早点出量化版本,赞美这些开源大模型
1
#22 opened about 22 hours ago
by
zmw911
感谢Deepseek
🤝 1
#21 opened about 23 hours ago
by
vayne1993
<a href=https://evil.com>hello</a>
#20 opened about 23 hours ago
by
tester9632587411
OpenAI and Claude don't make me download half a terabyte of weight just to ask a question smh
🧠🤯 26
7
#18 opened 1 day ago
by
Mikkkkoooo
Has the model's alignment with human ethics been strengthened compared to the previous generation?
2
#17 opened 1 day ago
by
likewendy
Bro....500多B的Flash,8卡H200已经上不了桌了吗[cry]
7
#16 opened 1 day ago
by
Saito-Karuha
Back to "attention is all you need"
🔥 1
#15 opened 1 day ago
by
shadowlilac
update README to add vLLM inference
🔥 1
#13 opened 1 day ago
by
riverclouds
Restore each indexer's K cache on incomplete compression steps
1
#12 opened 1 day ago
by
ZenAlexa
DeepSeek V4.1 Flash Lite
➕👍 32
14
#11 opened 1 day ago
by
KeinNiemand
为啥简单问题也强行输出小作文
🤯 1
3
#9 opened 1 day ago
by
qwq95195
Thanks for open-sourcing,感谢开源 DeepSeek V4.1 Flash
🚀🔥 16
1
#8 opened 1 day ago
by
lcc4567
Very impressive!
🧠 1
#7 opened 1 day ago
by
hgeist
Update README.md
#6 opened 1 day ago
by
zjxia
出来溜达一圈,等待社区反馈
#5 opened 1 day ago
by
JasonShane
梁圣!
🤗🚀 10
1
#4 opened 1 day ago
by
tom1ong
非常不错 感谢劳梁
🤗👍 18
#3 opened 1 day ago
by
zengzeng666
Thank you!
👍 18
#2 opened 1 day ago
by
KilincocomilK
liangsheng
🤗👍 78
2
#1 opened 1 day ago
by
TOURCMD