Deployment question: Single DGX‑Spark 128GB expert‑paging & agent tool‑calling
#4
by mitu159357 - opened
Hi lovesenko,
I want to run your DeepSeek‑V4‑Flash‑DSpark‑Abliterated on one single DGX‑Spark (128 GB unified RAM), using eugr/spark‑vllm‑b12x with expert‑paging SSD swap, single session only, 256K context limit, then connect it to local DeepSeek‑Harness for agent tool calls.
Quick questions:
- Is the DSpark speculative‑decoding head and tool‑calling function fully preserved after abliteration?
- Are there any known stability issues running this model on a single 128 GB DGX‑Spark?
- Is 1M‑token context possible for testing on one node?
Thanks a lot.