Deployment question: Single DGX‑Spark 128GB expert‑paging & agent tool‑calling

#4
by mitu159357 - opened

Hi lovesenko,
I want to run your DeepSeek‑V4‑Flash‑DSpark‑Abliterated on one single DGX‑Spark (128 GB unified RAM), using eugr/spark‑vllm‑b12x with expert‑paging SSD swap, single session only, 256K context limit, then connect it to local DeepSeek‑Harness for agent tool calls.

Quick questions:

  1. Is the DSpark speculative‑decoding head and tool‑calling function fully preserved after abliteration?
  2. Are there any known stability issues running this model on a single 128 GB DGX‑Spark?
  3. Is 1M‑token context possible for testing on one node?

Thanks a lot.

Sign up or log in to comment