Sm120/rtx6000pro supported?
Is sm120 supported?
Anyone got it to run on 2x rtx 6k have a recipe?
Thanks,
working on it right now, well i'm working on a 4x RTX PRO 6000 recipe. I'm not sure if this will fly on a 2x rig, since the weights themselves are bigger than 192GB.
Please share once the work is complete.
After spending a LOT of tokens with another AI helping me nail down the config, I finally got something working but it's a bit broken and some features aren't working and I had to patch the sglang and build my own container. Not exactly a smooth result. I'm going to see if I can make it easy to run with all the features enabled. I'm also only getting about 10 concurrent sessions at 200k context. Hopefully someone smarter than me, or with more tokens and time to iterate the config, will post something better than what I came up with.
https://github.com/0xSero/glm-5.3-flash-sglang-sm120 Has anyone actually tried this method?
I wouldn't use anything he puts up as he tends to take the work of others, claiming it as his own, without attribution. And he manifestly doesn't know what he is doing. He used my work without attribution 11 separate times over the course of the last month.
Model card got updated as well as config! Tuning stuff but it works with sglang now.