Tensor parallel

#13
by krampenschiesser - opened

For those with a lot of cheap gpu's.
My tensor parallel fix for stepfun also applies to laguna:
https://github.com/ggml-org/llama.cpp/pull/24554

model size params backend ngl sm fa mmap test t/s
laguna ?B Q5_K - Medium 82.02 GiB 117.56 B CUDA 999 tensor 1 0 pp512 1486.28 ± 19.16
laguna ?B Q5_K - Medium 82.02 GiB 117.56 B CUDA 999 tensor 1 0 tg128 68.79 ± 1.04
laguna ?B Q5_K - Medium 82.02 GiB 117.56 B CUDA 999 layer 1 0 pp512 854.33 ± 3.29
laguna ?B Q5_K - Medium 82.02 GiB 117.56 B CUDA 999 layer 1 0 tg128 50.05 ± 0.02

Sign up or log in to comment