vLLM-ready version of zerank-2

#13
by Knid - opened

Hi ZeroEntropy team and everyone πŸ‘‹

Thanks for open-sourcing zerank-2! We wanted to serve it with vLLM, which turned out to be a bit tricky (the server starts fine but the scores are quietly wrong unless a few details are right), so we published a converted checkpoint:

πŸ‘‰ polaria-tech/zerank-2-reranker-vllm

  • Same weights, same scores as predict() (checked against the original in fp32, up to 38k tokens), and /rerank directly returns the sigmoid(logit / 5) score.
  • About 2x faster than sentence-transformers on an H100 (0.62 s vs 1.19 s for 100 documents).
  • Works with the instruction field of vLLM's rerank API.

Conversion scripts, parity checks and benchmarks: github.com/polaria-tech/zerank-2-reranker-vllm.

Feedback welcome, and happy to adjust anything if you'd like it done differently!

Small follow-up: cc @dilawarm @ghita-ha @npip99 in case it's useful for former API users who need to self-host since the API shutdown. Happy to help if you'd like it linked from the model card or the docs.

Sign up or log in to comment