Integrate with Sentence Transformers

#1
by tomaarsen HF Staff - opened

Hello!

Preface

Congrats on the LightOn-rerank release! Nowadays, Sentence Transformers also supports multimodal pointwise rerankers, so I'd love to help integrate these. In the future, I'll also look at multimodal listwise rerankers, and yours should be a good example of models to support by then.

Heads up, this PR was AI-generated and human-reviewed.

Pull Request overview

  • Integrate this model with Sentence Transformers (v5.4.0+) as a CrossEncoder

Details

This PR adds config-only support for loading this model with Sentence Transformers, so it can be used for text and visual document reranking via the familiar CrossEncoder API: model.predict(...) and model.rank(...). No weights were changed.

The module pipeline is Transformer(any-to-any) -> LogitScore(true_token_id=9175, false_token_id=2665): Sentence Transformers loads the backbone via AutoModelForMultimodalLM, forces left padding, computes logits for the final position only (logits_to_keep=1), and LogitScore returns logit("Yes") - logit("No"), i.e. exactly the README scoring.

The trained system prompt and both user templates are baked into a dedicated additional_chat_templates/reranker.jinja, which is selected through processing_kwargs in sentence_bert_config.json. Sentence Transformers renders each (query, document) pair through that template: text documents produce the text template, documents containing an image produce the vision template with the image tokens first (a document with both renders the vision template plus a trailing Document: {text} block, mirroring the text format). The generation prompt emits the non-thinking <|im_start|>assistant\n<think>\n\n</think>\n\n suffix, matching where the model reads Yes/No logits. Passing model.predict(pairs, prompt="...") overrides the system prompt if anyone wants to experiment, but by default the trained prompt is used verbatim, with no user-side formatting needed.

Scores were verified against the README's plain-transformers implementation in fp32 on text-only, image-only, and mixed text+image batches: max |diff| is about 3e-5. This was checked on the released sentence-transformers 5.4.0 and 5.6.0 with transformers 5.4.0 and 5.13.1. The default (no-kwargs) load uses the checkpoint's bfloat16, and model_kwargs={"dtype": ..., "attn_implementation": ...} pass through for anyone who wants fp32 or flash attention.

Added files:

  • modules.json: the module pipeline (Transformer -> LogitScore)
  • sentence_bert_config.json: transformer_task, modality config, and chat template selection for the Transformer module
  • config_sentence_transformers.json: CrossEncoder metadata (identity activation, no prompts)
  • 1_LogitScore/config.json: the Yes/No token ids
  • additional_chat_templates/reranker.jinja: the trained reranking format as a chat template over query/document role messages

Modified files:

  • README.md: added the sentence-transformers tag and a "Using Sentence Transformers" usage section
from sentence_transformers import CrossEncoder

model = CrossEncoder("lightonai/LightOn-rerank-PW-0.8B")

query = "What is late interaction in neural information retrieval?"
documents = [
    "ColBERT computes token-level query-document interactions at search time...",
    "The Eiffel Tower is located on the Champ de Mars in Paris.",
]

pairs = [(query, doc) for doc in documents]
scores = model.predict(pairs)
print(scores)
# [-4.125  -7.1875]

rankings = model.rank(query, documents)
print(rankings)
# [{'corpus_id': 0, 'score': -4.125}, {'corpus_id': 1, 'score': -7.1875}]

Page images can be passed directly as documents (PIL.Image, URL, or file path), and text and image candidates can be mixed in one predict call.

Note that none of the existing usage is affected: the transformers and vLLM paths work exactly as before. This only adds an additional way to run the model in a familiar and common format. The same integration applies near-verbatim to PW-2B and PW-4B.

Happy to tweak anything you'd like changed. Please let me know if you have any questions or feedback!

  • Tom Aarsen
tomaarsen changed pull request status to open
NohTow changed pull request status to merged

Sign up or log in to comment