Instructions to use lightonai/LightOn-rerank-PW-0.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lightonai/LightOn-rerank-PW-0.8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="lightonai/LightOn-rerank-PW-0.8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("lightonai/LightOn-rerank-PW-0.8B") model = AutoModelForMultimodalLM.from_pretrained("lightonai/LightOn-rerank-PW-0.8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - sentence-transformers
How to use lightonai/LightOn-rerank-PW-0.8B with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("lightonai/LightOn-rerank-PW-0.8B") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lightonai/LightOn-rerank-PW-0.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lightonai/LightOn-rerank-PW-0.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/LightOn-rerank-PW-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/lightonai/LightOn-rerank-PW-0.8B
- SGLang
How to use lightonai/LightOn-rerank-PW-0.8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lightonai/LightOn-rerank-PW-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/LightOn-rerank-PW-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lightonai/LightOn-rerank-PW-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/LightOn-rerank-PW-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use lightonai/LightOn-rerank-PW-0.8B with Docker Model Runner:
docker model run hf.co/lightonai/LightOn-rerank-PW-0.8B
Integrate with Sentence Transformers
Hello!
Preface
Congrats on the LightOn-rerank release! Nowadays, Sentence Transformers also supports multimodal pointwise rerankers, so I'd love to help integrate these. In the future, I'll also look at multimodal listwise rerankers, and yours should be a good example of models to support by then.
Heads up, this PR was AI-generated and human-reviewed.
Pull Request overview
- Integrate this model with Sentence Transformers (v5.4.0+) as a
CrossEncoder
Details
This PR adds config-only support for loading this model with Sentence Transformers, so it can be used for text and visual document reranking via the familiar CrossEncoder API: model.predict(...) and model.rank(...). No weights were changed.
The module pipeline is Transformer(any-to-any) -> LogitScore(true_token_id=9175, false_token_id=2665): Sentence Transformers loads the backbone via AutoModelForMultimodalLM, forces left padding, computes logits for the final position only (logits_to_keep=1), and LogitScore returns logit("Yes") - logit("No"), i.e. exactly the README scoring.
The trained system prompt and both user templates are baked into a dedicated additional_chat_templates/reranker.jinja, which is selected through processing_kwargs in sentence_bert_config.json. Sentence Transformers renders each (query, document) pair through that template: text documents produce the text template, documents containing an image produce the vision template with the image tokens first (a document with both renders the vision template plus a trailing Document: {text} block, mirroring the text format). The generation prompt emits the non-thinking <|im_start|>assistant\n<think>\n\n</think>\n\n suffix, matching where the model reads Yes/No logits. Passing model.predict(pairs, prompt="...") overrides the system prompt if anyone wants to experiment, but by default the trained prompt is used verbatim, with no user-side formatting needed.
Scores were verified against the README's plain-transformers implementation in fp32 on text-only, image-only, and mixed text+image batches: max |diff| is about 3e-5. This was checked on the released sentence-transformers 5.4.0 and 5.6.0 with transformers 5.4.0 and 5.13.1. The default (no-kwargs) load uses the checkpoint's bfloat16, and model_kwargs={"dtype": ..., "attn_implementation": ...} pass through for anyone who wants fp32 or flash attention.
Added files:
modules.json: the module pipeline (Transformer->LogitScore)sentence_bert_config.json:transformer_task, modality config, and chat template selection for theTransformermoduleconfig_sentence_transformers.json:CrossEncodermetadata (identity activation, no prompts)1_LogitScore/config.json: the Yes/No token idsadditional_chat_templates/reranker.jinja: the trained reranking format as a chat template overquery/documentrole messages
Modified files:
README.md: added thesentence-transformerstag and a "Using Sentence Transformers" usage section
from sentence_transformers import CrossEncoder
model = CrossEncoder("lightonai/LightOn-rerank-PW-0.8B")
query = "What is late interaction in neural information retrieval?"
documents = [
"ColBERT computes token-level query-document interactions at search time...",
"The Eiffel Tower is located on the Champ de Mars in Paris.",
]
pairs = [(query, doc) for doc in documents]
scores = model.predict(pairs)
print(scores)
# [-4.125 -7.1875]
rankings = model.rank(query, documents)
print(rankings)
# [{'corpus_id': 0, 'score': -4.125}, {'corpus_id': 1, 'score': -7.1875}]
Page images can be passed directly as documents (PIL.Image, URL, or file path), and text and image candidates can be mixed in one predict call.
Note that none of the existing usage is affected: the transformers and vLLM paths work exactly as before. This only adds an additional way to run the model in a familiar and common format. The same integration applies near-verbatim to PW-2B and PW-4B.
Happy to tweak anything you'd like changed. Please let me know if you have any questions or feedback!
- Tom Aarsen