LT-OPD

Paper · GitHub

LT-OPD trains vision-language models with fewer visual tokens through on-policy self-distillation. This Qwen3.5-4B model combines CDPruner with an MLP that aggregates discarded tokens into their nearest retained tokens, preserving their original positions without adding tokens.

Quick Start

Use Python 3.12 and the CUDA environment in the installation guide.

git clone --depth 1 https://github.com/Yrxxxxxxxx1007/LT-OPD.git
cd LT-OPD
pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu126
pip install -e '.[eval]'
pip install flash-attn==2.8.3 --no-build-isolation

hf download yyy051007/LT-OPD \
  --local-dir models/LT-OPD

Load the model with LT-OPD's visual-token compression runtime:

import torch
from evaluation.runtime import build_route_query
from lt_opd import load_compression_runtime

runtime = load_compression_runtime(
    "models/LT-OPD",
    implementation="current",
    torch_dtype=torch.bfloat16,
    device_map={"": "cuda:0"},
)
question = "What is written on the sign?"
messages = [{
    "role": "user",
    "content": [
        {"type": "image", "image": "/path/to/image.jpg"},
        {"type": "text", "text": question},
    ],
}]
output = runtime.generate(
    [messages],
    route_queries_batch=[build_route_query(question)],
)
print(output["decoded_predictions"][0])

Citation

If you find our code, model, or dataset useful, please kindly cite our paper:

@article{li2026fewer,
  title={Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction},
  author={Li, Junxian and Yang, Ruixuan and Zhang, Tianao and Xu, Tiange and Dong, Weisheng and Zhang, Yulun},
  journal={arXiv preprint arXiv:2609.32353},
  year={2026}
}
Downloads last month
56
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yyy051007/LT-OPD

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(917)
this model

Paper for yyy051007/LT-OPD