SSVAL-7B

This is the released checkpoint of SSVAL (Spatial-Spectral Visual Anchor Learning) built on top of Qwen2.5-7B-Instruct with a CLIP ViT-L/14-336 vision tower. It is the artefact produced by the two-stage pipeline described in the paper:

  1. Projector pretraining on LLaVA-Pretrain LCS-558K.
  2. Visual instruction tuning on LLaVA-665K with SPARA + SPERA losses.

The checkpoint ships as a PEFT LoRA adapter + mm-projector: it is not a standalone model. To use it you must pair it with the base LLM Qwen/Qwen2.5-7B-Instruct and the vision tower openai/clip-vit-large-patch14-336.

Files

File Size What it is
adapter_model.safetensors ~620 MB LoRA adapter (rank 128, alpha 256) over the LLM's attention and MLP projections
non_lora_trainables.bin ~170 MB Multimodal projector and any other non-LoRA trainables saved separately
adapter_config.json <1 KB PEFT adapter configuration
config.json ~4 KB Full model configuration including vision-side and CPD settings
trainer_state.json ~3 MB Training-step / loss history (optional, kept for reproducibility)

Required base dependencies

# pip install: transformers==4.37.2, torch>=2.1, peft>=0.10,
#              Pillow, accelerate, einops

See the companion code repository msls38/SSVAL_code for the load_pretrained_model helper, evaluation scripts, and full dependency pins.

Vision encoder and VFM teacher

Component Choice
Vision tower (mm_vision_tower) clip-vit-large-patch14-336
Patch selection layer (mm_vision_select_layer) -2 (second-to-last)
Feature type (mm_vision_select_feature) patch
Image aspect ratio pad
Patch merge flat
MM projector type mlp2x_gelu
Image-token wrapping <image> only (no <im_start> / <im_end>)

VFM teacher (training-time only, not needed at inference).

Field Value
vfm_target dinov2-vit-b
num_vision_prompts 4
num_vfm_prompts 4
cpd_target_layers [14]
Multi-scale factors [1, 2, 4]

How to load and use

The companion repo's evaluation entry points expect the following on-disk layout:

checkpoints/
โ”œโ”€โ”€ Qwen2.5-7B-Instruct/         # base LLM, from Qwen/Qwen2.5-7B-Instruct
โ””โ”€โ”€ clip-vit-large-patch14-336/  # vision tower, from openai/clip-vit-large-patch14-336

Then, for single-image inference:

python SSVAL/src/inference.py \
    --model-path ./SSVAL-7B \
    --model-base ./checkpoints/Qwen2.5-7B-Instruct \
    --image-path ./my_image.jpg \
    --prompt "Describe this image."

For benchmark evaluation, e.g. MME:

python SSVAL/src/eval_mme.py \
    --model-path ./SSVAL-7B \
    --model-base ./checkpoints/Qwen2.5-7B-Instruct \
    --mme-path ./playground/benchmarks/MME/MME_Benchmark

The benchmark scripts run end-to-end without modification once the companion code is installed and the dataset / benchmark directories are populated as described in the repository README.

Citation

If you use this checkpoint, please cite the SSVAL paper (full bibtex to be added once the paper is publicly available).

License

  • The LoRA weights and configuration in this repository are released under the Apache License 2.0 (matching the SSVAL code repository).
  • The base LLM Qwen2.5-7B-Instruct is subject to its own license (Qwen2.5 LICENSE).
  • The vision tower is subject to the OpenAI CLIP LICENSE.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nmslmgl/SSVAL-7B

Base model

Qwen/Qwen2.5-7B
Adapter
(2573)
this model