vicuna_clip_patch

Mixture of Layers (MoL): patch_layer routing with k=1 and clip.

This release preserves the selected checkpoint weights without merging, retraining, quantizing, or changing their tensor values. See inference_config.json for the routing settings, base model, and conversation template.

Contents and loading

This contains the full sharded language-model checkpoint, MoL weights, and tokenizer. The vision/text encoder dependencies are constructed by MoL.

Use the custom MoL model implementation from https://github.com/SStoica12/MoL. These files are in the repository's original checkpoint format; generic AutoModelForCausalLM or PEFT-only loading is insufficient for multimodal routing. Download the repository snapshot to a local directory before using the MoL loader. Model source version at packaging: 398a93bd9e77a67cdd786d6d1ef9451b59ac8845 on layerwise_coeff; the release manifest identifies this as a working checkout, not a tested standalone Hub model. For Phi, the runtime must support the Phi model class as well as the SigLIP router.

Associated evaluation scripts live under scripts/v1_5/eval in the MoL checkout: vstar.sh, mmstar.sh, hrbench8k.sh, hrbench4k.sh, realworldqa.sh, naturalbench.sh, and charxiv.sh. Select this local package explicitly rather than relying on the scripts' historical default checkpoint arrays. Use the adapter base model and routing settings recorded in inference_config.json.

Evaluation and provenance

Original experiment directory: llava-v1.5-13b-moe-finetune_text-cond-attn-patchlayer-router_learnable-gate_lr1e-5-pretrain_lr2e-5-finetune. evaluation_evidence.json contains the recovered scores for this selected run. Some evidence is archival; these are not newly reproduced scores. A paper row may combine multiple checkpoints, so it must not be used as this model's scorecard.

source_config.json preserves the original configuration; release_manifest.json records the source, metadata-only configuration additions and file hashes. Training logs, optimizer states, and intermediate checkpoint directories are omitted. SHA256 copy verification and safetensors structure/index checks were performed; full model inference was not run as part of packaging.

Usage and dependencies

Intended for research on visual question answering and layer routing. Vision encoders and, for adapters, the base language model are external dependencies. The release does not grant new rights to the upstream models; consult their model cards and applicable terms. No new license is assigned by this packaging step.

Model card format: https://huggingface.co/docs/hub/model-cards

Downloads last month
13
Safetensors
Model size
13B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wjdghks950/vicuna_clip_patch

Finetuned
(11)
this model