ShotPilot Full-Five-Evidence Subject-v2 LoRA

This archive contains all three epoch checkpoints and reproducibility artifacts for the strict subject-v2 ablation.

Controlled Difference

The run uses the same 2,483 image pairs, prompt, full-five-evidence serialization, ordering, evaluation set, and optimization configuration as the preceding full-five-evidence run. Only the subject dimension was replaced on 1,742 covered rows. Train JSONL SHA256: 9a1d57c5210c4596d4ebce78a159105c85632d6ff0f91eeb81f1c7e5a2c4d586.

Training

  • Qwen3-VL-8B-Instruct
  • LoRA rank 64, alpha 32, dropout 0.05
  • 3 epochs; checkpoints 156/312/468
  • learning rate 5e-6, weight decay 0.1, warmup ratio 0.03, cosine schedule
  • global batch size 16, BF16, max sequence length 4096, image size 672, seed 42

Evaluation

Each checkpoint was evaluated on fixed Eval100_new with the fixed free-review prompt, deterministic decoding, and max_new_tokens=768. Ground-truth fallback and pair judging were disabled.

Primary outputs remain untouched. Repaired outputs replace only rows with loop_8gram_max >= 4 using deterministic no-repeat decoding and retain provenance.

Checkpoint Primary loops Repaired loops Repaired rows
156 52 0 52
312 10 0 10
468 10 0 10

Every primary and repaired output contains exactly 100 unique evaluation IDs. Repaired outputs are decoding interventions and must not be reported as untouched model output.

Archive

  • checkpoints/: adapter weights, adapter configs, trainer states, and checkpoint READMEs
  • nonweights.tar: train/eval JSONL, scripts, summaries, token audit, and logs
  • manifest.json: byte sizes and SHA256 checksums

Optimizer state, base-model weights, and images are excluded.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for purefall/shotpilot-e5-subject-v2-qwen3vl8b-lora

Adapter
(162)
this model