Laya-Vision

Image in, calibrated typed decision out. Laya-Vision gives Laya image states. A frozen SigLIP-base/16 tower is pooled to 49 tokens, which are projected into Laya's embedding space. The model then answers choice, noul (yes/no) and score questions about the image with calibrated probabilities. Text-only requests take Laya's normal text path.

Usage

Requires the laya_vision package from the project repo (pip install -e .).

import laya_vision

agent = laya_vision.load("ZeraG07/laya-vision", device="cuda")   # or "cpu"; private repo: set HF_TOKEN
out = agent.predict(
    {"image": "photo.jpg"},            # path, bytes, base64 or PIL image; optional "text": "..."
    {
        "land":  {"type": "choice", "instructions": "What land cover is shown?",
                  "criteria": ["river", "forest", "highway", "farmland"]},
        "water": {"type": "noul",   "instructions": "Is there water in this image?"},
        "green": {"type": "score",  "instructions": "How much vegetation is visible?",
                  "criteria": ["none", "some", "a lot"]},
    },
)
out["answers"]["water"]["noul"]   # calibrated P(yes)

HTTP server, using Laya's /v1/systemone contract with base64 images:

python -m laya_vision.serve --checkpoint ZeraG07/laya-vision --port 8000 --device cuda

Results

Held-out test splits. The acceptance criteria come from the project's TRAINING.md. All four checks pass.

Task Laya-Vision caption→Laya (BLIP caption + stock Laya) stock Laya
EuroSAT (choice) 0.970 0.389
Oxford-IIIT Pets (choice + "is this a …?") 0.765 0.280
VQAv2 yes/no 0.608 0.592
BoolQ (text) 0.747 0.756
AG News (text) 0.924 0.926
  • Image calibration: ECE@10 is 0.014.
  • Latency: p50 GPU latency for one question is 1.68× text-only Laya (RTX 4060 Ti).

Training

The recipe is "B" in TRAINING.md. Everything was trained on one 8 GB RTX 4060 Ti.

Stage 1 (alignment). COCO caption matching with easy negatives (444k rows), plus 10% typed-decision text rows.

  • Trained: the projector at LR 1e-4, and LoRA (r=16) on Laya's encoder.
  • Projector guards: an input LayerNorm, running standardization of the output, and a frozen modality embedding. They prevent the modality collapse seen with a projector-only stage 1 at LR 1e-3.

Stage 2 (decisions). Ran for 2 epochs with LoRA, from stage 1.

  • Image tasks: CIFAR-10, Oxford Pets, Food-101, EuroSAT, A-OKVQA, ScienceQA (image subset) and VQAv2 yes/no.
  • Text: 30% typed-decision rows.

Stage 2c. A continuation that adds 160k VQAv2 yes/no questions from the train split.

WiSE-FT. Laya's weights are interpolated toward stock Laya: stock + 0.85 · (fine-tuned − stock).

  • This limits text regression.
  • α was chosen on a BoolQ train-split slice and on the image calib splits, never on the test sets.

Calibration. Per-modality temperature scaling on the calib splits.

Inside the checkpoint:

  • vision.safetensors holds the pooler and projector. The SigLIP tower is frozen and downloaded from the Hub.
  • model.safetensors is a plain Laya state dict.
  • rl_agent_config.json["vision"] records the tower, the projector flags, wise_alpha and the image temperatures.

Limitations

  • Resolution. Images are seen at 224×224 as 49 pooled tokens, so fine detail is limited. VQAv2 is only slightly above the caption baseline.
  • Untested tasks. KonIQ (quality score) was not trained or evaluated. score questions about images work mechanically but are uncalibrated for image tasks.
  • Possible image overlap. The stage-1 COCO captions (Karpathy split) include some val2014 images, and the VQAv2 test questions also come from val2014. Answers never overlapped.

License

This model is released under CC BY-NC-SA 4.0 (non-commercial), the most restrictive licence among its training data (ScienceQA).

  • Laya and SigLIP are Apache-2.0.
  • Several training sets have unclear or research-only image licences (COCO/Flickr, CIFAR-10, Food-101).
  • See laya_vision/data/LICENSES.md in the project repo.
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZeraG07/laya-vision

Finetuned
(160)
this model