Instructions to use ZeraG07/laya-vision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use ZeraG07/laya-vision with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya-Vision
Image in, calibrated typed decision out. Laya-Vision gives Laya image states. A frozen SigLIP-base/16 tower is pooled to 49 tokens, which are projected into Laya's embedding space. The model then answers choice, noul (yes/no) and score questions about the image with calibrated probabilities. Text-only requests take Laya's normal text path.
Usage
Requires the laya_vision package from the project repo (pip install -e .).
import laya_vision
agent = laya_vision.load("ZeraG07/laya-vision", device="cuda") # or "cpu"; private repo: set HF_TOKEN
out = agent.predict(
{"image": "photo.jpg"}, # path, bytes, base64 or PIL image; optional "text": "..."
{
"land": {"type": "choice", "instructions": "What land cover is shown?",
"criteria": ["river", "forest", "highway", "farmland"]},
"water": {"type": "noul", "instructions": "Is there water in this image?"},
"green": {"type": "score", "instructions": "How much vegetation is visible?",
"criteria": ["none", "some", "a lot"]},
},
)
out["answers"]["water"]["noul"] # calibrated P(yes)
HTTP server, using Laya's /v1/systemone contract with base64 images:
python -m laya_vision.serve --checkpoint ZeraG07/laya-vision --port 8000 --device cuda
Results
Held-out test splits. The acceptance criteria come from the project's TRAINING.md. All four checks pass.
| Task | Laya-Vision | caption→Laya (BLIP caption + stock Laya) | stock Laya |
|---|---|---|---|
| EuroSAT (choice) | 0.970 | 0.389 | |
| Oxford-IIIT Pets (choice + "is this a …?") | 0.765 | 0.280 | |
| VQAv2 yes/no | 0.608 | 0.592 | |
| BoolQ (text) | 0.747 | 0.756 | |
| AG News (text) | 0.924 | 0.926 |
- Image calibration: ECE@10 is 0.014.
- Latency: p50 GPU latency for one question is 1.68× text-only Laya (RTX 4060 Ti).
Training
The recipe is "B" in TRAINING.md. Everything was trained on one 8 GB RTX 4060 Ti.
Stage 1 (alignment). COCO caption matching with easy negatives (444k rows), plus 10% typed-decision text rows.
- Trained: the projector at LR 1e-4, and LoRA (r=16) on Laya's encoder.
- Projector guards: an input LayerNorm, running standardization of the output, and a frozen modality embedding. They prevent the modality collapse seen with a projector-only stage 1 at LR 1e-3.
Stage 2 (decisions). Ran for 2 epochs with LoRA, from stage 1.
- Image tasks: CIFAR-10, Oxford Pets, Food-101, EuroSAT, A-OKVQA, ScienceQA (image subset) and VQAv2 yes/no.
- Text: 30% typed-decision rows.
Stage 2c. A continuation that adds 160k VQAv2 yes/no questions from the train split.
WiSE-FT. Laya's weights are interpolated toward stock Laya: stock + 0.85 · (fine-tuned − stock).
- This limits text regression.
- α was chosen on a BoolQ train-split slice and on the image calib splits, never on the test sets.
Calibration. Per-modality temperature scaling on the calib splits.
Inside the checkpoint:
vision.safetensorsholds the pooler and projector. The SigLIP tower is frozen and downloaded from the Hub.model.safetensorsis a plain Laya state dict.rl_agent_config.json["vision"]records the tower, the projector flags,wise_alphaand the image temperatures.
Limitations
- Resolution. Images are seen at 224×224 as 49 pooled tokens, so fine detail is limited. VQAv2 is only slightly above the caption baseline.
- Untested tasks. KonIQ (quality score) was not trained or evaluated.
scorequestions about images work mechanically but are uncalibrated for image tasks. - Possible image overlap. The stage-1 COCO captions (Karpathy split) include some val2014 images, and the VQAv2 test questions also come from val2014. Answers never overlapped.
License
This model is released under CC BY-NC-SA 4.0 (non-commercial), the most restrictive licence among its training data (ScienceQA).
- Laya and SigLIP are Apache-2.0.
- Several training sets have unclear or research-only image licences (COCO/Flickr, CIFAR-10, Food-101).
- See
laya_vision/data/LICENSES.mdin the project repo.
- Downloads last month
- -
Model tree for ZeraG07/laya-vision
Base model
convaiinnovations/laya