Instructions to use steven0226/defectforge-visa-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use steven0226/defectforge-visa-lora with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("sd2-community/stable-diffusion-2-inpainting,diffusers/stable-diffusion-xl-1.0-inpainting-0.1", torch_dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("steven0226/defectforge-visa-lora") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
DefectForge VisA Inpainting LoRAs
Per-object SD2 and SDXL inpainting LoRA adapters trained from 10 real anomalous VisA
training images for each of pcb1 and capsules.
繁中摘要:每個物件只使用 10 張真實瑕疵訓練圖,分別訓練 SD2/SDXL inpainting LoRA。權重必須搭配 ROI crop、binary mask 與 blend-back 流程使用;直接對 整張 AOI 影像做一般文字生圖,不是本模型的支援用法。
Files and base locks
lora_sd2/
pcb1/final/
capsules/final/
lora_sdxl/
pcb1/final/
capsules/final/
Each final directory includes the UNet adapter, learned token embedding adapter(s), tokenizer state, and final training configuration.
| Family | Base repository | Immutable revision | Resolution |
|---|---|---|---|
| SD2 | sd2-community/stable-diffusion-2-inpainting |
5f74973cbb64c8568780732c17f43eb269d63a0d |
512 |
| SDXL | diffusers/stable-diffusion-xl-1.0-inpainting-0.1 |
115134f363124c53c7d878647567d04daf26e41e |
1024 |
Trigger tokens
The tokens are frozen unsupervised pseudo-types, not official VisA labels.
| Object | Token | Frozen training-mask components |
|---|---|---|
| pcb1 | <pcb1-type0> |
16 |
| pcb1 | <pcb1-type1> |
7 |
| capsules | <capsules-type0> |
9 |
| capsules | <capsules-type1> |
3 |
Supported inference path
Use the repository pipeline so inference crops around the placed defect mask, runs the matching base model and adapter at its native resolution, then blends the generated ROI back into the original image. This preserves the surrounding AOI frame.
uv run python src/synthetic/generate_diffusion.py `
--config configs/generate_sd2.yaml `
--object pcb1 `
--n 1 `
--refine
The equivalent Python flow is:
# 1. Load a frozen normal image and a placement mask.
# 2. Expand the mask bounding box by the config crop_ratio.
# 3. Resize that ROI + mask to the base model's native resolution.
# 4. Inpaint with the object-specific LoRA and trigger token.
# 5. Resize the generated ROI back and feather/Poisson blend only inside the mask.
#
# The versioned implementation and exact metadata contract live in:
# src/synthetic/generate_diffusion.py
Do not bypass crop-to-ROI and blend-back by sending a full-resolution industrial frame directly to a generic text-to-image call; that is outside this release's tested interface.
Training data and leakage boundary
- Each object uses 10 real anomalous images selected with seed 42.
- Selections are published in
splits/fewshot_selection.jsonin the GitHub repository. - The single test partition is VisA
2cls_highshottest. - No frozen test image, mask, or embedding is read during adapter training or generation.
- Every test image/mask SHA-256 is published in
splits/test_blocklist.json.
Limitations
- Ten training images per object create a high risk of overfitting and limited defect diversity.
- Trigger tokens are pseudo-types and may mix multiple visual failure mechanisms.
- SDXL
pcb1-type0can hallucinate component- or insect-like shapes, while capsulestype0can repeatedly produce jewellery-, button-, lens-, or mechanical-ring-like objects; search improves boundary fit but does not guarantee semantic correctness. - Inpainting quality is sensitive to mask geometry, crop ratio, guidance scale, and background domain.
- These adapters cover only
pcb1andcapsules; they are not general industrial anomaly models. - Outputs remain subject to the base models' Open RAIL++-M use restrictions.
License chain
Source Code 與第三方 Artifact 的完整界線見 THIRD_PARTY_NOTICES.md。
| 資產 | License | DefectForge 義務 |
|---|---|---|
| VisA 原始 Dataset | CC BY 4.0 | 標示 VisA 與其論文;Hugging Face Dataset 不得包含原始影像 |
sd2-community/stable-diffusion-2-inpainting |
CreativeML Open RAIL++-M | 保留用途限制,並揭露 preservation mirror |
diffusers/stable-diffusion-xl-1.0-inpainting-0.1 |
CreativeML Open RAIL++-M | 保留用途限制 |
facebook/dinov2-base |
Apache-2.0 | 標示模型與 DINOv2 論文 |
| DefectForge Synthetic Images | CC BY 4.0 | 視為 VisA 衍生內容;保留 VisA attribution,並揭露 Diffusion base model License |
| DefectForge LoRA Weights | CreativeML Open RAIL++-M | 繼承對應 base model 的限制,並附上 License 連結 |
| DefectForge Source Code | MIT | MIT 僅授權程式碼,不包含 Dataset 與 Model Weights |
Citation and provenance
Methodology, checksums, training reports, and limitations are published in kuotunyu/defectforge-visa-synthetic-data. GitHub Citation Metadata 見 CITATION.cff. The project is an independent open-source replication and is not affiliated with or endorsed by NVIDIA.
- Downloads last month
- -
Model tree for steven0226/defectforge-visa-lora
Base model
stabilityai/stable-diffusion-xl-base-1.0