Image-to-Text
Transformers
Safetensors
qwen2_5_vl
image-text-to-text
qwen2.5-vl
web-design
html-generation
lora
designforge
text-generation-inference
Instructions to use ivan123-123/preparing.ai with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ivan123-123/preparing.ai with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("image-to-text", model="ivan123-123/preparing.ai")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ivan123-123/preparing.ai") model = AutoModelForMultimodalLM.from_pretrained("ivan123-123/preparing.ai", device_map="auto") - Notebooks
- Google Colab
- Kaggle
preparing.ai - DesignForge Qwen2.5-VL 3B (screenshot + brief -> HTML)
Fine-tune of Qwen/Qwen2.5-VL-3B-Instruct that turns a webpage screenshot plus a short design
brief into a complete self-contained HTML/CSS/JS document.
- Full merged model (fp16):
ivan123-123/preparing.ai - LoRA adapter only:
ivan123-123/preparing.ai-lora
Usage
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
"ivan123-123/preparing.ai", torch_dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained("ivan123-123/preparing.ai")
SYSTEM = ("You are DesignForge, an expert front-end engineer and UI designer. Given a webpage "
"screenshot and a design brief, you write one complete, polished, responsive, "
"self-contained HTML document (HTML, CSS and JavaScript as needed). Output only the HTML.")
msgs = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": [
{"type": "image", "image": image},
{"type": "text", "text": "Design brief:\n" + brief + "\n\nGenerate the complete webpage."},
]},
]
prompt = processor.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=2500, do_sample=True,
temperature=0.6, top_p=0.9, top_k=50)
html = processor.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0]
open("page.html", "w", encoding="utf-8").write(html)
Inference settings
- Use sampling (
temperature=0.6, top_p=0.9, top_k=50). Greedy decoding on fp16 can degenerate to repeated characters on some inputs. - The processor ships with the training pixel budget:
min_pixels=100352, max_pixels=301056(≈1008x784). Keep these so images are framed as they were during training. - Expected output length: ~1.7k-2.6k tokens (one HTML document, no markdown fences).
Training / evaluation
- 200 optimizer steps, grad-accum 4, LR 2e-4 cosine, LoRA r=16 on a P100, bf16/4-bit NF4.
- 504 hand-written brief + screenshot pairs (WebSight-style), 500 train / 4 holdout.
- Held-out: 3/3 structurally valid and rendered HTML, mean SSIM 0.67 vs the reference screenshot.
- Artifact test (4-bit + LoRA, sampling): PASS.
- Downloads last month
- -
Model tree for ivan123-123/preparing.ai
Base model
Qwen/Qwen2.5-VL-3B-Instruct