preparing.ai - DesignForge Qwen2.5-VL 3B (screenshot + brief -> HTML)

Fine-tune of Qwen/Qwen2.5-VL-3B-Instruct that turns a webpage screenshot plus a short design brief into a complete self-contained HTML/CSS/JS document.

  • Full merged model (fp16): ivan123-123/preparing.ai
  • LoRA adapter only: ivan123-123/preparing.ai-lora

Usage

import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "ivan123-123/preparing.ai", torch_dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained("ivan123-123/preparing.ai")

SYSTEM = ("You are DesignForge, an expert front-end engineer and UI designer. Given a webpage "
          "screenshot and a design brief, you write one complete, polished, responsive, "
          "self-contained HTML document (HTML, CSS and JavaScript as needed). Output only the HTML.")

msgs = [
    {"role": "system", "content": SYSTEM},
    {"role": "user", "content": [
        {"type": "image", "image": image},
        {"type": "text", "text": "Design brief:\n" + brief + "\n\nGenerate the complete webpage."},
    ]},
]
prompt = processor.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=2500, do_sample=True,
                     temperature=0.6, top_p=0.9, top_k=50)
html = processor.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0]
open("page.html", "w", encoding="utf-8").write(html)

Inference settings

  • Use sampling (temperature=0.6, top_p=0.9, top_k=50). Greedy decoding on fp16 can degenerate to repeated characters on some inputs.
  • The processor ships with the training pixel budget: min_pixels=100352, max_pixels=301056 (≈1008x784). Keep these so images are framed as they were during training.
  • Expected output length: ~1.7k-2.6k tokens (one HTML document, no markdown fences).

Training / evaluation

  • 200 optimizer steps, grad-accum 4, LR 2e-4 cosine, LoRA r=16 on a P100, bf16/4-bit NF4.
  • 504 hand-written brief + screenshot pairs (WebSight-style), 500 train / 4 holdout.
  • Held-out: 3/3 structurally valid and rendered HTML, mean SSIM 0.67 vs the reference screenshot.
  • Artifact test (4-bit + LoRA, sampling): PASS.
Downloads last month
-
Safetensors
Model size
4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ivan123-123/preparing.ai

Adapter
(301)
this model