Paint-on-glass LoRA

Paint-on-glass LoRA is a temporal-style LoRA for MiniMax H3 Ref2VA. Its target is the moving, continuously repainted behavior of hand-painted animation rather than only a static oil-paint appearance.

Main comparison

The main comparison reel provides a side-by-side overview of the base MiniMax H3 output and the Paint-on-glass LoRA across representative subjects, motion types, and camera movements.

Open the main comparison video

Individual comparisons

Each clip below can be opened and played independently. These qualitative tests cover subject motion, camera tracking, environmental motion, water, wind, and temporal painted variation. The original prompts were not retained, so this repository does not claim reconstructed prompts for these outputs.

01 β€” Goldfish 02 β€” White fish

Open video

Open video
03 β€” Candle 04 β€” Countryside walk

Open video

Open video
05 β€” Handheld selfie 06 β€” Running puppy

Open video

Open video
07 β€” Snow 08 β€” Water

Open video

Open video
09 β€” Grassland run

Open video

Trigger phrase

pboilx

Use the trigger phrase at the beginning of the inference prompt. When using Ref2VA, explicitly refer to the uploaded reference image as <Picture 1> in the prompt. If additional reference images are supplied, refer to them with their corresponding identifiers, such as <Picture 2>.

Recommended prompt structure:

pboilx. Use <Picture 1> as the visual reference for the composition, subject, and visual style. [Describe the desired subject movement, camera work, and environmental movement in detail.]

The bracketed section is not a fixed prompt. Replace it with instructions for the intended motion. Do not copy the comparison labels above as prompts; the exact prompts used for those videos were not retained.

Confirmed training record

Field Archived value
Base model family MiniMax H3
Training type Ref2VA
Trainer fal minimax/h3/ref2va/trainer
LoRA rank 16
Training steps 1500
Learning rate 0.0002
Frames per training sample 73
Training frame rate 24 fps
Resolution preset Medium
Aspect ratio 16:9
Reference conditioning probability 0.9
Training samples 59
Trigger phrase pboilx
Audio used for training No
Reference strategy One automatic middle-frame self-reference per clip

Only fields supported by the final fal training request, archived weight filename, trainer output, debug manifest, preprocessing report, or retained experiment record are listed. Optimizer, learning-rate scheduler, random seed, and the exact pixel dimensions represented by the medium resolution preset were not exposed by the retained request.

Repository contents

weights/
  MiniMax_H3_Paint-on-glass_rank16_step1500.safetensors
config/
  trainer_output_config.json
  training_config.json
  inference_config.json
workflows/
  H3lora_minimax_r2v.json
preprocessing/
  README.md
  debug_manifest.json
  preprocessing_report.json
evaluation/
  qualitative_findings.md
  comparison_videos/
    Main_H3LoRAcomparison.mp4
    01_GlodFish.mp4
    02_WhiteFish.mp4
    03_candle.mp4
    04_walk.mp4
    05_GoldHair.mp4
    06_dog.mp4
    07_snow.mp4
    08_water.mp4
    09_run.mp4
checksums/
  SHA256SUMS

Inference notes

workflows/H3lora_minimax_r2v.json is the actual ComfyUI workflow used for one validation run. It uses MiniMax H3 Ref2VA, the LoRA at model strength 1.0, res_multistep, 20 sampling steps, a 6-second duration input, 24 fps, and a 16:9 0.4-megapixel output (resolved by the workflow to 864 x 480). The duration expression resolves this run to 158 frames.

The workflow preserves the original local LoRA filename MiniMax_H3_pboilx_temporal_rank16_step1500.safetensors, while the same archived artifact is named MiniMax_H3_Paint-on-glass_rank16_step1500.safetensors. Rename the archived file to the workflow filename after placing it in ComfyUI, or select the archived filename in the LoraLoaderModelOnly node. The node's custom title still says β€œStep 100,” but its selected file is the step-1500 LoRA.

The model was evaluated manually on varied reference images including people, animals, landscapes, camera movement, environmental motion, and abstract animation. Increasing LoRA strength generally increased the visible painted temporal variation. At high strength, subject motion could become more rigid and generated audio could be reduced or lost.

The videos above are qualitative comparisons. A complete seed-to-output evaluation index and quantitative metric table are not currently included.

Dataset note

The training media and original caption files are not included in this model archive. debug_manifest.json records the 59 processed samples seen by the trainer. Its captions include the trigger phrase as recorded after trainer preprocessing, and its decoded-video and conditioning paths refer to temporary trainer outputs that are not bundled here.

License

No redistribution license has been assigned yet. Confirm the base-model terms and the rights status of the training sources before making this repository or its weights public.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support