Instructions to use ij/PixelTune-MiniMax-H3-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
- Notebooks
- Google Colab
- Kaggle
PixelTune · MiniMax H3 pixel-art LoRA
Experimental 1000-step checkpoint. This is the current PixelTune v3 adapter used in our tested custom H3 inference workflow. It is not a completed 5000-step run and has not passed final aesthetic or motion-quality approval.
현재 웹 생성에 사용 중인 1000스텝 LoRA입니다. 이후 준비한 클립별 캡션 수정, 실제 gradient 비중 재보정, 정적 픽셀 시간 손실 및 사용자 선별 데이터로의 재학습은 이 체크포인트에 포함되지 않습니다.
Download and load
Download pixeltune_h3_full_canvas4_v3_step_01000.safetensors (FP32 adapter weights, 155,058,176 parameters).
The base model and its Qwen encoder/video and audio VAEs are separate downloads.
Tested base: Comfy-Org/MiniMax-H3, revision
e5eb578a89295337b8ff433a035929ce0279e0b6, diffusion_models/minimax_h3_fl2va_pruned_bf16.safetensors.
Original encoder/VAE revision: MiniMaxAI/MiniMax-H3 at
42ed227ee7df40d41602854ae760620d6eb651fe.
The file contains 416 canonical *.lora_A.weight / *.lora_B.weight tensors,
rank 32, alpha 32, with Comfy QKV layout. It is not a merged model or a
standard PEFT save_pretrained directory. Automatic ComfyUI loading has not
been verified. Do not assume compatibility with a differently ordered QKV base.
Our real workflow uses DiffSynth-Studio commit 974cfa37f27ac55eba3b6d10efa21f876900572d
and an already loaded MiniMaxH3Pipeline named pipe:
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
path = hf_hub_download("ij/PixelTune-MiniMax-H3-LoRA", "pixeltune_h3_full_canvas4_v3_step_01000.safetensors")
state = {name: tensor.to(device=pipe.device, dtype=pipe.torch_dtype)
for name, tensor in load_file(path, device="cpu").items()}
pipe.load_lora(pipe.dit, state_dict=state, alpha=1.0)
This is the adapter-loading step, not a standalone inference program. Start from the upstream H3 pipeline with the matching Comfy-layout weights. The tested custom workflow uses 50 inference steps, CFG 1, video shift 12, audio shift 3 and native unchanged RoPE. Strength 1.0 is the current default; 0.5 can be compared, but no optimal strength has been established.
Prompt and resolution conditions
Training used a 4× nearest-neighbor pixel grid. Declare the native canvas and generate at four times those dimensions. For example, 128×128 native pixels correspond to a 512×512 generated image. Retained native sides are multiples of 8. Training output was capped at 1280 pixels per side and 1,048,576 pixels in area. Composition was preserved with minimal border trimming (up to four native edge pixels); it was not resized to a detail crop.
integrated_multimodal_description: [Shot 1] pxgrid4 native_res_128x128 source_fps_10. Native pixel canvas: 128 by 128 pixels. Each native pixel is a 4 by 4 solid-color square; nearest-neighbor enlargement to 512 by 512 pixels. Original source sampling rate: 10 fps. The full scene composition is preserved. 2D pixel-art animation. A small fox sits in a moonlit forest, gently blinking and moving its tail. Fixed camera and a stable background.
overall_soundscape: The clip is silent, with no dialogue, ambient sound or sound effects.
non_diegetic_music: N/A
Use source_fps_8p3333 for a fractional FPS condition such as 25/3.
These are training prompt features, not guarantees of native pixel
alignment, pose cadence or physical motion speed. Our tested pipeline
renders raw video at 24fps and accepts lengths 17n + 5 (22, 39, 56, …).
Longer inference is available in the app, but this checkpoint was trained
on short clips; long-video quality has not been validated. Audio decoding
was omitted in the tested silent-animation workflow.
Training and known limitations
- Run:
pixel_h3_full_canvas4_v3, completed checkpoint step 1000. - 153 active training clips from 51 source videos; 21 validation clips.
- Rank/alpha: 32/32; learning rate: 5e-5; BF16 frozen base, FP32 adapter.
- Loss terms: flow matching plus grid, decoded RGB and within-hold consistency.
- Coefficients: grid 1.256806, RGB 1.662547, hold 2.761293.
- The coefficients above are not measured gradient percentages. Earlier calibration had limited coverage after resolution filtering.
- Some training captions described whole sources rather than the exact short training window. Those issues were identified after this run.
- The later owner-reviewed 23-source selection was not used to train this checkpoint. Static-pixel temporal loss and revised calibration are pending.
- Fine-detail appearance, motion, static-region stability and exact pixel boundaries may be inconsistent. Lower validation loss does not prove better aesthetic generation.
See adapter_manifest.json for exact hashes and
evaluation/ for actual loss graphs and validation records.
The training log contains a resumed interval; the loss summary excludes
superseded records and contains 1000 effective steps.
Actual raw samples — AI-generated
These are unprocessed step-1000 outputs, seed 777, 50 inference steps, 39 frames at 24fps. No grid snapping, palette reduction or pose holding was applied to these previews. MP4 previews use BT.709 YUV encoding. Prompts and measured metrics are included beside the videos.
320×176 native · source 0020
128×128 native · source 0054
144×168 native · source 0049
Pixel postprocessing is separate
Exact pixels and palette limits shown by the app are enforced by a separate postprocessor, not solely by these weights. It extracts actual representative frames using explicit 5/4/4/4 groups, applies Pixel Art Fixer, a shared clip palette without dithering and nearest-neighbor 4× enlargement. Pose FPS is chosen separately and determines output duration: pose count / pose FPS. The raw generator remains 24fps. Exact playback FPS does not establish that the model learned the requested semantic motion speed.
License and attribution
Powered by MiniMax H3. This is a modified model derivative. The original MiniMax H3 Community License Agreement, including its territorial, distribution and use restrictions, applies; see NOTICE. This release does not grant rights beyond that agreement. The original model and encoder/VAEs are not redistributed here.
Training data: trojblue/test-HunyuanVideo-pixelart-videos. Base model: Comfy-Org/MiniMax-H3, originally MiniMaxAI/MiniMax-H3. The dataset's license does not replace the model's license.
- Downloads last month
- -