Prism Alpha for WanGP
INT8 ConvRot weights for the Prism Alpha preview, with the matching UMT5 conditioning encoder, original-precision BF16 video/audio codecs, tokenizer assets and a paired LightX2V accelerator.
Prism takes a starting image and a prompt and generates video with a 48 kHz mono soundtrack. It requires the Prism integration in WanGP, including its shared INT8 ConvRot loader. These quantized files are not drop-in Diffusers checkpoints.
Files
| File or directory | Contents |
|---|---|
Prism_Alpha_int8_convrot.safetensors |
Single-file Alpha transformer; both video experts, audio transformer and bridges |
UMT5-XXL-Transformers/ |
INT8 ConvRot conditioning encoder, config and tokenizer |
prism/video_vae_bf16.safetensors |
Native BF16 video codec |
prism/audio_vae_bf16.safetensors |
Native BF16 audio codec |
loras_accelerators/ |
Paired HIGH/LOW 260412 rank-256 LightX2V I2V adapters, unchanged from the pinned source |
profiles/prism/ |
WanGP accelerator profile |
prism/provenance.json |
Pinned upstream sources and preparation provenance |
SHA256SUMS.json |
Local hashes and sizes for release verification |
The BF16 main transformer was verified as a single file. The three UMT5 shards were merged without changing tensor bytes or dtypes before quantization. Main transformer and encoder BF16 intermediates are not included in this release. The 1,340 transformer Linear weights and 168 encoder Linear weights are quantized; remaining tensors retain their original precision. In particular, all 24 UMT5 relative-position embedding tables retain their original BF16 bytes.
For manual installation, place the transformer in the WanGP checkpoint root, retain the UMT5-XXL-Transformers/ and prism/ folder hierarchy under that root, and place the two adapters in the configured Prism LoRA folder. Select INT8 ConvRot for both transformer and text encoder. A WanGP build with Prism integration downloads these assets through its normal model and accelerator selection.
LightX2V accelerator
Select LightX2V 260412 Rank256 - 8 Steps in WanGP's accelerator profiles. It loads both adapters with multipliers 1;0 0;1, selects Full Video Attention, sets 8 video steps, flow shift 5 and video guidance 2 on the first step (1 thereafter). Audio takes four Euler updates per video step with guidance 5.
This adapts the FreeVideo Light recipe to WanGP shared attention and a BF16 conditioning cache. It does not claim numerical, quality or speed parity with FreeVideo. The Wan2.2 adapter pair is mapped to Prism's video experts; other adapters have not been validated.
To use the base model again, clear the adapters and select Euler with 50 steps, guidance 5 and flow shift 7. The upstream reference is 1280x720, 205 frames and 24 fps. Full Video Attention is recommended. Sparse remains experimental and damaged late frames in a tested 480p sample.
Validation and limitations
The actual WanGP headless loader and generator were exercised with these checkpoints and both adapters. An 848x480, 49-frame, 8-step render using SageAttention 2, head split 2 and MMGP profile 4 completed in 2m07s including loading and saving on the test machine, with 4.37 GiB peak allocated VRAM and 5.62 GiB reserved. This is one measured run, not a general throughput promise. Its 49 decoded frames were visually coherent, and audio was finite and unclipped. The audio was quiet; speech intelligibility and lip synchronization have not been established.
Cached conditional audio matched the corresponding joint pass exactly. Cancellation followed by another generation was reproducible, step caches were released, and unloading adapters restored native transformer output exactly. Full-length 720p visual quality has not been validated here. Quality depends on the image, prompt and seed; smaller resolutions and shorter clips are experimental for this preview.
Sources and licensing
- Tencent-Hunyuan/Prism, revision
8f0c2633f40ccf89e26f928c6d1771616caeb510: Prism MIT license with inherited MOVA Apache-2.0 terms; seeLICENSEandNOTICE. - FrancisRing/Prism: original Alpha and supporting weights.
- Kijai/WanVideo_comfy: source of the unmodified LightX2V adapters. Underlying Wan2.2 and LightX2V components retain their respective upstream Apache-2.0 terms.
- FlashML-org/FreeVideo: accelerated sampling recipe, Apache-2.0; see
licenses/FreeVideo-LICENSE.
The preparation and verification scripts in tools/ are intended to run in the matching WanGP checkout. Transformer quantization is exposed by the handler through --save-quantized --convrot --test. Encoder conversion uses the same shared converter in a temporary post-load hook during a headless test queue; embeddings must remain unquantized. Source hashes and conversion checks are retained in prism/.
Model tree for DeepBeepMeep/Prism
Base model
FrancisRing/Prism