Prism Alpha for WanGP

INT8 ConvRot weights for the Prism Alpha preview, with the matching UMT5 conditioning encoder, original-precision BF16 video/audio codecs, tokenizer assets and a paired LightX2V accelerator.

Prism takes a starting image and a prompt and generates video with a 48 kHz mono soundtrack. It requires the Prism integration in WanGP, including its shared INT8 ConvRot loader. These quantized files are not drop-in Diffusers checkpoints.

Files

File or directory Contents
Prism_Alpha_int8_convrot.safetensors Single-file Alpha transformer; both video experts, audio transformer and bridges
UMT5-XXL-Transformers/ INT8 ConvRot conditioning encoder, config and tokenizer
prism/video_vae_bf16.safetensors Native BF16 video codec
prism/audio_vae_bf16.safetensors Native BF16 audio codec
loras_accelerators/ Paired HIGH/LOW 260412 rank-256 LightX2V I2V adapters, unchanged from the pinned source
profiles/prism/ WanGP accelerator profile
prism/provenance.json Pinned upstream sources and preparation provenance
SHA256SUMS.json Local hashes and sizes for release verification

The BF16 main transformer was verified as a single file. The three UMT5 shards were merged without changing tensor bytes or dtypes before quantization. Main transformer and encoder BF16 intermediates are not included in this release. The 1,340 transformer Linear weights and 168 encoder Linear weights are quantized; remaining tensors retain their original precision. In particular, all 24 UMT5 relative-position embedding tables retain their original BF16 bytes.

For manual installation, place the transformer in the WanGP checkpoint root, retain the UMT5-XXL-Transformers/ and prism/ folder hierarchy under that root, and place the two adapters in the configured Prism LoRA folder. Select INT8 ConvRot for both transformer and text encoder. A WanGP build with Prism integration downloads these assets through its normal model and accelerator selection.

LightX2V accelerator

Select LightX2V 260412 Rank256 - 8 Steps in WanGP's accelerator profiles. It loads both adapters with multipliers 1;0 0;1, selects Full Video Attention, sets 8 video steps, flow shift 5 and video guidance 2 on the first step (1 thereafter). Audio takes four Euler updates per video step with guidance 5.

This adapts the FreeVideo Light recipe to WanGP shared attention and a BF16 conditioning cache. It does not claim numerical, quality or speed parity with FreeVideo. The Wan2.2 adapter pair is mapped to Prism's video experts; other adapters have not been validated.

To use the base model again, clear the adapters and select Euler with 50 steps, guidance 5 and flow shift 7. The upstream reference is 1280x720, 205 frames and 24 fps. Full Video Attention is recommended. Sparse remains experimental and damaged late frames in a tested 480p sample.

Validation and limitations

The actual WanGP headless loader and generator were exercised with these checkpoints and both adapters. An 848x480, 49-frame, 8-step render using SageAttention 2, head split 2 and MMGP profile 4 completed in 2m07s including loading and saving on the test machine, with 4.37 GiB peak allocated VRAM and 5.62 GiB reserved. This is one measured run, not a general throughput promise. Its 49 decoded frames were visually coherent, and audio was finite and unclipped. The audio was quiet; speech intelligibility and lip synchronization have not been established.

Cached conditional audio matched the corresponding joint pass exactly. Cancellation followed by another generation was reproducible, step caches were released, and unloading adapters restored native transformer output exactly. Full-length 720p visual quality has not been validated here. Quality depends on the image, prompt and seed; smaller resolutions and shorter clips are experimental for this preview.

Sources and licensing

  • Tencent-Hunyuan/Prism, revision 8f0c2633f40ccf89e26f928c6d1771616caeb510: Prism MIT license with inherited MOVA Apache-2.0 terms; see LICENSE and NOTICE.
  • FrancisRing/Prism: original Alpha and supporting weights.
  • Kijai/WanVideo_comfy: source of the unmodified LightX2V adapters. Underlying Wan2.2 and LightX2V components retain their respective upstream Apache-2.0 terms.
  • FlashML-org/FreeVideo: accelerated sampling recipe, Apache-2.0; see licenses/FreeVideo-LICENSE.

The preparation and verification scripts in tools/ are intended to run in the matching WanGP checkout. Transformer quantization is exposed by the handler through --save-quantized --convrot --test. Encoder conversion uses the same shared converter in a temporary post-load hook during a headless test queue; embeddings must remain unquantized. Source hashes and conversion checks are retained in prism/.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DeepBeepMeep/Prism

Finetuned
(4)
this model