3D VAE Step2
图像与视频混合训练的左目生成右目,训练至 14,000 steps 的原始 full checkpoint。
代码与完整准备流程:WangHaiyan378/3D_VAE,发布时对应提交 77f53f5。
- 权重文件:
step_14000.pt,17,030,129,902 bytes。 - Wan2.1-T2V-1.3B DiT,全量微调;32 通道输入由噪声 latent(16)与左目 latent(16)拼接。
- checkpoint 包含
model、optim、step、finetune、cfg,保留完整优化器状态。 - 配套
stage2_mix.yaml是代码仓库的可迁移配置模板;原始训练配置保存在 checkpoint 的cfg中,包含原训练环境路径,使用时请按本机路径调整。 - 本仓库只提供该阶段 checkpoint;官方 Wan/VAE/T5 权重和文本缓存需要按代码仓库说明准备。
- 文件是项目自定义 PyTorch checkpoint,需要先安装并在代码仓库内使用;不是 Diffusers 或 Transformers 的
from_pretrained格式。原始配置包含项目 Config 对象,项目加载器使用weights_only=False。
下载及使用
在代码仓库根目录、安装依赖后执行:
from huggingface_hub import hf_hub_download
from pathlib import Path
import shutil
downloaded = hf_hub_download('sdfwasdf/3D_VAE_Step2_model', 'step_14000.pt')
target = Path('output/mix_1.3b_full/checkpoints/latest.pt')
target.parent.mkdir(parents=True, exist_ok=True)
shutil.copyfile(downloaded, target)
按代码仓库 README 准备官方权重与文本缓存,再运行:
python -m video.stereo_mix.infer.infer_video --config configs/stage2_mix.yaml --ckpt output/mix_1.3b_full/checkpoints/latest.pt --input left.mp4 --output right.mp4 --num-frames 81 --steps 30
本阶段从 Step1 的 step_4000.pt 初始化。可直接作为代码仓库 Stage3 深度条件训练的 model.init_from。
文件字节数与 SHA-256 见 checkpoint_info.json 和 SHA256SUMS。
Model tree for sdfwasdf/3D_VAE_Step2_model
Base model
Wan-AI/Wan2.1-T2V-1.3B