Jogg-Avatar 14B

Jogg-Avatar 14B is an audio-driven 720p avatar video generation model based on Wan2.1-T2V-14B. It adds audio conditioning and LoRA adapters to the Wan video diffusion model.

Source code and complete inference instructions: chanjing-ai/Jogg-Avatar

Checkpoint

Directory Base model Parameters stored Weight dtype
Jogg-Avatar-14B/ Wan2.1-T2V-14B Audio modules, input projection, and LoRA adapters BF16

The Wan2.1 base model and Wav2Vec audio encoder are not duplicated here. Download them separately:

Repository Layout

.
├── Jogg-Avatar-14B/
│   ├── config.json
│   └── diffusion_pytorch_model.safetensors
├── LICENSE
├── README.md
└── SHA256SUMS

Download

mkdir -p models
hf download cicada-ai/jogg-avatar \
  --include "Jogg-Avatar-14B/*" \
  --local-dir models

Then follow the inference guide.

中文说明

本仓库只提供基于 Wan2.1-T2V-14B 的 Jogg-Avatar 14B 音频驱动权重。 权重包含音频条件模块、输入投影和 LoRA 参数,不重复分发 Wan2.1 基座模型与 Wav2Vec。完整环境、模型目录和推理命令请参考 开源代码仓库

License

Released under the Apache License 2.0. The license and usage terms of the Wan2.1 base model also apply.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cicada-ai/jogg-avatar

Finetuned
(76)
this model