Instructions to use joyfox/JoyFox-LiveTalk-DH-1.3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use joyfox/JoyFox-LiveTalk-DH-1.3B with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("joyfox/JoyFox-LiveTalk-DH-1.3B", torch_dtype=torch.bfloat16, device_map="cuda") pipe.to("cuda") prompt = "A man with short gray hair plays a red electric guitar." image = load_image( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png" ) output = pipe(image=image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Notebooks
- Google Colab
- Kaggle
JoyFox LiveTalk-DH 1.3B
Any Image Β· Any Voice Β· Bring Every Persona to Life
JoyFox LiveTalk-DH 1.3B is a custom-trained model for high-quality, audio-driven human video generation. It was retrained from the official SoulX-FlashHead 1.3B Lite base model, with a focus on high-resolution portrait detail, identity consistency, lip synchronization, and long-sequence stability.
Provide one portrait image and one driving audio track to generate a human video that preserves the original composition, identity, and background. Inference uses the complete reference image without center cropping or adding white padding.
Visual Comparisons
Each comparison uses the same reference image, driving audio, and inference
settings. The input retains its original aspect ratio and is scaled only to a
VAE-compatible aligned resolution. No cropping, padding, or recomposition is
applied. The left side shows SoulX-FlashHead-1_3B Lite; the right side shows
JoyFox-LiveTalk-DH.
Natural Portrait
| Reference Image | Side-by-Side Comparison |
|---|---|
|
Studio Portrait
| Reference Image | Side-by-Side Comparison |
|---|---|
|
Model Highlights
| Capability | Design | Benefit |
|---|---|---|
| π Identity Preservation | Spatial VAE reference conditioning | Preserves facial structure, silhouette, clothing colors, and background layout |
| πΌοΈ Original Composition | Multi-aspect-ratio buckets and automatic size matching | Uses the full reference image without cropping or white padding |
| π Audio Driven | Multi-layer wav2vec2 features and audio cross-attention | Improves synchronization between lip motion, speech rhythm, and facial movement |
| π¬ Continuous Generation | 97-frame training and streaming segment continuity | Improves long-sequence motion and transitions between generated segments |
| β¨ High-Definition Detail | Retrained at a 960-pixel baseline | Enhances portrait texture and fine detail at high resolution |
| β‘ Efficient Deployment | 1.3B Transformer with BF16 weights | Balances generation quality with inference cost |
| π Simplified Dependencies | Official LTX Video VAE | Requires neither a text encoder nor a CLIP image encoder |
JoyFox LiveTalk-DH
JoyFox LiveTalk-DH is the high-resolution, multi-aspect-ratio training and inference method used by this release. It does not rely on an additional text encoder or CLIP image encoder. Instead, the reference image, audio, and video timeline are jointly modeled within one generation network.
Multi-Aspect-Ratio Bucketing
Training samples are not stretched into a fixed 512Γ512 square. The system selects the closest resolution bucket from the source video aspect ratio and aligns the spatial dimensions to the compression requirements of the LTX VAE. The 960-pixel baseline provides more effective pixels while allowing landscape, portrait, and square content to retain more natural geometry.
Composition Preservation
At inference time, the original image dimensions are read and the output aspect
ratio is selected automatically. For non-standard inputs, direct mode
prioritizes the source aspect ratio and selects the nearest dimensions aligned
to a multiple of 32. The full image is scaled proportionally without center
cropping, padding, or synthetic background extension.
Spatial VAE Reference Conditioning
The official LTX Video VAE encodes the reference image into a 128-channel video latent space. The reference latent acts as first-frame motion context and is expanded into video conditioning that enters the Transformer together with the noise latents. Compared with a single image embedding, this spatial condition retains facial structure, subject contours, colors, and background layout.
Audio Cross-Attention
wav2vec2 extracts multi-layer temporal features from the driving speech. A local audio window is constructed for each video frame and converted into context tokens by the audio projection network. Cross-attention in the Transformer layers aligns lip motion with the current, preceding, and following speech segments.
High-Resolution Long-Sequence Training
The model was retrained with a 960-pixel baseline and 97-frame clips. Training jointly optimizes spatial detail, inter-frame motion, and audio synchronization, enabling more stable long sequences while preserving identity. The model is stored in BF16 to reduce inference memory and storage requirements while maintaining generation quality.
Streaming Segment Generation
Long-audio inference uses segmented generation with historical motion-frame continuity. Each new segment reuses preceding motion context while a sliding window continuously updates the audio features. This supports videos longer than a single training clip and reduces motion discontinuities at segment boundaries.
Scope and Limitations
The current training distribution is dominated by close-up portraits, so close-up and bust compositions are the most reliable. Automatic aspect-ratio selection and composition preservation reduce geometric distortion for non-square inputs, but results for half-body, full-body, multi-person, and complex-background scenes still depend on the coverage of those cases in the training data.
Inference Example
The following standalone example initializes the model components and generates an audio-driven video:
import math
from pathlib import Path
import torch
from diffusers import AutoencoderKLLTXVideo, FlowMatchEulerDiscreteScheduler
from huggingface_hub import snapshot_download
from PIL import Image
from videox_fun.models import FlashHeadAudioEncoder, FlashHeadTransformer3DModel
from videox_fun.pipeline import FlashHeadPipeline
from videox_fun.utils.utils import (
get_image_latent,
merge_video_audio,
save_videos_grid,
)
# Model repository, inputs, and output
model_dir = Path(
snapshot_download(repo_id="joyfox/JoyFox-LiveTalk-DH-1.3B")
)
reference_image = model_dir / "assets/reference-natural-portrait.jpg"
driving_audio = model_dir / "assets/driving-audio.wav"
output_video = Path("joyfox_livetalk_dh.mp4")
# wav2vec2 is a separate official audio encoder downloaded on first use.
audio_encoder_dir = snapshot_download(
repo_id="facebook/wav2vec2-base-960h",
allow_patterns=[
"config.json",
"preprocessor_config.json",
"model.safetensors",
],
)
# Find the nearest 32-aligned size while preserving the source aspect ratio.
with Image.open(reference_image) as image:
source_width, source_height = image.size
source_ratio = source_width / source_height
target_pixels = 960**2
ideal_width = (target_pixels * source_ratio) ** 0.5
ideal_height = (target_pixels / source_ratio) ** 0.5
width_center = max(1, round(ideal_width / 32))
height_center = max(1, round(ideal_height / 32))
candidates = []
for width_step in range(max(1, width_center - 4), width_center + 5):
for height_step in range(max(1, height_center - 4), height_center + 5):
width = width_step * 32
height = height_step * 32
aspect_error = abs(math.log((width / height) / source_ratio))
area_error = abs(math.log((width * height) / target_pixels))
candidates.append((aspect_error + 0.02 * area_error, height, width))
_, height, width = min(candidates)
dtype = torch.bfloat16
device = "cuda"
transformer = FlashHeadTransformer3DModel.from_pretrained(
model_dir / "Model_Lite",
transformer_additional_kwargs={
"transformer_subpath": "./",
"dict_mapping": {
"in_dim": "in_channels",
"dim": "hidden_size",
},
},
low_cpu_mem_usage=True,
torch_dtype=dtype,
)
vae = AutoencoderKLLTXVideo.from_pretrained(
model_dir / "VAE_LTX",
low_cpu_mem_usage=True,
).to(dtype)
audio_encoder = FlashHeadAudioEncoder(str(audio_encoder_dir), "cpu")
scheduler = FlowMatchEulerDiscreteScheduler(
num_train_timesteps=1000,
shift=5.0,
use_dynamic_shifting=False,
base_shift=0.5,
max_shift=1.15,
base_image_seq_len=256,
max_image_seq_len=4096,
)
pipeline = FlashHeadPipeline(
transformer=transformer,
vae=vae,
scheduler=scheduler,
audio_encoder=audio_encoder,
).to(device)
ref_image = get_image_latent(
str(reference_image),
sample_size=[height, width],
resize_mode="direct",
)
generator = torch.Generator(device=device).manual_seed(42)
with torch.inference_mode():
video = pipeline(
ref_image=ref_image,
audio_path=str(driving_audio),
height=height,
width=width,
segment_frame_length=33,
num_inference_steps=4,
audio_guide_scale=1.0,
audio_encode_mode="stream",
shift=5.0,
fps=25,
max_frames_num=500,
stochastic_sampling=True,
generator=generator,
).videos
output_video.parent.mkdir(parents=True, exist_ok=True)
save_videos_grid(video, str(output_video), fps=25)
merge_video_audio(
video_path=str(output_video),
audio_path=str(driving_audio),
)
direct mode uses the complete reference image. Together with automatic size
selection, it preserves the source aspect ratio while aligning the output
dimensions to multiples of 32. It does not introduce cropping, white borders,
or blurred background fill.
Source and License
This model was retrained from the official Lite release of Soul-AILab/SoulX-FlashHead-1_3B and uses the Apache-2.0 license. Use and redistribution must also comply with the original model license and the licenses governing the training data.
- Downloads last month
- -
Model tree for joyfox/JoyFox-LiveTalk-DH-1.3B
Base model
Soul-AILab/SoulX-FlashHead-1_3B