DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion

Official model weights for the paper DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion.

Paper (coming soon) | Project Page | Code

Despite recent progress, direct text-driven 4D object generation remains challenging yet highly desirable. In this paper, we introduce DiTex4D, a native text-to-4D generation framework that enables both text-driven 4D generation from scratch and 3D animation from static mesh. Built upon large-scale pre-trained 3D generation models, our framework DiTex4D avoids intermediate text-to-video pipelines and costly per-object optimization. Specifically, (i) we achieve 4D spatiotemporal consistency via inflating 3D attention with mixed-4D RoPE and a tailored correlated noise injection strategy. (ii) To enable 3D animation, we introduce a mask-based diffusion model conditioned on multi-view global context to maintain strict consistency with the initial frame. We further fine-tune the framework for 4D interpolation to synthesize high-frame-rate sequences with smoother motion. Extensive experiments demonstrate that DiTex4D can achieve higher-quality, semantically aligned, and spatiotemporally coherent 4D object generation, surpassing most existing state-of-the-art text-to-4D generation methods.

Citation

@article{chen2026ditex4d,
  author  = {Chen, Xiaozhe and Rong, Mengqi and Liu, Jian and Shen, Shuhan},
  title   = {DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion},
  journal = {ECCV},
  year    = {2026},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sugercxz/DiTex4D-text-to-4d

Finetuned
(1)
this model