declare-lab
/

TangoFlux

Inference Endpoints

Model card Files Files and versions Community

soujanyaporia commited on 11 days ago

Commit

714699f

·

verified ·

1 Parent(s): 3d7b891

Update README.md

Files changed (1) hide show

README.md +59 -1

README.md CHANGED Viewed

@@ -2,4 +2,62 @@
 license: mit
 datasets:
 - cvssp/WavCaps
----

 license: mit
 datasets:
 - cvssp/WavCaps
+---
+<h1 align="center">✨
+<br/>
+TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
+<br/>
+✨✨✨
+</h1>
+<div align="center">
+  <img src="https://github.com/declare-lab/TangoFlux/blob/main/assets/tf_teaser.png" alt="TangoFlux" width="1000" />
+<br/>
+<div style="display: flex; gap: 10px; align-items: center;">
+  <a href="https://openreview.net/attachment?id=tpJPlFTyxd&name=pdf">
+    <img src="https://img.shields.io/badge/Read_the_Paper-blue?link=https%3A%2F%2Fopenreview.net%2Fattachment%3Fid%3DtpJPlFTyxd%26name%3Dpdf" alt="arXiv">
+  </a>
+  <a href="https://huggingface.co/declare-lab/TangoFlux">
+    <img src="https://img.shields.io/badge/TangoFlux-Huggingface-violet?logo=huggingface&link=https%3A%2F%2Fhuggingface.co%2Fdeclare-lab%2FTangoFlux" alt="Static Badge">
+  </a>
+  <a href="https://tangoflux.github.io/">
+    <img src="https://img.shields.io/badge/Demos-declare--lab-brightred?style=flat" alt="Static Badge">
+  </a>
+  <a href="https://huggingface.co/spaces/declare-lab/TangoFlux">
+    <img src="https://img.shields.io/badge/TangoFlux-Huggingface_Space-8A2BE2?logo=huggingface&link=https%3A%2F%2Fhuggingface.co%2Fspaces%2Fdeclare-lab%2FTangoFlux" alt="Static Badge">
+  </a>
+  <a href="https://huggingface.co/datasets/declare-lab/CRPO">
+    <img src="https://img.shields.io/badge/TangoFlux_Dataset-Huggingface-red?logo=huggingface&link=https%3A%2F%2Fhuggingface.co%2Fdatasets%2Fdeclare-lab%2FTangoFlux" alt="Static Badge">
+  </a>
+  <a href="https://github.com/declare-lab/TangoFlux">
+    <img src="https://img.shields.io/badge/Github-brown?logo=github&link=https%3A%2F%2Fgithub.com%2Fdeclare-lab%2FTangoFlux" alt="Static Badge">
+  </a>
+</div>
+</div>
+## Model Overview
+TangoFlux consists of FluxTransformer blocks which are Diffusion Transformer (DiT) and Multimodal Diffusion Transformer (MMDiT), conditioned on textual prompt and duration embedding to generate audio at 44.1kHz up to 30 seconds. TangoFlux learns a rectified flow trajectory from audio latent representation encoded by a variational autoencoder (VAE). The TangoFlux training pipeline consists of three stages: pre-training, fine-tuning, and preference optimization. TangoFlux is aligned via CRPO which iteratively generates new synthetic data and constructs preference pairs to perform preference optimization.
+## Getting Started
+## Citation
+```bibtex
+@article{Hung2025TangoFlux,
+  title = {TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization},
+  author = {Chia-Yu Hung and Navonil Majumder and Zhifeng Kong and Ambuj Mehrish and Rafael Valle and Bryan Catanzaro and Soujanya Poria},
+  year = {2025},
+  url = {https://openreview.net/attachment?id=tpJPlFTyxd&name=pdf},
+  note = {Available at OpenReview}
+}
+```