🎤 LTX-2.3 22B Distilled — Audio-to-Video / Lip Sync

📌 Overview / Tổng quan

This repository contains the backup weights for the LTX-2.3 22B Distilled model, optimized for Audio-to-Video (Lip Sync) tasks. All symlinks have been resolved to ensure safe and easy downloads.

🚀 Specifications / Đặc điểm kỹ thuật

  • Architecture: LTX-Video (Distilled 8-steps)
  • Format: GGUF (Q4_K_M) + Safetensors
  • Text Encoder: Gemma 3 12B IT (bfloat16)
  • Hardware Requirement: Recommended for Google Colab with T4 (16GB VRAM) or L4 (22.5GB VRAM) GPU.

📁 Directory Structure / Cấu trúc thư mục

  • ltx-2.3-22b-distilled-Q4_K_M_light.gguf: Main Transformer weights.
  • gemma-3-12b-it-qat-q4_0-unquantized/: Full Text Encoder.
  • *.safetensors: Essential modules (VAE, Vocoder, Spatial Upscaler x2).

Maintained by vinh-gogo. Follow for more AI tutorials at Hinn Channel.

Downloads last month
1,131
GGUF
Model size
19B params
Architecture
ltxv
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support