Instructions to use WepeNerd/Obscura_Remova with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use WepeNerd/Obscura_Remova with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Lightricks/LTX-2.3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("WepeNerd/Obscura_Remova") prompt = "Remove the window curtains from the foreground." input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png") image = pipe(image=input_image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - LTX-2
How to use WepeNerd/Obscura_Remova with LTX-2:
# Install the LTX-2 pipelines git clone https://github.com/Lightricks/LTX-2.git cd LTX-2 uv sync --frozen
# Download the weights from this repo, plus the Gemma text encoder hf download WepeNerd/Obscura_Remova --local-dir models/Obscura_Remova hf download google/gemma-3-12b-it-qat-q4_0-unquantized --local-dir models/gemma-3-12b
# Text/image-to-video with the LoRA on the HQ two-stage base pipeline uv run python -m ltx_pipelines.ti2vid_two_stages_hq \ --checkpoint-path path/to/checkpoint.safetensors \ --distilled-lora path/to/distilled_lora.safetensors 0.8 \ --spatial-upsampler-path path/to/spatial_upsampler.safetensors \ --gemma-root models/gemma-3-12b \ --lora models/Obscura_Remova/<weights>.safetensors 1.0 \ --prompt "your prompt here" \ --output-path output.mp4 # For image-to-video, add: --image path/to/image.jpg 0 0.8 - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
breakdown of video guide token at specific resolution ...
Trying to remove an actress to separate the background in a 2k plate.
Test runs in low resolution test workflow just fine with a mask input fed into the latent to zero the guide strength within the mask, otherwise Obscura still did actually put out a human in every run. The masked workflow works up towards 1344x704, any higher resolution will lead to a freshly generated video that just follows the text prompt, no apparent elements are still pulled from the reference video.
Do not know if this is a bug or to be expected within LTX2.3, just wanted to let you know.
Revisiting this, it doesn't seem to be a resolution only gate, but the point before Obscura breaks and plain video generation from text prompt starts is dependent on resolution and frame count.
So, the 1344x704 with 41frames will work, equally 768x448 with 113frames, on my machine, running LTX2.3 dev model MLX within ComfyUI an a MacBookPro M5 Max with 128GB unified memory.
After extensive testing, Obscura Remova IC-LoRA seems to break down and stop working above around 5200 latent video tokens for my use case with multiple shots and an actress in the foreground of a handheld camera shot utilizing a guide mask. So, since doing ObscuraRemova removals in chunks usually breaks temporal stability, this leads to needing to lower the resolution to get usable results.
latent_frames = (frames - 1) / 8 + 1
tokens = (width / 32) × (height / 32) × latent_frames
For 113 frames, this resulted in a 768x448 window, for another shot with 145 frames, it was 704x384 producing stable Obscura Remova Outputs.