Instructions to use microsoft/radedit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use microsoft/radedit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("microsoft/radedit", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Fix keep_mask handling in inversion_reverse_process_two_masks
Thanks for releasing RadEdit. While using it to generate counterfactual chest X-rays for auditing a medical vision-language model (https://github.com/m871-akram/cxr-grounding), I found two problems in how pipeline.py (main, e8ebd31) handles keep_mask in inversion_reverse_process_two_masks. Both touch the same lines, so they are fixed together here.
1. The kept region is pasted back one step too noisy.
After reverse_step, xt is at the previous timestep, but the code pastes back timestep_to_latents[:, idx + 1], the inverted latent of the current one. The kept region therefore ends as the noisy latent of the smallest timestep instead of x0 (more noise with fewer steps). The + 1 seems to avoid a NaN: at the last inversion step the variance is 0, and the error-correction line writes NaN into timestep_to_latents[:, 0]. The fix skips that correction when the variance is 0 (so index 0 holds x0) and pastes back [:, idx].
2. A keep mask is overridden.
A second if keep_mask is not None: block pastes the inverted latent back everywhere outside the edit mask. So any keep mask, even an empty one, gives the same image as keep = 1 - edit, and an empty keep mask differs from passing none. Algorithm 3 of the paper only restores the region inside m_keep and lets the area outside m_edit change through unconditional generation, which the loop already does. The fix removes that block.
Reproduction: script, public NIH ChestX-ray14 film 00000001_000.png, seed 0, skip ratio 0.3, CPU is enough. Mean absolute gray-level difference (0-255) more than 32 px outside a box edit mask:
| Check | Expected | Current | This PR |
|---|---|---|---|
| keep = 1 - edit vs the film's VAE round trip (4 / 8 / 16 steps) | ~0 | 38.0 / 16.6 / 8.0 | 0.81 / 0.36 / 0.11 |
| keep = 1 - edit vs empty keep mask | > 0 | 0.000 | 5.42 |
| empty keep mask vs no keep mask | 0 | 8.78 | 0.000 |
| no keep mask: current vs this PR (whole image) | 0 | 0.000 (larges |
Behaviour change: calls without keep_mask give identical outpusk, the region outside both masks can now change, as in the paper; tokeep everything outside the edit mask, pass keep = 1 - edit, which is now restored exactly.
Versions: diffusers 0.40.0, transformers 5.17.0, torch 2.8.0. Happy to split this into two PRs or adjust anything.