Instructions to use AbstractPhil/geolip-beatrix-sana with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Sana
How to use AbstractPhil/geolip-beatrix-sana with Sana:
# Load the model and infer image from text import torch from app.sana_pipeline import SanaPipeline from torchvision.utils import save_image sana = SanaPipeline("configs/sana_config/1024ms/Sana_1600M_img1024.yaml") sana.from_pretrained("hf://AbstractPhil/geolip-beatrix-sana") image = sana( prompt='a cyberpunk cat with a neon sign that says "Sana"', height=1024, width=1024, guidance_scale=5.0, pag_guidance_scale=2.0, num_inference_steps=18, ) - Diffusers
How to use AbstractPhil/geolip-beatrix-sana with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Efficient-Large-Model/Sana_600M_512px_diffusers", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("AbstractPhil/geolip-beatrix-sana") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
geolip-beatrix-sana
Experiments on the way to conditioning Sana, a fast text-to-image diffusion model, on Beatrix, a byte-level language model from the geolip line, through a trained connector. Sana 600M at 512 px is the test bed: small enough to train and measure in minutes. The repo also keeps the experiments that led here. By default; SANA of this size looks, horrible. SD15-grade quality at best.
Each experiment has its own folder under experiments/ with a README (the question, the recipe, the rule fixed before
the run, the result), meta.json, and its configuration, weights, evaluation and logs where it has them.
Experiments
| folder | date | what | result |
|---|---|---|---|
e001_flavor_test |
2026-10-03 | Flavor test on stock Sana: the words, the dial, a random control | DIAL: steering the text conditioning moves the mood in 128 of 128 cases (+0.795 +- 0.039 per unit); a random direction of the same size moves it 14% as much; the words separate upbeat from downbeat in 255 of 256 pairs |
e002_flavor_dial_heldout |
2026-10-03 | The flavor dial, held out: each scene steered by the other half's direction | HOLDS: steered by a direction built only from the other half of the scenes, +0.788 +- 0.039 per unit (128 of 128; in-sample +0.795); the two halves' directions agree to cosine 0.997 |
e003_trainer_parity |
2026-10-03 | Sana in the trainer: the diffusion-pipe adapter against the diffusers pipeline | PASS: text conditioning identical, the transformer identical up to ordinary float rounding, and one image through the trainer's preview sampler within 0.32 brightness levels of the stock pipeline on average |
e004_lora_mood_up |
2026-10-03 | Mood LoRA: upbeat images, neutral captions | FLAVOR LORA: mood effect +1.05 +- 0.15 at scale 1, 97% of 32 held-out cells the expected way |
e005_lora_neutral_control |
2026-10-03 | Control LoRA: the model's own neutral images, neutral captions | control: mood effect +0.30 +- 0.09 at scale 1; CONTROL QUIET against e004 |
e006_lora_mood_down |
2026-10-03 | Mood LoRA, mirrored: downbeat images, neutral captions | FLAVOR LORA: mood effect -2.07 +- 0.16 at scale 1, 97% of 32 held-out cells the expected way |
e007_lora_mood_up_reseed |
2026-10-03 | Mood LoRA, replicate: a fresh draw of upbeat images | FLAVOR LORA: mood effect +1.16 +- 0.16 at scale 1, 94% of 32 held-out cells the expected way |
e008_lora_mood_up_lr_5e-5 |
2026-10-03 | Mood LoRA at half the learning rate | FLAVOR LORA: mood effect +0.96 +- 0.13 at scale 1, 97% of 32 held-out cells the expected way |
e009_lora_mood_up_lr_2e-4 |
2026-10-03 | Mood LoRA at double the learning rate | FLAVOR LORA: mood effect +1.23 +- 0.15 at scale 1, 97% of 32 held-out cells the expected way |
Reads across the LoRA sequence (rules fixed before the runs)
- The control LoRA (e005, the model's own neutral images) against the upbeat LoRA (e004): CONTROL QUIET (quiet = at most a third of e004's effect).
- e004 minus e005, cell by cell: +0.746 +- 0.107, 97% of 32 cells positive: THE MOOD COMES FROM THE IMAGES.
- The mirror (e006, downbeat images): effect -2.065 against e004's +1.050: FLAVOR LORA in the downward direction.
- A fresh draw of upbeat images (e007) against e004: REPLICATES.
| learning rate | final effect (mean +- SE) | first epoch beyond 3 SE |
|---|---|---|
| 5e-05 | +0.956 +- 0.133 | 2 |
| 0.0001 | +1.050 +- 0.153 | 2 |
| 0.0002 | +1.227 +- 0.149 | 2 |
How the mood experiments are measured
The mood judge. Each image is scored on its pixels only, with CLIP ViT-L/14: 100 x (the mean cosine similarity to three upbeat phrases, "a cheerful, upbeat image", "a happy, joyful scene", "a bright, uplifting photo", minus the same for three downbeat phrases, "a gloomy, downbeat image", "a sad, melancholy scene", "a dark, depressing photo"). On the stock model, writing upbeat words into a prompt raises this score by 1.81 over the neutral prompt (experiment e001). Content kept is the CLIP image cosine between an image and the no-LoRA image of the same prompt and seed.
The LoRA experiments train on 24 everyday scenes and are scored on 8 scenes they never saw (4 seeds each, 32 paired cells): each cell compares the same prompt and seed with and without the LoRA.
Tools
Training: diffusion-pipe (the AbstractEyes fork, model type sana),
driven by anima-trainer (notebooks/sana_colab_train.ipynb). The LoRAs
are in diffusers format: pipe.load_lora_weights("experiments/<id>/lora/epochN", weight_name="adapter_model.safetensors").
References
- Xie, Chen, Chen, Cai, Tang et al., "SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers" (2024). https://arxiv.org/abs/2410.10629
- Gandikota, Materzynska, Zhou, Torralba, Bau, "Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models" (2023). https://arxiv.org/abs/2311.12092
- Turner, Thiergart, Leech, Udell, Vazquez, Mini, MacDiarmid, "Steering Language Models With Activation Engineering" (2023). https://arxiv.org/abs/2308.10248
- Yu, Luo, Wang, Zhao, "Uncovering the Text Embedding in Text-to-Image Diffusion Models" (2024). https://arxiv.org/abs/2404.01154
- Brack, Friedrich, Hintersdorf, Struppek, Schramowski, Kersting, "SEGA: Instructing Text-to-Image Models using Semantic Guidance" (NeurIPS 2023). https://arxiv.org/abs/2301.12247
- Kwon, Jeong, Uh, "Diffusion Models already have a Semantic Latent Space" (ICLR 2023). https://arxiv.org/abs/2210.10960
- Hertz, Mokady, Tenenbaum, Aberman, Pritch, Cohen-Or, "Prompt-to-Prompt Image Editing with Cross Attention Control" (2022). https://arxiv.org/abs/2208.01626
- Yang, Feng, Huang, "EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models" (2024). https://arxiv.org/abs/2401.04608
- Radford, Kim, Hallacy et al., "Learning Transferable Visual Models From Natural Language Supervision" (2021). https://arxiv.org/abs/2103.00020
Licences
Sana's weights are Apache-2.0; its Gemma-2-2B-IT text encoder is under Google's Gemma Terms of Use. The LoRAs and results here are Apache-2.0.
- Downloads last month
- -
Model tree for AbstractPhil/geolip-beatrix-sana
Unable to build the model tree, the base model loops to the model itself. Learn more.