Krea2 Identity Long Caption R32 3ep
This repository contains a Krea 2 LoRA trained for long-caption adherence,
artist identity, character identity, and broader visual coverage. It is not a
standalone diffusion model and requires the krea/Krea-2-Raw base model.
The default file, krea2_identity_long_caption_r32_3ep.safetensors, was
updated on 2026-07-20 with the completed artist-balanced continuation. This
run initialized from the original completed 3-epoch identity LoRA and trained
for another 3 epochs on refreshed captions and a rebalanced dataset.
Model details
- Base model:
krea/Krea-2-Raw - Network implementation: Musubi Tuner
networks.lora_krea2 - Rank / alpha:
32 / 32 - Main learning rate:
2e-4 - Text-fusion learning-rate ratio:
0.02 - LoRA+ ratio:
2.5 - Optimizer: CAME
- Scheduler: warmup-stable-decay
- Gradient accumulation:
8 - Precision: BF16
- Epochs / optimizer steps:
3 / 2529 - Repeated training items per epoch:
162,633 - Final average loss reported by the run:
0.08539 - W&B run:
sumomo/krea2_r32_ep3_continue/grnngwuw
Per-device batches for the 768, 1024, 1280, and 1536 resolution views were
26 / 17 / 11 / 8 on 3 GPUs. With gradient accumulation 8, their effective
batch sizes per optimizer update were 624 / 408 / 264 / 192.
Excluded module patterns:
first
last\.linear
tmlp\..*
txtmlp\..*
tproj\.1
Text-fusion modules matching txtfusion\..* were trained at the low learning
rate rather than frozen.
Training data notes
The continuation uses refreshed natural-language captions and a rebalanced mixture of curated illustration, artist, character, figure/3D, Chinese 20k, and approximately 10,000 aesthetic real-world images. Low-representation artist subsets receive additional repeats, while generic figure/3D and todo subsets are downweighted. The dataset spans 768, 1024, 1280, and 1536 training views; the Chinese 20k subset does not use the 768 view.
Artist and character identities are represented in natural-language captions,
including @artist identity tokens where applicable. The real-world subset
uses refreshed detailed captions. Structured AI-generated images are kept
distinct from photographic descriptions in the captioning rules.
Usage
Load krea2_identity_long_caption_r32_3ep.safetensors as a Krea 2 LoRA with
Musubi Tuner's networks.lora_krea2 network implementation. It is intended for
the raw Krea 2 base model and does not replace the base DiT, VAE, or text
encoder.
The model does not use one universal trigger word. Prompts can use artist and
character identities present in the training captions, including artist tokens
written in the @artist_name form.
Experimental PatchGAN test checkpoint
krea2_identity_long_caption_r32_3ep_patchgan_test_step500.safetensors is an
older experimental test checkpoint derived from the previous primary LoRA. It
was trained with a DINOv3 PatchGAN objective to test whether adversarial tuning
can restore skin, hair, fabric, metal, glass, and other fine real-world textures
that may be softened by illustration-heavy SFT data.
This file is a complete replacement LoRA state. Do not stack it with the current default LoRA. It was selected at 500 discriminator updates (166 generator updates) because later checkpoints began adding meaningless high-frequency detail.
- Stored precision: BF16
- SHA256:
527e6d6a1150f9edba91574a08f26e721937498b54f1968040726c26dc527569
This checkpoint is provided strictly as a test. It may exaggerate pores and wrinkles, add scratches or colored speckles, alter facial appearance, and over-detail anime linework or decorations.
Integrity
Default model:
SHA256 f3764f57f0d558176172b9466988b382842d09da0037a77a28cafa83cf59516e
Limitations and license
The training distribution remains illustration-heavy despite the real-world supplement. Artist and character rendering quality varies with representation, visual complexity, and prompt composition.
Use of this LoRA is subject to the license and usage terms of the Krea 2 base
model. The base model declares an other license on Hugging Face.
Model tree for SumomoLee/Krea2-Identity-Long-Caption-R32-3ep
Base model
krea/Krea-2-Raw