🌿 Craftly-Image-Turbo (12.8B)

⚑ Ultra-Fast High-Fidelity Diffusion Transformer by Md Mushfiqur Rahim

Craftly-Image-Turbo is a state-of-the-art distilled 8-step visual synthesis foundation model engineered for real-time inference, photorealistic detail, anime aesthetic excellence, and complex prompt execution.

Open in Colab Open in Kaggle Marimo Studio 8 Steps Resolution


🎨 Showcase Gallery

Craftly Nature Banner
🌸 Anime Girl in Nature 🌿 4K Macro Mossy Forest πŸ›οΈ Minimalist Travertine Pavilion

(Explore the full 50-image 4K showcase gallery in the πŸ“ images/ directory).


πŸš€ One-Click Interactive Studios (Colab, Kaggle, Marimo)

You can launch and generate images immediately using any of our ready-to-run studios in the notebooks/ directory:

1. 🟑 Google Colab Studio

  • Notebook: notebooks/CRAFTLY_GOOGLE_COLAB.ipynb
  • Features: Interactive UI sliders, curated prompt presets (Anime Girl, 4K Nature, Minimalist Zen, Cyberpunk City), and automatic image download to your computer.
  • Hardware: Runs smoothly on free T4 GPU, V100, L4, or A100.

2. πŸ”΅ Kaggle Studio

  • Notebook: notebooks/CRAFTLY_KAGGLE.ipynb
  • Features: Optimized for Kaggle Dual T4 or P100 GPUs, automatic VRAM clearing, batch image generation loop, and output to /kaggle/working/.

3. 🟣 Marimo Reactive Web App


πŸ’» Python Quickstart

import torch
from diffusers import DiffusionPipeline

REPO_ID = "CraftlyrobotMushfiqur/testy"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float16

# 1. Load sovereign pipeline
pipe = DiffusionPipeline.from_pretrained(
    REPO_ID,
    custom_pipeline="pipeline_craftly",
    torch_dtype=dtype,
    trust_remote_code=True
).to(device)

# 2. Ultra-Fast 8-Step Synthesis
prompt = (
    "A breathtaking masterpiece of a beautiful anime girl standing in a sunlit lush green forest meadow, "
    "gentle breeze fluttering her long silky hair and delicate white dress, glowing cherry blossoms and emerald leaves "
    "floating in the air, crystal clear sparkling stream beside her, Makoto Shinkai and Studio Ghibli aesthetic, "
    "cinematic volumetric golden hour sunbeams, radiant expressive eyes, peaceful smile, vibrant 4k anime wallpaper, hyper-detailed"
)

with torch.inference_mode():
    image = pipe(
        prompt=prompt,
        num_inference_steps=8,
        guidance_scale=0.0,
        width=1024,
        height=1024
    ).images[0]

image.save("craftly_output.png", quality=95)
print("βœ… Saved to craftly_output.png")

πŸ—οΈ Architecture Specifications

  • Pipeline Class: CraftlyImagePipeline
  • Transformer Diffusion Backbone: CraftlyImageTransformer2DModel (28 Layers, 48 Attention Heads, 64-channel Latent Input)
  • Text Conditioning Backbone: CraftlyVisionModel (36 Layers, 2560 Hidden Dim)
  • Latent Autoencoder: 16-channel Spatial VAE
  • Denoising Trajectory: Flow-Matching Euler Discrete Scheduler (Optimized 8-step Distillation)
  • Native Resolution: 1024Γ—1024 pixels

βš–οΈ Citation & Attribution

@article{craftlyimageturbo2026,
  title={Craftly-Image-Turbo: High-Velocity Visual Generation via Distilled Flow-Matching},
  author={Md Mushfiqur Rahim and Craftly AI Research Team},
  year={2026},
  publisher={Hugging Face}
}

πŸ“ Sovereign Architecture Computation Graph

Craftly-Image-Turbo Architecture Flow

Computation graph rendered from the real module graph (CraftlyTransformerBlock x28, CraftlyTextFusionBlock x2, CraftlyVisionEncoder x36). Open craftly_model_viewer.html for the interactive version.


Downloads last month
69
Safetensors
Model size
13B params
Tensor type
F32
Β·
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support