Craftly-Image-2.1

🌟 Next-Generation Large-Scale Text-to-Image Foundation Model by Craftly AI Labs

Craftly-Image-2.1 is a state-of-the-art multimodal text-to-image diffusion transformer foundation model engineered for photorealistic synthesis, typography rendering, and complex spatial generation.


πŸ—οΈ Architecture Overview

  • Core Framework: Latent Diffusion Transformer (DiT) with Flow-Matching Schedulers
  • Text Conditioning Backbone: CraftlyVision-7B-Instruct multimodal encoder
  • Latent Autoencoder: High-fidelity 16-channel Spatial VAE
  • Native Spatial Support: Dynamic aspect ratio synthesis up to 2K resolution

πŸš€ Quickstart & Inference

import torch
from diffusers import DiffusionPipeline

# Load full Craftly-Image-2.1 pipeline
pipe = DiffusionPipeline.from_pretrained(
    "CraftlyrobotMushfiqur/image",
    torch_dtype=torch.bfloat16
)
pipe.to("cuda")

prompt = "A breathtaking futuristic cyberpunk laboratory with neon reflections, ultra high resolution 8K"
image = pipe(prompt, num_inference_steps=30, guidance_scale=4.5).images[0]
image.save("craftly_output.png")

βš–οΈ Citation & Attribution

@article{craftlyimage2026,
  title={Craftly-Image-2.1: Advanced Visual Synthesis via Deep Multimodal Conditioning},
  author={Md Mushfiqur Rahim and Craftly AI Research Team},
  year={2026},
  publisher={Hugging Face}
}
Downloads last month
206
Safetensors
Model size
7B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support