๐Ÿ‘๏ธ๐ŸŽจ SaShi 1.0 Vision โ€” Sovereign Indian Multimodal AI Ecosystem

SaShi 1.0 Vision is an all-in-one unified Visual Intelligence & Art engine developed by ftmdeveloperz006 in Varanasi (Banaras), Uttar Pradesh, India (Two Brothers).

Model Architecture Resolution Origin


๐ŸŒŸ The 3-in-1 Unified Multimodal Engine

SaShi 1.0 Vision consolidates three essential visual AI capabilities into one unified framework:

                                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                                โ”‚   ๐Ÿ‘๏ธ SaShi 1.0 Vision   โ”‚
                                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ                           โ–ผ                           โ–ผ
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ”‚ ๐ŸŽจ Text-to-Image โ”‚        โ”‚ ๐Ÿ–ผ๏ธ Image-to-Imageโ”‚        โ”‚ ๐Ÿ” Image-to-Text โ”‚
        โ”‚      (T2I)       โ”‚        โ”‚      (I2I)       โ”‚        โ”‚   (I2T/Vision)   โ”‚
        โ”‚ 8K Ultra-HD Art  โ”‚        โ”‚ Style & Upscale  โ”‚        โ”‚ OCR & Visual QA  โ”‚
        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
  1. Text-to-Image (T2I): Generate photorealistic 8K images, Indian cultural monuments, temples, futuristic concepts, and DSLR portraits from natural language prompts.
  2. Image-to-Image (I2I): Transform existing photos, modify styles, convert sketches to photorealistic art, and enhance lighting.
  3. Image-to-Text (I2T / Vision): Visual Question Answering (VQA), reading Devanagari/English text from receipts and signboards (OCR), and understanding complex diagrams.

๐Ÿ› ๏ธ Unified Python Code (T2I, I2I, and I2T)

import torch
from PIL import Image
from diffusers import StableDiffusionPipeline, StableDiffusionImg2ImgPipeline, EulerAncestralDiscreteScheduler
from transformers import pipeline

device = "cuda" if torch.cuda.is_available() else "cpu"

# 1. Text-to-Image (T2I)
t2i_pipe = StableDiffusionPipeline.from_pretrained(
    "ftmdeveloperz006/SaShi-1.0-Vision",
    torch_dtype=torch.float16
).to(device)

def sashi_t2i(prompt):
    return t2i_pipe(prompt + ", 8k uhd photorealistic", num_inference_steps=30).images[0]

# 2. Image-to-Image (I2I)
i2i_pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
    "ftmdeveloperz006/SaShi-1.0-Vision",
    torch_dtype=torch.float16
).to(device)

def sashi_i2i(init_image, prompt, strength=0.75):
    return i2i_pipe(prompt=prompt, image=init_image, strength=strength).images[0]

# 3. Image-to-Text (I2T / Vision Analysis)
vqa_pipe = pipeline("image-to-text", model="Salesforce/blip-image-captioning-large", device=0 if device=="cuda" else -1)

def sashi_i2t(image):
    return vqa_pipe(image)[0]["generated_text"]

๐Ÿ‡ฎ๐Ÿ‡ณ Origin & Creators

  • Name Origin: Sa-Shi (Named after the Two Brothers who conceived and trained this AI model).
  • Location: Varanasi (Kashi / Banaras), Uttar Pradesh, India.
Downloads last month
28
Safetensors
Model size
0.9B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support