YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

🎬 Jack Studio: Master Technical Archive & Engineering Playbook

Autonomous 4K Character Photography & Cinema Generation Suite
Subject: jackthelastcreator
Author & Architect: Antigravity AI & Jack
Last Updated: September 2026
Repository: /Users/gameboy/Documents/Dev Apps/Comfy + Modal
Hugging Face: chijiokejackson35/jack-lora


πŸ“‘ Table of Contents

  1. Executive Summary & System Architecture
  2. The LoRA Training Saga: What Worked, What Failed & Why
  3. Dataset Curation & Captioning Strategy
  4. The 4K Anti-Wax Realism Pipeline (ComfyUI + SeedVR2)
  5. Autonomous Pinterest & Video Intelligence Engines
  6. Cloud Infrastructure & Serverless Deployment
  7. Full Reproduction Playbook (Step-by-Step)
  8. Production File & Resource Directory

1. Executive Summary & System Architecture

This project is an end-to-end, serverless generative AI studio designed to produce 100% photorealistic 4K imagery and 97-frame cinematic videos of Jack (jackthelastcreator). The system accepts a Pinterest link or text prompt, reverse-engineers the camera optics, lighting, and composition using Google Gemini 3.8 Flash, generates high-fidelity latent representations using Tongyi Z-Image Turbo conditioned on a custom LoRA, upscales the result with SeedVR2 VideoUpscaler, and animates still frames into 4K video using Lightricks LTX-Video 2.3 22B.

flowchart TD
    A["Pinterest URL / Prompt"] --> B["Pinterest Scraper (web_ui/pinterest_gemini_engine.py)"]
    B --> C["Google Gemini 3.8 Flash Vision"]
    C -->|Optical Prompt + Negative + Aspect Ratio| D["Headless ComfyUI on Modal A100"]
    
    subgraph ComfyUI_Engine ["ComfyUI 4K Pipeline"]
        E["Z-Image Turbo BF16 (8 Steps, FlowMatch)"] --> G["Custom LoRA (comfy_jack_zimage_v1_recipe)"]
        F["Qwen 2.5 VL Text Encoder (FP8)"] --> G
        G --> H["Base KSampler (Euler Simple, Denoise 1.0)"]
        H --> I["VAE Decode (z_image_ae.safetensors)"]
        I --> J["SeedVR2 VideoUpscaler (DiT 3B FP8, input_noise_scale 0.18)"]
    end
    
    D --> ComfyUI_Engine
    J --> K["4K Master Still (PNG)"]
    K --> L["Motion Gemini Engine (web_ui/motion_gemini_engine.py)"]
    L --> M["LTX-Video 2.3 22B Distilled (Modal A100-80GB)"]
    M --> N["97-Frame 4K MP4 Video (24fps)"]

2. The LoRA Training Saga: What Worked, What Failed & Why

2.1 The V1 Success Formula

The initial model (comfy_jack_lora_v1.safetensors) produced incredible facial likeness. When we audited its original configuration (ai-toolkit/config/jack_zimage_turbo.json), we uncovered the exact parameters responsible:

  • Architecture: Pure Linear Attention LoRA (linear: 32, linear_alpha: 32).
  • Convolution Layers: 0 (conv: null, no convolution layers).
  • Lokr / Full Rank: None (lokr_full_rank: false).
  • Resolution: Multi-resolution bucketing [512, 768, 1024].
  • Dataset: 35 clean, high-resolution portrait headshots.
  • Steps & Epochs: 1,500 steps at batch size 1 (~42.8 epochs).

2.2 The Failed Experiments

Subsequent attempts to expand the dataset to multi-angle shots degraded likeness:

  1. The 3,000-Step Run (jack_zimage_turbo_3000 / jack_zimage_master_lora):
    • What was changed: Added conv: 16, enabled lokr_full_rank: true, locked resolution to square [1024], trained to 3,000 steps.
    • Result: Severe facial elongation, pinched nose, waxy/plastic skin, rigid lighting adherence, loss of prompt flexibility.
  2. The 750-Step Trial:
    • What was changed: Lowered step count to prevent overcooking, but left conv: 16 and lokr_full_rank: true enabled.
    • Result: Skull remained vertically stretched; skin looked airbrushed.

2.3 Root Cause Analysis

Parameter V1 Recipe (Success) 3,000-Step Run (Failure) Technical / Mathematical Reason
linear vs conv linear: 32, conv: 0 linear: 32, conv: 16 Convolutional LoRA layers inject rank updates into spatial feature maps. In diffusion architectures, this distorts the spatial cross-attention geometry of the face, stretching the skull and narrowing the nose bridge. Pure linear LoRA only tunes the attention projection matrices (Q, K, V).
Lokr Full Rank false true Full rank matrix reconstruction over-parameterized the adapter, memorizing camera lens focal lengths and burning rigid spatial proportions into the weights.
Resolution Bucketing [512, 768, 1024] [1024] (Locked Square) Locking training to square 1024 force-stretched vertical phone portraits and 16:9 landscape b-roll, teaching the model unnatural face proportions. Bucketing preserves true aspect ratios.
Dataset Dilution 35 portraits (100% face) 37 wide body/desk shots + 5 faces When 88% of images had small or obscured faces, the trigger token jackthelastcreator learned computer monitors, office chairs, and desks rather than facial morphology.
Step Count & Epochs 1,500 steps (32 epochs) 3,000 steps (64 epochs) Z-Image Turbo is an 8-step distilled flowmatch model. 64 epochs at lr: 1e-4 causes latent saturation (waxy plastic specular shine and blown highlights).

2.4 The Master Recipe (jack_zimage_v1_recipe)

We combined the pristine facial fidelity of V1 with multi-angle capability:


3. Dataset Curation & Captioning Strategy

3.1 The 75/25 Anchor Rule

Located at lora_training/master_dataset/, containing exactly 47 curated images:

  • Face Anchors (75% / 35 images): The original V1 portrait dataset (DSC_2667 to DSC_2723). Pristine lighting, sharp eye focus, varying facial angles.
  • Multi-Angle Coverage (25% / 12 images):
    • Side Profile: DSC_2691_crop.png, king_of_kings_crop.png.
    • 45Β° Working Angle: broll_front45_01.png to 04.png.
    • Overhead Desk View: broll_overhead_01.png to 04.png.

3.2 Captions & Angle Scoping

Every image is paired with a .txt file containing the trigger word jackthelastcreator. Captions explicitly describe the camera angle and environment so the model decouples facial identity from the scene:

  • Portrait Example: "A sharp eye-level portrait of jackthelastcreator, young Black man with clean fade haircut and short trimmed chin beard, neutral grey background."
  • Overhead Desk Example: "A high-angle bird's-eye overhead shot of jackthelastcreator seated at a dark desk typing on a laptop, hands on keyboard, modern workspace."

4. The 4K Anti-Wax Realism Pipeline (ComfyUI + SeedVR2)

4.1 Why Distilled Models Look Waxy

Z-Image Turbo is distilled down to 8 sampling steps. While fast, 8-step flowmatch schedulers lack the Brownian motion variance of traditional 30-step Euler schedulers, tending to produce overly smoothed, plastic skin.

4.2 SeedVR2 Micro-Texture Injection

To eliminate the waxy look, we engineered a two-stage pipeline:

  1. Stage 1 (Z-Image Turbo): Generates the core scene, facial structure, and composition at native resolution (1024x1024 or 768x1152).
  2. Stage 2 (SeedVR2 VideoUpscaler 4K Restoration):
    • File: jack_zimage_seedvr2_api.json
    • Model: seedvr2_ema_3b_fp8_e4m3fn.safetensors + ema_vae_fp16.safetensors.
    • input_noise_scale: 0.18: Setting this parameter to 0.18 (up from 0.06) forces the 3B diffusion transformer to inject high-frequency photographic noise, restoring natural skin pores, beard stubble, cloth weave, and 35mm film grain.
    • LoRA Strength: Kept at 1.0 for frontal/side portraits, and tuned to 0.85 for extreme overhead bird's-eye shots to prevent facial weights from distorting downward angles.

4.3 The "90-Degree Flat Lay" Prompting Trap & Resolution

  • The Problem: Prompting "Top-down overhead 90-degree flat lay of jack sitting at desk" caused catastrophic anatomical failures (floating disembodied arms, mutant third arms growing from desks, severed torsos).
  • The Cause: In diffusion training sets, "flat lay" is almost exclusively paired with stationary products, clothing, or food.
  • The Fix:
    • For angled desk shots: Prompt as "A high-angle cinematic shot from above and behind jackthelastcreator seated at a desk, visible shoulders, back of neck, and hands typing on keyboard".
    • For true perpendicular overhead shots: Prompt as "A true 90-degree direct overhead bird's-eye view looking straight down from the ceiling at a dark wooden desk at night. At the bottom of the frame, the crown of jackthelastcreator's head with clean buzzcut fade, neck, shoulders in an ochre t-shirt, arms typing on laptop".
    • Aggressive negative prompt: "mutated hands, extra limbs, floating arms, third arm, extra fingers, malformed limbs, disembodied hands, extra keyboard".

5. Autonomous Pinterest & Video Intelligence Engines

5.1 Pinterest & Gemini 3.8 Flash Vision Engine

Located at web_ui/pinterest_gemini_engine.py:

  1. Scraper: Resolves shortlinks (pin.it/...) and Pinterest pin pages, bypassing redirects and extracting the uncompressed source image (i.pinimg.com/736x/...).
  2. Gemini 3.8 Flash Analysis:
    • Reverse-engineers lighting style, color gels, wardrobe, lens focal length, depth of field, and aspect ratio.
    • Substitutes the subject with jackthelastcreator (young Black man, clean fade, short trimmed goatee).
    • Injects optical realism keywords (Kodak Portra 400, 35mm DSLR capture, visible epidermal pores, natural skin sheen) while banning generic AI buzzwords (photorealistic, 8k, hyperrealistic).
    • Automatically enforces camera perspective rules for desk and overhead shots.

5.2 LTX-2.3 22B Image-to-Video Animator

Located at web_ui/motion_gemini_engine.py and modal_generate_jack_video.py:

  1. Analyzes a still image to determine physically plausible camera motion (slow push-in, subtle pan) and human micro-movements (breathing, eye blinks, typing fingers).
  2. Deploys LTX-Video 2.3 22B Distilled on Modal A100-80GB GPU.
  3. Renders 97 frames at 24fps (~4.04 seconds) with SeedVR2 spatial upscaling.

6. Cloud Infrastructure & Serverless Deployment

6.1 Modal Volumes & Zero-Idle Cost Architecture

All models and outputs are persisted across dedicated Modal Volumes on workspace the-john-secret:

  • jack-zimage-seedvr2-volume: Foundation diffusion models, VAEs, text encoders, ComfyUI custom nodes, and converted LoRAs.
  • ostris-outputs: AI-Toolkit training checkpoints, database, and intermediate validation samples.
  • flux-lora-models: LTX-Video base checkpoints and text encoders.

Zero Idle Cost ($0.00): All Modal worker functions spin down to 0 replicas immediately after execution. The deployed web studio (jack-studio-web) is serverless ASGI; it consumes zero GPU credits while waiting for user requests.

6.2 Headless ComfyUI Process Tree

Our execution scripts (modal_headless_zimage_seedvr2.py and modal_cloud_studio.py) run ComfyUI headlessly:

  1. Boot ComfyUI locally inside the container on 127.0.0.1:8188.
  2. Symlink models from /cache directly into ComfyUI's internal folder hierarchy.
  3. Submit workflow JSON to POST /prompt.
  4. Poll GET /history for task completion.
  5. Extract rendered image bytes, write to output, and terminate the process tree (os.killpg).

6.3 The Cloud Web App & Next.js Mobile Frontend

  • Live Cloud Web Studio (Permanent URL):
    πŸ‘‰ https://the-john-secret--jack-studio-web-web-endpoint.modal.run
  • Next.js Fullstack App: Located at jack-studio-next/. Built with React 19, Tailwind CSS, TypeScript, and Lucide icons for mobile browser testing.

7. Full Reproduction Playbook (Step-by-Step)

If you ever need to retrain or redeploy from scratch, execute these exact steps:

Step 1: Train the Master LoRA on Modal A100

# Launch detached training on cloud GPU
.venv311/bin/modal run --detach modal_train_v1_recipe.py
  • Monitors 1,500 steps at 1.7 it/s (15 mins).
  • Checkpoints save to volume ostris-outputs every 250 steps.
  • Automatically converts final checkpoint to ComfyUI format and saves to /cache/models/loras/comfy_jack_zimage_v1_recipe.safetensors.
  • Dispatches Telegram alert upon completion.

Step 2: Test Headless 4K Generation

# Test Frontal Portrait
.venv311/bin/modal run modal_headless_zimage_seedvr2.py \
  --prompt "A sharp cinematic portrait of jackthelastcreator in a modern studio, warm directional key lighting, visible skin pores" \
  --output "rendered_clips/test_portrait.png"

# Test True 90-Degree Bird's-Eye Desk View
.venv311/bin/modal run modal_headless_zimage_seedvr2.py \
  --prompt "A true 90-degree direct overhead bird's-eye view looking straight down from the ceiling at a dark wooden desk at night. Crown of head of jackthelastcreator with clean buzzcut fade at bottom edge, arms typing on laptop, warm desk lamp, notebook, raw 35mm film photograph" \
  --negative-prompt "mutated hands, extra limbs, floating arms, third arm, extra fingers, 3d render, plastic skin" \
  --width 768 --height 1152 \
  --output "rendered_clips/test_overhead.png"

Step 3: Deploy the Serverless Cloud Studio Web App

# Deploy permanent web app to Modal
.venv311/bin/modal deploy modal_cloud_studio.py
  • Endpoint: https://the-john-secret--jack-studio-web-web-endpoint.modal.run
  • Runs 24/7 at $0 cost while idle.

8. Production File & Resource Directory

File / Resource Location / URI Description
Production LoRA chijiokejackson35/jack-lora / /cache/models/loras/ comfy_jack_zimage_v1_recipe.safetensors (162.2 MB)
Original V1 LoRA chijiokejackson35/jack-lora / /cache/models/loras/ comfy_jack_lora_v1.safetensors (162.2 MB)
Master Training Config jack_zimage_v1_recipe.yaml Full training recipe specification
Modal Training Runner modal_train_v1_recipe.py Cloud training daemon with Telegram alerts
ComfyUI 4K Workflow jack_zimage_seedvr2_api.json Z-Image Turbo + SeedVR2 API execution graph
Headless 4K Runner modal_headless_zimage_seedvr2.py Python CLI for running headless renders
Pinterest Gemini Engine web_ui/pinterest_gemini_engine.py Pin scraper & Gemini 3.8 Flash optical prompt extractor
Motion Gemini Engine web_ui/motion_gemini_engine.py Motion dynamics analysis for video animation
Cloud Studio Web App modal_cloud_studio.py Serverless FastAPI web server on Modal
Next.js Studio App jack-studio-next/ Mobile-first Next.js React frontend
Master Training Dataset lora_training/master_dataset/ 47 curated images + 47 matching .txt captions
Telegram Alerts Bot Bot: 8934320817:..., Chat: 6291627175 Automatic completion & error dispatch
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support