YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

FLUX.1-dev Handler for H200 GPU

High-quality image generation using FLUX.1-dev optimized for NVIDIA H200 GPU with FP8 quantization.

Features

  • FP8 Quantization: Native FP8 support on Hopper architecture (~1.5-2.4x speedup)
  • BF16 Precision: Optimal for H200 Tensor Cores
  • torch.compile: CUDA graph optimization (~1.3x speedup)
  • VAE Slicing: Memory-efficient batch processing
  • Large Batch Size: H200's 141GB VRAM enables 16+ images in parallel

Performance

GPU FP8 Time per Image Batch Size
H200 141GB Yes ~3-4 seconds 16
H100 80GB Yes ~4-5 seconds 4
L40S 48GB Yes ~5-6 seconds 2-4
A100 80GB No ~8-10 seconds 4

GPU Requirements

GPU FP8 Support Mode Notes
H200 141GB Yes (CC 9.0) Full GPU Optimal performance
H100 80GB Yes (CC 9.0) Full GPU Full FP8 support
L40S 48GB Yes (CC 8.9) Full GPU / Offload FP8 enabled
A100 80GB No (CC 8.0) Full GPU BF16 only
A10G 24GB No (CC 8.6) CPU Offload BF16 only

API Usage

Single Image Generation

{
  "inputs": "A professional portrait of a business person in a modern office",
  "parameters": {
    "num_inference_steps": 28,
    "guidance_scale": 3.5,
    "width": 1344,
    "height": 768
  }
}

Batch Generation

{
  "inputs": [
    "A sunset over mountains",
    "A city skyline at night",
    "A peaceful forest scene"
  ],
  "parameters": {
    "num_inference_steps": 28,
    "guidance_scale": 3.5
  }
}

Response Format

Single:

{
  "image": "data:image/png;base64,..."
}

Batch:

[
  {"image": "data:image/png;base64,..."},
  {"image": "data:image/png;base64,..."}
]

Parameters

Parameter Default Description
num_inference_steps 28 Denoising steps (20-50 recommended)
guidance_scale 3.5 Prompt adherence (3.0-7.0)
width 1344 Image width
height 768 Image height
seed -1 Random seed (-1 for random)

Environment Variables

Variable Default Description
TORCH_COMPILE_MODE reduce-overhead Compile mode (reduce-overhead/default/max-autotune/false)
FIXED_BATCH_SIZE 16 Batch size for CUDA graphs
NUM_INFERENCE_STEPS 28 Default inference steps
ENABLE_FP8 auto FP8 quantization (auto/true/false)

FP8 Control

  • auto (default): Enable FP8 for GPUs with Compute Capability >= 8.9
  • true: Force enable FP8 (will fail on unsupported GPUs)
  • false: Disable FP8 (use BF16 only)

Deployment

  1. Create a new model repository on Hugging Face Hub
  2. Upload handler.py and requirements.txt
  3. Create a Dedicated Endpoint with H200/H100/L40S GPU
  4. Set Task type to "Custom"
  5. Add HF_TOKEN secret for gated model access

Comparison with Other Handlers

Handler FP8 Steps Batch Speed (H200) Use Case
flux-schnell No 4 10 ~1.5s Quick previews
flux-dev-no-pulid No 28 4 ~8s Standard quality
flux-dev-h200 Yes 28 16 ~3-4s H200 optimized
flux-dev-pulid No 28 1 ~10s Character consistency

Technical Details

FP8 Quantization

FP8 (8-bit floating point) quantization is applied to the transformer weights using torchao:

from torchao.quantization import quantize_, float8_weight_only
quantize_(pipe.transformer, float8_weight_only())

Benefits:

  • ~50% memory reduction for transformer weights
  • ~1.5-2.4x inference speedup on Hopper GPUs
  • Minimal quality loss (typically imperceptible)

Requirements:

  • Compute Capability >= 8.9 (H200, H100, L40S, RTX 4090)
  • torchao >= 0.5.0

VAE Slicing

VAE slicing is enabled for memory-efficient batch processing:

pipe.enable_vae_slicing()

This processes images one at a time during VAE decoding, reducing peak memory usage.

References

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support