YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
FLUX.1-dev Handler for H200 GPU
High-quality image generation using FLUX.1-dev optimized for NVIDIA H200 GPU with FP8 quantization.
Features
- FP8 Quantization: Native FP8 support on Hopper architecture (~1.5-2.4x speedup)
- BF16 Precision: Optimal for H200 Tensor Cores
- torch.compile: CUDA graph optimization (~1.3x speedup)
- VAE Slicing: Memory-efficient batch processing
- Large Batch Size: H200's 141GB VRAM enables 16+ images in parallel
Performance
| GPU | FP8 | Time per Image | Batch Size |
|---|---|---|---|
| H200 141GB | Yes | ~3-4 seconds | 16 |
| H100 80GB | Yes | ~4-5 seconds | 4 |
| L40S 48GB | Yes | ~5-6 seconds | 2-4 |
| A100 80GB | No | ~8-10 seconds | 4 |
GPU Requirements
| GPU | FP8 Support | Mode | Notes |
|---|---|---|---|
| H200 141GB | Yes (CC 9.0) | Full GPU | Optimal performance |
| H100 80GB | Yes (CC 9.0) | Full GPU | Full FP8 support |
| L40S 48GB | Yes (CC 8.9) | Full GPU / Offload | FP8 enabled |
| A100 80GB | No (CC 8.0) | Full GPU | BF16 only |
| A10G 24GB | No (CC 8.6) | CPU Offload | BF16 only |
API Usage
Single Image Generation
{
"inputs": "A professional portrait of a business person in a modern office",
"parameters": {
"num_inference_steps": 28,
"guidance_scale": 3.5,
"width": 1344,
"height": 768
}
}
Batch Generation
{
"inputs": [
"A sunset over mountains",
"A city skyline at night",
"A peaceful forest scene"
],
"parameters": {
"num_inference_steps": 28,
"guidance_scale": 3.5
}
}
Response Format
Single:
{
"image": "data:image/png;base64,..."
}
Batch:
[
{"image": "data:image/png;base64,..."},
{"image": "data:image/png;base64,..."}
]
Parameters
| Parameter | Default | Description |
|---|---|---|
num_inference_steps |
28 | Denoising steps (20-50 recommended) |
guidance_scale |
3.5 | Prompt adherence (3.0-7.0) |
width |
1344 | Image width |
height |
768 | Image height |
seed |
-1 | Random seed (-1 for random) |
Environment Variables
| Variable | Default | Description |
|---|---|---|
TORCH_COMPILE_MODE |
reduce-overhead | Compile mode (reduce-overhead/default/max-autotune/false) |
FIXED_BATCH_SIZE |
16 | Batch size for CUDA graphs |
NUM_INFERENCE_STEPS |
28 | Default inference steps |
ENABLE_FP8 |
auto | FP8 quantization (auto/true/false) |
FP8 Control
auto(default): Enable FP8 for GPUs with Compute Capability >= 8.9true: Force enable FP8 (will fail on unsupported GPUs)false: Disable FP8 (use BF16 only)
Deployment
- Create a new model repository on Hugging Face Hub
- Upload
handler.pyandrequirements.txt - Create a Dedicated Endpoint with H200/H100/L40S GPU
- Set Task type to "Custom"
- Add
HF_TOKENsecret for gated model access
Comparison with Other Handlers
| Handler | FP8 | Steps | Batch | Speed (H200) | Use Case |
|---|---|---|---|---|---|
| flux-schnell | No | 4 | 10 | ~1.5s | Quick previews |
| flux-dev-no-pulid | No | 28 | 4 | ~8s | Standard quality |
| flux-dev-h200 | Yes | 28 | 16 | ~3-4s | H200 optimized |
| flux-dev-pulid | No | 28 | 1 | ~10s | Character consistency |
Technical Details
FP8 Quantization
FP8 (8-bit floating point) quantization is applied to the transformer weights using torchao:
from torchao.quantization import quantize_, float8_weight_only
quantize_(pipe.transformer, float8_weight_only())
Benefits:
- ~50% memory reduction for transformer weights
- ~1.5-2.4x inference speedup on Hopper GPUs
- Minimal quality loss (typically imperceptible)
Requirements:
- Compute Capability >= 8.9 (H200, H100, L40S, RTX 4090)
- torchao >= 0.5.0
VAE Slicing
VAE slicing is enabled for memory-efficient batch processing:
pipe.enable_vae_slicing()
This processes images one at a time during VAE decoding, reducing peak memory usage.
References
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support