πŸš€ AFP-GIC: Controllable Generative Image Compression

IEEE Access paper arXiv Live demo Source code Reproducible capsule

Official pretrained model release for Adaptive Fused Prior Transfer for Controllable Generative Image Compression, published in IEEE Access (2026).

Compress your images into small bitstreams and reconstruct them when needed. AFP-GIC is a neural image codec with five bitrate settings in one pretrained model. Choose a setting, encode your image into an .afp file, and decode it into a reconstructed PNG without needing the original image.

Try it in your browser or run it locally. Use the live demo without installing anything, or follow the Python guide below to download the weights and run the codec on your own images. Local use requires the AFP-GIC inference code and its dependencies; CPU and CUDA GPU execution are supported.

πŸ€— Try Your Own Images

AFP-GIC demo showing original and reconstructed images with bitrate controls and downloadable results.

Upload your own image in the Hugging Face Space, select an operating point, compress, and decompress the downloaded bitstream. The public demo runs on CPU by default.

Try the live demo β†’

✨ Highlights

  • Make limited bandwidth go further: encode images at very low bitrates for workflows where transmission and storage budgets are tight.
  • Keep images natural-looking at small file sizes: adapt reconstruction to each image's content, guiding texture synthesis toward the source rather than using the same fixed visual guidance for every image.
  • Adjust the file-size budget without juggling models: choose among five operating points using one pretrained model, with no weight reload between settings.
  • Compress once, reconstruct when needed: save a portable bitstream for storage or transfer, then decode it separately with the compatible model. The decoder does not need the original image.
  • Try it on your own workflow: upload images in the live demo or use the Python examples to save compressed files, reconstruct images, and inspect actual file sizes and bitrates.

🧠 Compress and Decompress Your Own Images

Use the pretrained weights hosted here with the official AFP-GIC PyTorch/CompressAI implementation. Follow steps 1-4 for your own images; the dataset evaluation commands below are optional.

1. Install the Inference Code

Run these commands in a terminal. All subsequent examples assume your working directory is the cloned AFP_GIC repository root.

git clone https://github.com/yifeipet/AFP_GIC.git
cd AFP_GIC
conda create -n afp-gic python=3.9 -y
conda activate afp-gic
python -m pip install -r public_release/requirements.txt
python -m pip install huggingface_hub

The release uses Python 3.9, PyTorch 2.1.0, and torchvision 0.16.0. Use a PyTorch build compatible with your GPU and driver. CPU inference is also supported.

2. Download the Pretrained Model

Run this in Python from the repository root. It puts the weights in the location expected by the evaluation script. Subsequent calls reuse the downloaded file when unchanged.

from huggingface_hub import hf_hub_download

hf_hub_download(
    repo_id="yifeipet/AFP-GIC",
    filename="afp_gic_release.pth.tar",
    local_dir="checkpoint/afp_gic_release/model",
)

The weights include the frozen prior component, so no separate AdaCode checkpoint is needed.

3. Load the Codec

Run this setup once in your Python session before encoding or decoding. It uses the downloaded checkpoint from step 2. CUDA is selected when available; set device = "cpu" to explicitly use the CPU.

from pathlib import Path
import struct
import sys
import torch

runtime = Path("public_release/runtime").resolve()
sys.path.insert(0, str(runtime))
import eval_public_release as afp

device = "cuda:0" if torch.cuda.is_available() else "cpu"
config = afp.load_infer_config(
    str(runtime / "config/afp_gic_release.yaml"), device
)
model = afp.build_comp_model(config).to(device)
model.load_learned_weight(
    ckpt_path="checkpoint/afp_gic_release/model/afp_gic_release.pth.tar"
)
model.codec_setup()
model.eval()

4. Encode and Decode

Encode: replace input.png with your image path. Choose quality from 0 to 4. This writes an actual compressed bitstream, not a PNG renamed with a different extension. This example does not resize the input.

quality = 0
image = afp.read_real_tensor("input.png")
with torch.no_grad():
    parts = model.compress(image, quality_ind=quality)["string_list"]

bitstream = Path("compressed.afp")
bitstream.write_bytes(
    b"".join(struct.pack("<I", len(part)) + part for part in parts)
)
height, width = image.shape[-2:]
print(f"Compressed size: {bitstream.stat().st_size:,} bytes")
print(f"File bitrate: {bitstream.stat().st_size * 8 / (height * width):.6f} bpp")

Decode: read the saved .afp file and produce a reconstructed PNG. In a separate Python process, first run the model setup in step 3, then run the following block. No original image or separate quality argument is required for decoding.

data = Path("compressed.afp").read_bytes()
parts, offset = [], 0
for _ in range(3):
    if offset + 4 > len(data):
        raise ValueError("Truncated bitstream")
    size = struct.unpack_from("<I", data, offset)[0]
    offset += 4
    if size == 0 or offset + size > len(data):
        raise ValueError("Invalid payload length")
    parts.append(data[offset:offset + size])
    offset += size
if offset != len(data):
    raise ValueError("Unexpected trailing data")

with torch.no_grad():
    reconstruction, _, _ = model.decompress(parts)
afp.img_utils.imwrite("reconstruction.png", reconstruction)
File Purpose
input.png Your source image; required only for encoding
compressed.afp Compressed bitstream for storage or transfer
reconstruction.png Decoded RGB image at the original dimensions

The .afp file is decoded with AFP-GIC, not a ZIP utility or an ordinary image viewer. Its operating-point information is stored in the stream. File bitrate above includes the 12-byte length-prefix framing. Decode your own generated streams with compatible weights and software. Larger images require more memory.

πŸ§ͺ Dataset Evaluation

Kodak

Place the 24 original Kodak PNG images directly in datasets/kodak/. These terminal commands evaluate a dataset; they are separate from the single-image Python example above.

One operating point:

python public_release/test.py -d cuda:0 --dataset kodak --qualities 0

All five operating points:

python public_release/test.py -d cuda:0 --dataset kodak --qualities 0 1 2 3 4

Use -d cpu for CPU execution. Install a PyTorch build appropriate for your hardware, keeping the release's pinned versions. See the GitHub instructions for setup and metric protocols. Metric dependencies may download their weights on first use.

CLIC2020 and DIV2K

Place original images in the following directories:

datasets/
|-- kodak/
|-- CLIC/
|   `-- clic_test_images/
`-- DIV2K_valid_HR/
    `-- DIV2K_valid_HR/
python public_release/test.py -d cuda:0 --dataset clic2020_test --qualities 0 1 2 3 4
python public_release/test.py -d cuda:0 --dataset div2k_valid_hr --qualities 0 1 2 3 4

Evaluation Outputs

Set a custom output directory with:

python public_release/test.py -d cuda:0 --dataset kodak --qualities 0 --results-root results/kodak_demo

Outputs include:

results/kodak_demo/
|-- summary_all.csv
|-- comparison_pivot.csv
`-- kodak/afp_gic_release/q0/
    |-- <image_name>.png
    |-- _bitrates.csv
    |-- per_image_metrics.csv
    |-- _metrics.json
    `-- summary.json

per_image_metrics.csv records in-loop metrics; _metrics.json reports metrics on saved reconstructions. Use the paper's metric protocol when comparing results. The single-image example explicitly saves a .afp file; the dataset runner focuses on reconstructions, bitrate records, and evaluation metrics.

πŸŽ›οΈ Operating Points

Quality index 0 1 2 3 4
Nominal target bpp 0.050 0.075 0.100 0.125 0.150

These are nominal targets, not guaranteed per-image bitrates. Actual bitrates depend on image content. All five operating points use the same pretrained model.

⚑ Efficiency

Method Inference parameters Encoder latency Decoder latency
DC-VIC 151.7M 61.61 ms 98.27 ms
AFP-GIC 120.6M 81.34 ms 80.47 ms

20.5% fewer inference parameters; 18.1% lower decoder latency. Measurements use an RTX 4090 and 100 DIV2K patches of 256 x 256 pixels, as reported in Table 5. Parameter counts include frozen components. These are not online-demo response times.

πŸ“Š Reconstructed Images and Metrics

GitHub Releases provide 2,760 reconstructed images and associated metrics: 24 Kodak, 428 CLIC2020, and 100 DIV2K images at five operating points. These support baseline comparisons under matched evaluation protocols without rerunning the model.

πŸ”¬ Reproducible Capsule

Run fresh Kodak evaluation at all five operating points in the published Code Ocean capsule, with its configured environment, input images, pretrained weights, and result comparisons.

🀝 License and Attribution

The existing AFP-GIC license notice and third-party notices are retained. Original AFP-GIC additions are provided for research and evaluation use; third-party and derived components retain their applicable terms. This upload does not grant a new blanket MIT or Apache license over pretrained weights or third-party components. See the linked upstream projects for their terms.

πŸ“ Citation

@article{pei2026adaptive,
  title   = {Adaptive Fused Prior Transfer for Controllable Generative Image Compression},
  author  = {Pei, Yifei and Liu, Ying and Ling, Nam},
  journal = {IEEE Access},
  year    = {2026},
  doi     = {10.1109/ACCESS.2026.3737467},
  url     = {https://ieeexplore.ieee.org/document/11712133}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for yifeipet/AFP-GIC