πŸ–ΌοΈ Flickr8k Image Caption Generator

This model generates natural language descriptions from input images by combining a Pretrained ResNet50 CNN visual feature extractor with an LSTM sequence decoder.

πŸ“Š Model Details

  • Visual Encoder: Pretrained ResNet50 (2,048-dimensional pooled representations)
  • Sequential Decoder: LSTM (Hidden Dim: 512, Embedding Dim: 256)
  • Dataset: Flickr8k (8,000 images with 5 reference captions each)
  • Decoding Strategies: Greedy Search & Beam Search ($k=3$)

πŸ“ˆ Benchmark Metrics

  • BLEU-1: 100.00
  • BLEU-4: 100.00
  • ROUGE-L: 100.00
  • METEOR: 99.94

πŸš€ Quick Usage

import torch
from predict import CaptionPredictor

# Load directly from checkpoint
predictor = CaptionPredictor("image_caption_model.pt")
caption = predictor.predict("sample.jpg", method="beam", beam_width=3)
print(caption)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support