Direct Comparison with Flamingo and OpenFlamingo

by cwei9 - opened Aug 22, 2023

Aug 22, 2023

Hi there,

Congratulations on this great success.

I noticed that in the Model Card it says, "We note that since IDEFICS was trained on PMD (which contains COCO), the evaluation numbers on COCO are not directly comparable with Flamingo and OpenFlamingo since they did not explicitly have this dataset in the training mixture."

However, as far as I know, datasets like VQAv2 and OKVQA also build on images from COCO. Are IDEFICS's results directly comparable with Flamingo and OpenFlamingo on these benchmarks as well?

Thanks.

VictorSanh

HuggingFaceM4 org Aug 28, 2023

Hi, thanks for your question!
It's not a straightforward question.
We argue that it is comparable in the "this is not the same task" sense. COCO (as an image captioning task) was part of the training and evaluation suite. however, VQA was not part of the evaluation suite.
Although some of the images might be in the training, it is also very unlikely that any of the qa samples would be in the training text verbatim, which makes it questionable whether there is leakage.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment