Metric: bertscore

#### Description

BERTScore leverages the pre-trained contextual embeddings from BERT and matches words in candidate and reference sentences by cosine similarity. It has been shown to correlate with human judgment on sentence-level and system-level evaluation. Moreover, BERTScore computes precision, recall, and F1 measure, which can be useful for evaluating different language generation tasks. See the [README.md] file at https://github.com/Tiiiger/bert_score for more information.

How to load this metric directly with the datasets library:

					from datasets import load_metric

@inproceedings{bert-score,
}