Back to all metrics
Metric: bertscore πŸ“‰
Update on GitHub

How to load this metric directly with the πŸ€—/datasets library:

Copy to clipboard
from datasets import load_metric metric = load_metric("bertscore")


BERTScore leverages the pre-trained contextual embeddings from BERT and matches words in candidate and reference sentences by cosine similarity. It has been shown to correlate with human judgment on sentence-level and system-level evaluation. Moreover, BERTScore computes precision, recall, and F1 measure, which can be useful for evaluating different language generation tasks. See the [] file at for more information.


  title={BERTScore: Evaluating Text Generation with BERT},
  author={Tianyi Zhang* and Varsha Kishore* and Felix Wu* and Kilian Q. Weinberger and Yoav Artzi},
  booktitle={International Conference on Learning Representations},