Instructions to use Mindie/CRAG-Evaluator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mindie/CRAG-Evaluator with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Mindie/CRAG-Evaluator")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Mindie/CRAG-Evaluator") model = AutoModelForSequenceClassification.from_pretrained("Mindie/CRAG-Evaluator", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CRAG Retrieval Evaluator
This model is a fine-tuned version of google-t5/t5-large for query-document relevance evaluation.
The model was developed as part of an implementation and experimentation project based on Corrective Retrieval Augmented Generation (CRAG) And experiment for efficiency of hard negative data made by using retriever model. (we used Contriever model by facebook)
Model Description
This model uses google-t5/t5-large as its base model and was fully fine-tuned on a dataset containing positive examples and hard negative examples.
The primary purpose of the model is to evaluate the relevance of a retrieved document with respect to a given query.
Base Model
- Base model:
google-t5/t5-large - Architecture: T5-Large
- Parameters: approximately 770M
- Fine-tuning: Full fine-tuning
- Language: English
The original T5-Large model was developed by Google and is distributed under the Apache 2.0 license.
Motivation
This model was developed to improve the retrieval evaluator used in the CRAG framework.
The original CRAG framework uses a trained retrieval evaluator to assess whether retrieved passages are relevant to a given query and uses the evaluation result to determine how the retrieved knowledge should be handled.
This implementation focuses on improving the evaluator by training it with hard negative examples.
Training Data
The model was fine-tuned using a processed retrieval dataset containing positive query-document pairs and hard negative query-document pairs.
Hard negatives were constructed from retrieval results rather than being sampled randomly. This was intended to provide negative examples that are more difficult to distinguish from relevant documents.
The training dataset was specifically constructed for experiments with the CRAG retrieval evaluator.
[Mindie/nq-hard-negative]
Training Procedure
The model was initialized from google-t5/t5-large and fully fine-tuned on the training dataset.
No parameter-efficient fine-tuning method such as LoRA was used.
Training Configuration
for training almost same hyperparameters were used except epoch8 -> 7 (from model training used by CRAG)
- Epochs: 7
- Batch size:6
- Learning rate: 1e-4
- Maximum sequence length: 512
- Optimizer: AdamW
- Scheduler: get_linear_schedule_with_warmup
- Hardware: A100
Intended Use
This model is primarily intended for research and experimentation involving:
- retrieval evaluation,
- query-document relevance classification,
- retrieval-augmented generation,
- corrective retrieval,
- hard-negative training,
- reranking and retrieval analysis.
The model can be used as a standalone query-document relevance evaluator or as a component of a larger RAG pipeline.
Limitations
The model was trained using a specific hard-negative construction procedure and dataset.
Therefore, its performance may depend on the characteristics of the retrieval system and datasets used during training.
The model should not be assumed to generalize equally well to all retrieval domains or retrieval systems.
Evaluation
The model was evaluated on a separate evaluation dataset constructed from the WikiQA development split.
Positive question-sentence pairs were obtained from the original WikiQA relevance annotations, while negative examples were randomly sampled to obtain the desired positive-to-negative ratio.
Results
The model was evaluated on a separate evaluation dataset and compared with the evaluator used in the original CRAG implementation.
| Model | Precision | Recall | F1 | Accuracy |
|---|---|---|---|---|
| CRAG Evaluator | 0.4408 | 0.2693 | 0.3344 | 0.5885 |
| This Model | 0.9263 | 0.4131 | 0.5714 | 0.7607 |
Confusion Matrix
| Model | TP | FP | TN | FN |
|---|---|---|---|---|
| CRAG Evaluator | 108 | 137 | 507 | 293 |
| This Model | 176 | 14 | 663 | 250 |
The proposed evaluator shows substantially higher precision and a lower false positive rate compared with the original CRAG evaluator. It also achieves higher recall, F1 score, and overall accuracy on the evaluation dataset.
In particular, the number of false positives was reduced from 137 to 14, while the number of true positives increased from 108 to 176.
Relation to CRAG
This model is inspired by the retrieval evaluator used in the CRAG framework.
Original CRAG implementation:
The main modification explored in this model is the use of a training dataset containing hard negative examples for training the retrieval evaluator.
This model is an independent fine-tuned model and is not an official model released by the authors of the CRAG paper.
How to Use
from transformers import T5ForSequenceClassification, T5Tokenizer
model_name = "Mindie/CRAG-Evaluator"
tokenizer = T5Tokenizer.from_pretrained(model_name)
model = T5ForSequenceClassification.from_pretrained(model_name)
query = "YOUR QUERY"
document = "YOUR DOCUMENT"
text = f"{query} [SEP] {document]"
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=512
)
outputs = model(**inputs)
#the output.logits would show the relevance of query and document. The value would be around -1 ~ 1 -1: no relevance 1: has relevance
### Notes
* The model is intended to evaluate the relevance between a query and a retrieved document.
* The input format should follow the format used during training.
* The maximum input length is `512` tokens.
## Citation
If you use this model, please cite the original T5 paper and the CRAG paper.
### T5
```bibtex
@article{raffel2020exploring,
title={Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer},
author={Raffel, Colin and Shazeer, Noam and Roberts, Adam and Lee, Katherine and Narang, Sharan and Matena, Michael and Zhou, Yanqi and Li, Wei and Liu, Peter J.},
journal={Journal of Machine Learning Research},
volume={21},
number={140},
pages={1--67},
year={2020}
}
CRAG
https://arxiv.org/abs/2401.15884
License
This model is based on google-t5/t5-large.
Please refer to the base model's license and the licenses of the datasets used for fine-tuning when using or redistributing this model.
Acknowledgements
This work was developed based on the CRAG framework and uses google-t5/t5-large as its base model.
- Base model:
google-t5/t5-large - CRAG implementation:
HuskyInSalt/CRAG - Training dataset: 'Mindie/nq-hard-negative'
- Evaluation dataset: 'Mindie/wikiqa-eval'
- Downloads last month
- 20
Model tree for Mindie/CRAG-Evaluator
Base model
google-t5/t5-large