CRAG Retrieval Evaluator

This model is a fine-tuned version of google-t5/t5-large for query-document relevance evaluation.

The model was developed as part of an implementation and experimentation project based on Corrective Retrieval Augmented Generation (CRAG) And experiment for efficiency of hard negative data made by using retriever model. (we used Contriever model by facebook)

Model Description

This model uses google-t5/t5-large as its base model and was fully fine-tuned on a dataset containing positive examples and hard negative examples.

The primary purpose of the model is to evaluate the relevance of a retrieved document with respect to a given query.

Base Model

  • Base model: google-t5/t5-large
  • Architecture: T5-Large
  • Parameters: approximately 770M
  • Fine-tuning: Full fine-tuning
  • Language: English

The original T5-Large model was developed by Google and is distributed under the Apache 2.0 license.

Motivation

This model was developed to improve the retrieval evaluator used in the CRAG framework.

The original CRAG framework uses a trained retrieval evaluator to assess whether retrieved passages are relevant to a given query and uses the evaluation result to determine how the retrieved knowledge should be handled.

This implementation focuses on improving the evaluator by training it with hard negative examples.

Training Data

The model was fine-tuned using a processed retrieval dataset containing positive query-document pairs and hard negative query-document pairs.

Hard negatives were constructed from retrieval results rather than being sampled randomly. This was intended to provide negative examples that are more difficult to distinguish from relevant documents.

The training dataset was specifically constructed for experiments with the CRAG retrieval evaluator.

[Mindie/nq-hard-negative]

Training Procedure

The model was initialized from google-t5/t5-large and fully fine-tuned on the training dataset.

No parameter-efficient fine-tuning method such as LoRA was used.

Training Configuration

for training almost same hyperparameters were used except epoch8 -> 7 (from model training used by CRAG)

  • Epochs: 7
  • Batch size:6
  • Learning rate: 1e-4
  • Maximum sequence length: 512
  • Optimizer: AdamW
  • Scheduler: get_linear_schedule_with_warmup
  • Hardware: A100

Intended Use

This model is primarily intended for research and experimentation involving:

  • retrieval evaluation,
  • query-document relevance classification,
  • retrieval-augmented generation,
  • corrective retrieval,
  • hard-negative training,
  • reranking and retrieval analysis.

The model can be used as a standalone query-document relevance evaluator or as a component of a larger RAG pipeline.

Limitations

The model was trained using a specific hard-negative construction procedure and dataset.

Therefore, its performance may depend on the characteristics of the retrieval system and datasets used during training.

The model should not be assumed to generalize equally well to all retrieval domains or retrieval systems.

Evaluation

The model was evaluated on a separate evaluation dataset constructed from the WikiQA development split.

Positive question-sentence pairs were obtained from the original WikiQA relevance annotations, while negative examples were randomly sampled to obtain the desired positive-to-negative ratio.

Results

The model was evaluated on a separate evaluation dataset and compared with the evaluator used in the original CRAG implementation.

Model Precision Recall F1 Accuracy
CRAG Evaluator 0.4408 0.2693 0.3344 0.5885
This Model 0.9263 0.4131 0.5714 0.7607

Confusion Matrix

Model TP FP TN FN
CRAG Evaluator 108 137 507 293
This Model 176 14 663 250

The proposed evaluator shows substantially higher precision and a lower false positive rate compared with the original CRAG evaluator. It also achieves higher recall, F1 score, and overall accuracy on the evaluation dataset.

In particular, the number of false positives was reduced from 137 to 14, while the number of true positives increased from 108 to 176.

Relation to CRAG

This model is inspired by the retrieval evaluator used in the CRAG framework.

Original CRAG implementation:

HuskyInSalt/CRAG

The main modification explored in this model is the use of a training dataset containing hard negative examples for training the retrieval evaluator.

This model is an independent fine-tuned model and is not an official model released by the authors of the CRAG paper.

How to Use

from transformers import T5ForSequenceClassification, T5Tokenizer

model_name = "Mindie/CRAG-Evaluator"

tokenizer = T5Tokenizer.from_pretrained(model_name)
model = T5ForSequenceClassification.from_pretrained(model_name)

query = "YOUR QUERY"
document = "YOUR DOCUMENT"

text = f"{query} [SEP] {document]"

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=512
)

outputs = model(**inputs)

#the output.logits would show the relevance of query and document. The value would be around -1 ~ 1  -1: no relevance 1: has relevance

### Notes

* The model is intended to evaluate the relevance between a query and a retrieved document.
* The input format should follow the format used during training.
* The maximum input length is `512` tokens.


## Citation

If you use this model, please cite the original T5 paper and the CRAG paper.

### T5

```bibtex
@article{raffel2020exploring,
  title={Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer},
  author={Raffel, Colin and Shazeer, Noam and Roberts, Adam and Lee, Katherine and Narang, Sharan and Matena, Michael and Zhou, Yanqi and Li, Wei and Liu, Peter J.},
  journal={Journal of Machine Learning Research},
  volume={21},
  number={140},
  pages={1--67},
  year={2020}
}

CRAG

https://arxiv.org/abs/2401.15884

License

This model is based on google-t5/t5-large.

Please refer to the base model's license and the licenses of the datasets used for fine-tuning when using or redistributing this model.

Acknowledgements

This work was developed based on the CRAG framework and uses google-t5/t5-large as its base model.

  • Base model: google-t5/t5-large
  • CRAG implementation: HuskyInSalt/CRAG
  • Training dataset: 'Mindie/nq-hard-negative'
  • Evaluation dataset: 'Mindie/wikiqa-eval'
Downloads last month
20
Safetensors
Model size
0.7B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Mindie/CRAG-Evaluator

Finetuned
(137)
this model

Paper for Mindie/CRAG-Evaluator