YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

ProSRL: Prototype-guided Semantic Representation Learning for Cross-Project Defect Prediction

This repository provides the source code and experimental resources for ProSRL (Prototype-guided Semantic Representation Learning).

Requirements

ProSRL is implemented in Python3.8 and is recommended to be executed on Linux, ideally Ubuntu 20.04.6 LTS.

Dataset

ProSRL is evaluated on publicly available datasets for cross-project defect prediction.

The provided data.zip contains the PROMISE dataset used by ProSRL.

After downloading the repository, extract `data.zip` into the project directory:

```bash
unzip data.zip

Running ProSRL

The complete ProSRL workflow consists of two main stages:

  1. Data preprocessing
  2. Model training and evaluation

1. Data Preprocessing

Before training ProSRL, the datasets need to be preprocessed.

The preprocessing scripts are located in the preprocess/ directory.

Step 1: Generate the Global Dictionary

First, run GenerateGlobalDictionary.py to generate the global dictionary required for subsequent preprocessing:

cd preprocess
python GenerateGlobalDictionary.py

The corresponding script is:

GenerateGlobalDictionary.py

Step 2: Preprocess the Datasets

After generating the global dictionary, run preprocess.py to preprocess the datasets:

python preprocess.py

The corresponding script is:

preprocess.py

Please complete both preprocessing steps before running the training scripts.

2. Model Training and Evaluation

After data preprocessing is completed, run the corresponding training script for the desired source-target project pair.

For example:

cd ..
python run_mmd_mp.py

The run_mmd_mp.py script performs the model training and evaluation within the same experimental process.

For each source-target project pair, the model is trained independently. The target-domain labels are not used for model optimization and are used only for evaluating the defect prediction performance.

The main evaluation metrics include:

  • F-measure
  • AUC
  • G-measure
  • MCC (Matthews Correlation Coefficient)

3. Baseline Implementations

For reproducibility, we provide links to publicly available implementations of the baseline methods used in our experiments. When an official implementation is unavailable, we follow the implementation details and experimental settings reported in the corresponding papers.

Method Implementation
TCA+ Reimplemented based on the original paper; see baseline/
CodeBERT Microsoft CodeBERT
CodeT5+ Salesforce CodeT5+
TLSTM Reimplemented based on the original paper; see baseline/
MVHR-DP Official implementation
TFDP Official implementation
TGCN Official implementation

The links above refer to publicly available implementations that can be directly accessed from the corresponding repositories. For methods without an available official implementation, we reimplemented the methods based on their original publications and experimental descriptions.

For TCA+ and TLSTM, no publicly available source code was provided by the original authors. Therefore, we implemented these methods based on their respective papers. Their implementations are included in the baseline/ directory for reference and reproducibility.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support