RiskLoop Representative Model Checkpoints

This repository hosts the reproduced representative model checkpoints for RiskLoop, an NLP system for automated contract risk detection based on Legal-BERT fine-tuning on the CUAD (Contract Understanding Atticus Dataset) benchmark.

Provenance & Reproduction Note: These checkpoints are reproduced representative checkpoints created under strictly controlled, frozen training conditions (Condition A) and evaluated in RiskLoop's frozen Phase 6 test evaluation. They are NOT the original historical official checkpoints from earlier development phases, as those original historical binaries were lost. These reproduced models represent the exact frozen representative checkpoints used to generate the Phase 6 decision analysis and final test metrics.


1. Model Architecture & Training Details

  • Backbone Model: nlpaueb/legal-bert-base-uncased
  • Model Type: Single-task span classification models (QA-style start/end logit heads over transformer sequence outputs).
  • Training Strategy: Condition A (single-task fine-tuning with fixed seeds, 4 epochs, max sequence length 512, document stride 256).

2. Model Checkpoint Registry & Integrity Hashes

The repository contains three frozen model binary files (.pt PyTorch state dictionaries):

Checkpoint Filename Target Contract Task Experimental Condition Random Seed File Size (Bytes) SHA-256 Hash
run_03_best_model.pt Cap On Liability Condition A 44 438,001,498 8689ef00a0718713f4ae8bdf8b48d3441b5d9a2ccd13ae322185862551c6f674
run_04_best_model.pt Anti-Assignment Condition A 42 438,001,498 ea6a75254798342e0cc3925dd8adca1ba9b5bb60c6a631ddf14057748c01947c
run_07_best_model.pt Termination For Convenience Condition A 42 438,001,498 a8a28902ef3dbc28e706135e8599f01cfa8c0e9c4c05cc081818e3027deda119

3. Frozen Phase 6 Test Evaluation Metrics

The reproduced representative checkpoints achieved the following metrics on the held-out frozen CUAD test evaluation set (Phase 6):

Task Name Representative Checkpoint Test F1 Score Test ROC-AUC Test PR-AUC
Cap On Liability run_03_best_model.pt 0.7727273 0.9960492 0.9147721
Anti-Assignment run_04_best_model.pt 0.8083624 0.9944473 0.9146714
Termination For Convenience run_07_best_model.pt 0.7483871 0.9939074 0.7926948

4. Inference Score Interpretation

In RiskLoop's inference engine and Streamlit demo interface:

  • Output scores are uncalibrated raw logit deltas ($\text{logit}_1 - \text{logit}_0$).
  • High positive scores indicate strong model activation for clause presence.
  • Scores are not calibrated probabilities or confidence percentages.

5. Usage & Legal Disclaimer

  • Intended Use: Portfolio evaluation, academic research, and interactive open-source demonstration of contract risk detection.
  • Legal Disclaimer: This software and model predictions do NOT constitute legal advice, formal contract audit, or legal guarantee. Users should consult qualified legal professionals for actual contract review.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support