YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Legal NER Model β v1 Baseline
Devil1710/Legal-NER-v1-baseline
What This Model Does
Takes a legal contract clause as input and identifies 5 types of legally important entities:
PARTY β who is involved
DATE β when things happen
AMOUNT β financial values
TERM β key license conditions
JURISDICTION β which law/court applies
Quick Start
from transformers import AutoModelForTokenClassification, AutoTokenizer
import torch
model = AutoModelForTokenClassification.from_pretrained(
"Devil1710/Legal-NER-v1-baseline")
tokenizer = AutoTokenizer.from_pretrained(
"Devil1710/Legal-NER-v1-baseline")
model.eval()
text = "Apple Inc agrees to pay $500,000 by January 2025."
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
preds = torch.argmax(logits, dim=-1)[0]
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
for token, pred in zip(tokens, preds):
label = model.config.id2label[pred.item()]
if label != "O":
print(f"{token:20s} {label}")
Training Data
| Source | Entity Types | Examples |
|---|---|---|
| CUAD | PARTY, DATE, JURISDICTION | 2127 |
| LEDGAR | AMOUNT (spaCy+regex) | 1175 |
| LEDGAR | TERM (keywords) | 674 |
| Total | all 5 entities | 2850 train |
Data Strategy
- CUAD answers filtered to max 8-15 words per entity
- Jurisdiction spans narrowed using keyword extraction
- AMOUNT spans verified to contain digits/currency symbols
- TERM spans extracted via legal keyword matching
- Placeholder dates (____) removed
- Stratified 70/15/15 split
Model Architecture
Base : nlpaueb/legal-bert-base-uncased
Parameters : 108,900,107
Head : Linear(768 β 11 labels)
Dropout : 0.2 (hidden + attention)
Labels : 11 (O + B/I for 5 entities)
Max length : 320 tokens
Training Configuration
Epochs : 8 (early stopping patience=3)
Best epoch : 6
Learning rate : 3e-5
Batch size : 8 per GPU Γ 2 GPUs Γ 2 accum = 32
Weight decay : 0.01
Warmup steps : 71 (10% of total)
Loss function : Standard CrossEntropyLoss
FP16 : True
Grad checkpoint : True
Performance on Test Set
| Entity | Precision | Recall | F1 |
|---|---|---|---|
| TERM | 0.833 | 0.941 | 0.884 |
| JURISDICTION | 0.788 | 0.863 | 0.824 |
| AMOUNT | 0.587 | 0.746 | 0.657 |
| DATE | 0.440 | 0.659 | 0.528 |
| PARTY | 0.000 | 0.000 | 0.000 |
| Micro F1 | 0.634 |
Train / Val / Test Summary
| Split | Micro F1 |
|---|---|
| Train | 0.663 |
| Validation | 0.595 |
| Test | 0.634 |
Train-Test gap = 0.029 β no overfitting β
Known Limitations
PARTY F1 = 0.000
CUAD PARTY answers are role names: "Distributor", "Licensee", "Company". These are indistinguishable from regular nouns. Model trained on Run 1 could not learn any signal for PARTY detection.
Fix in v2: replace CUAD PARTY with spaCy ORG/PERSON entities from LEDGAR β proper names like "Apple Inc", "Microsoft Corporation" provide clear visual signal.
DATE Precision = 0.440
DATE spans from CUAD include surrounding words: "7th day of September , 1999 ." instead of "September 1999". Model learned wide boundaries.
Fix in v2: filter DATE spans starting with lowercase words. Keep only spans starting with digits or month names.
What Works Well
- TERM F1 0.884 β keyword-extracted spans are clean
- JURISDICTION F1 0.824 β narrowing function effective
- No overfitting β train/test gap only 2.9%
- Runs on CPU β backend compatible
Superseded By
Devil1710/Legal-NER-v2 (in progress) Fixes: PARTY proper names, DATE boundary cleanup Target: Micro F1 > 0.75
Labels
LABEL_LIST = [
"O",
"B-PARTY", "I-PARTY",
"B-DATE", "I-DATE",
"B-AMOUNT", "I-AMOUNT",
"B-TERM", "I-TERM",
"B-JURISDICTION", "I-JURISDICTION"
]
- Downloads last month
- 5