T1 Text Keyword Layer

Server-side keyword pre-filter for ECE Connect text moderation.

Components

File Purpose
hurtlex_EN.tsv HurtLex English lexicon (raw TSV)
hurtlex_categories.json Which HurtLex categories to load (ddf, re, is)
manual_slurs.json 42 curated slurs across 10 categories
threat_patterns.json 9 regex patterns for direct threats

Tier

T1 — runs on 100% of text uploads on the Connect server.

Not on-device. There is no T0 text model.

Measured performance

On 300 Kaggle hate speech tweets:

Component Hate recall
Keyword layer alone 74%
toxic-bert alone 81%
Combined (rescue) 90%

The keyword layer catches explicit slurs. toxic-bert rescues 62% of keyword misses (coded language, context-dependent hate).

Usage

import csv, json, re
from huggingface_hub import hf_hub_download

hurtlex_path = hf_hub_download('ECE-Software/t1-text-keywords', 'hurtlex_EN.tsv')
slurs_path   = hf_hub_download('ECE-Software/t1-text-keywords', 'manual_slurs.json')
threat_path  = hf_hub_download('ECE-Software/t1-text-keywords', 'threat_patterns.json')
cats_path    = hf_hub_download('ECE-Software/t1-text-keywords', 'hurtlex_categories.json')

HURTLEX_CATS = json.load(open(cats_path))['categories']
MANUAL       = json.load(open(slurs_path))
THREATS      = json.load(open(threat_path))['patterns']

keywords = {}
with open(hurtlex_path, encoding='utf-8') as f:
    for row in csv.DictReader(f, delimiter='\t'):
        cat = row.get('category', '').strip().lower()
        w = row.get('lemma', '').strip().lower()
        if cat in HURTLEX_CATS and w:
            keywords[w] = HURTLEX_CATS[cat]

for cat, words in MANUAL.items():
    for w in words:
        keywords[w] = f'manual_{cat}'

patterns = {w: re.compile(rf'\b{re.escape(w)}\b') for w in keywords}
threat_res = [re.compile(p, re.IGNORECASE) for p in THREATS]
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support