Instructions to use YesNLP/chc-folk-relevance-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use YesNLP/chc-folk-relevance-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="YesNLP/chc-folk-relevance-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("YesNLP/chc-folk-relevance-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CHC Folk-Description Relevance Classifier (v10)
A seed-conditioned cross-encoder that decides whether a sentence is a folk description of a given narrow ability of the Cattell-Horn-Carroll (CHC) model: an everyday description of someone exercising that ability, as opposed to an incidental mention or an off-topic sentence. It is the relevance filter used to build the corpus in How Everyday Language Describes Cognitive Abilities: A Folk Vocabulary Resource for the Cattell-Horn-Carroll Taxonomy (Noh & Ramani, submitted). Code: github.com/Jiho-YesNLP/chc-folk-vocabulary.
The released model is an average of three checkpoints trained with different random seeds (s42/, s43/, s44/). The paper's filter and every number below use the mean of their folk_description probabilities.
Model description
| Property | Value |
|---|---|
| Base model | microsoft/deberta-v3-base (bottom 3 layers and embeddings frozen) |
| Architecture | Cross-encoder, 3-way sequence classification |
| Labels | off_topic (0), incidental_mention (1), folk_description (2), judged relative to the anchored ability |
| Input | [CLS] anchor [SEP] passage [SEP], where the anchor is the ability name and 5 of its seed sentences joined by [SEP] |
| Max length | 256 tokens |
| Ensemble | mean folk_description probability of the three seed checkpoints |
| Operating threshold | 0.55 |
The anchor makes the score depend on which ability a sentence is paired with: the same sentence should score high for the ability it describes and low for a neighboring one. The seed sentences (ten hand-written everyday sentences per ability) are in data/raw/chc_project_annotations.json in the code repository.
Training data
Training targets are soft labels: the average of two independent LLM raters' labels (Claude Opus 5.5 in conversation and gpt-6-astra through the API, both following the same annotation rubric) on the ranked annotation set, the 50 top-ranked retrieval candidates for each of 71 CHC narrow abilities (3,550 passages from Reddit, κ = 0.527 between raters). After removing 28 passages whose text also appears in the evaluation set, the dataset has 4,980 anchored examples (80/10/10 split by passage): positives, annotated negatives, and hard negatives in which a folk description of one ability is paired with a different ability.
Training details
| Hyperparameter | Value |
|---|---|
| Epochs | 8 (best validation folk-class F1 at epoch 8 for all three seeds) |
| Batch size | 16 |
| Learning rate | 2e-5, linear warmup over 10% of steps |
| Weight decay | 0.01 |
| Loss | cross-entropy against soft targets, inverse-frequency class weights |
| Seeds | 42, 43, 44 |
| Hardware | NVIDIA RTX 3090 |
| Framework | PyTorch 2.6.0 (CUDA 12.4), Transformers 4.49.0 |
Evaluation
Held-out random annotation set: 50 passages sampled uniformly from each ability's top-1,000 retrieval pool (3,550 passages, disjoint from training). A passage is positive for a rater when that rater labels it a folk description of the target ability (170 for Opus 5.5, 166 for gpt-6-astra, 4.7%); both raters' labels are pooled. Each passage is scored with 5 seeds per anchor averaged over 4 seed draws, and the three checkpoints are averaged. The threshold is the F1 maximum of a 0.05–0.95 sweep.
| Metric (folk_description, 3-model average, threshold 0.55) | Value | 95% CI |
|---|---|---|
| Precision | 0.382 | [0.331, 0.433] |
| Recall | 0.557 | [0.492, 0.620] |
| F1 | 0.453 | [0.400, 0.501] |
| Average precision (pooled) | 0.403 | |
| F1 against Opus 5.5 / gpt-6-astra | 0.448 / 0.457 |
CIs are percentile bootstrap intervals over passages (2,000 resamples, seed 42). scripts/p3_eval_classifier.py configs/p3_filter.yaml in the code repository reproduces this table. Each s*/test_metrics.json holds that checkpoint's metrics on its internal test split (drawn from the rank-biased ranked set, so not comparable to the table above).
Usage
Exact reproduction of the paper's scores uses scripts/p3_filter_passages.py from the code repository (score_ability with n_seeds: 5, seed_subsets: 4, seed: 42). The snippet below implements the same scoring:
import random
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
REPO = "YesNLP/chc-folk-relevance-classifier"
members = []
for sub in ("s42", "s43", "s44"):
tok = AutoTokenizer.from_pretrained(REPO, subfolder=sub)
model = AutoModelForSequenceClassification.from_pretrained(REPO, subfolder=sub).eval()
members.append((tok, model))
FOLK = 2 # label id of folk_description
def prob_folk(uid, name, seeds, passages, n_seeds=5, seed_subsets=4, base_seed=42):
"""Mean folk_description probability per passage: 4 anchor draws x 3 checkpoints."""
total = [0.0] * len(passages)
for k in range(seed_subsets):
rng = random.Random(f"{base_seed}-{uid}-{k}")
drawn = rng.sample(seeds, min(n_seeds, len(seeds)))
for tok, model in members:
anchor = f" {tok.sep_token} ".join([name, *drawn])
enc = tok([anchor] * len(passages), passages, truncation=True, max_length=256,
padding=True, return_tensors="pt")
with torch.no_grad():
p = torch.softmax(model(**enc).logits, dim=-1)[:, FOLK]
total = [t + q.item() for t, q in zip(total, p)]
return [t / (seed_subsets * len(members)) for t in total]
seeds = [ # the ability's ten seed sentences (data/raw/chc_project_annotations.json)
"After watching the machine jam a few times, she figured out on her own what kept setting it off.",
"Give her a handful of examples and she'll tell you what they all secretly have in common.",
# ...
]
scores = prob_folk("Gf-I", "Induction", seeds,
["He plays a couple of rounds and just works out the rule underneath."])
print([s >= 0.55 for s in scores])
Intended use and limitations
- Built to filter retrieval candidates for a research corpus, where precision and recall are traded at the 0.55 threshold; at that point about 38% of passages it accepts are false positives, so downstream extraction re-checks every passage.
- Trained and evaluated on English Reddit sentences retrieved for the 71 non-tentative CHC narrow abilities in Schneider & McGrew's (2017) definitions sheet; performance on other text sources or abilities is untested.
- Training labels come from two LLM raters, not human annotators; their agreement on the training set is moderate (κ = 0.527).
- Scores depend on the anchor: use the ability name plus its seed sentences, as in training. A bare ability name or different seeds shifts the probabilities.
Citation
@unpublished{noh2026everyday,
title = {How Everyday Language Describes Cognitive Abilities: A Folk Vocabulary Resource for the {Cattell-Horn-Carroll} Taxonomy},
author = {Noh, Jiho and Ramani, Shwetaben},
year = {2026},
note = {Manuscript submitted for publication}
}
Model tree for YesNLP/chc-folk-relevance-classifier
Base model
microsoft/deberta-v3-baseEvaluation results
- precisionself-reported0.382
- recallself-reported0.557
- f1self-reported0.453