FactCascade: Utility-Guided Routing for Selective Factuality Verification

This repository provides the trained frozen utility-router checkpoint and the code/configuration needed to load it. It is a small NumPy histogram gradient booster (31 features, 100 trees, maximum depth 2), not a verifier language model. All learned tree splits/leaf values, histogram thresholds and train-only imputation/scaling parameters are in router.json.

The router estimates the expected signed change in correctness when a Bespoke decision replaces a cheap-verifier fallback. The frozen FactCascade allocation combines equal-weight empirical midranks of this estimate and the raw FT5/DeBERTa score gap, then selects a half-up quota of 920/3337 per component, with stable row-key ties. config.json gives the exact feature order, thresholds, quota and pinned verifier revisions.

Repository contents and quick start

Files Purpose
router.json Existing trained utility-router weights and DEV-only preprocessing
config.json Frozen31feature order,policy,thresholds and upstream verifier revisions
factcascade/ NumPy feature,checkpoint loading,rank fusion,quota,replacement,metrics and adapter code
MODEL_REFERENCE.md Model names,exact checkpoint revisions,pinned Hugging Face links and license scope
docs/USAGE.md Download/setup,inputs,batch routing and selective replacement integration
docs/FEATURES.md All31slot definitions and DEV/external adapter differences
examples/selective_replacement.py Runnable synthetic feature-to-routing-to-replacement example
VALIDATION.json, USAGE_VALIDATION.json Original serialization parity and synthetic usage checks,respectively
hf download nhnk/FactCascade --local-dir FactCascade
cd FactCascade
python -m pip install -r requirements.txt
python -B example.py
python -B examples/selective_replacement.py

The synthetic example uses invented scores/text and executes no verifier. Real deployment supplies the pinned native FT5/B192 observations and then calls Bespoke on the selected rows. Changing packet segmentation or batch boundaries changes the scientific operation;see the usage guide. The original checkpoint revision remains 08bd580bb9f81908295354a678b4c048ac72e40a;this extension changes documentation/examples,not trained parameters.

Load the existing trained checkpoint

Download the repository files, install numpy==1.26.4, and run:

import numpy as np
from factcascade.utility_model import UtilityModel

model = UtilityModel.load("router.json")
# X must contain the exact ordered 31 frozen features; zeros only illustrate shape.
X = np.zeros((1, 31), dtype=np.float64)
utility = model.predict(X)

example.py runs this loading example. No fitting is needed. The JSON loader validates the schema, finite preprocessing parameters, feature order and tree structure. There is no executable pickle checkpoint in this public repository. This custom NumPy router does not use Transformers AutoModel.

For full selective verification, callers must produce the canonical features from the frozen verifier scores and evidence geometry, combine utility/disagreement ranks within the intended component, and apply the saved replacement policy. A feature vector from a different adapter, chunking rule or score convention is not interchangeable. Included API/adapters explain the expected record fields; this lightweight release does not execute the three verifier models.

Training and evaluation scope

The saved router was trained on 3,337 DEV rows in 1,397 natural groups. Its target is per-row signed correctness change under unweighted squared loss; this surrogate does not directly optimize component-macro balanced accuracy. The final fusion policy was selected after observing DEV and external results. Results on exposed cohorts must be treated as retrospective; this release makes no claim that fusion is optimal or generally superior to disagreement or predicted-correctness routing.

This checkpoint is the original frozen full-DEV utility model. It is distinct from later diagnostic OOF/source-heldout fits and from the separate saved FCE predicted-correctness heads. Those baseline heads and verifier weights are not included.

Serialization validation

VALIDATION.json records exact parameter-structure and CPU prediction parity with the original frozen checkpoint. Validation includes 1,024 synthetic rows, NaN/infinity imputation and every learned split boundary, plus equal fusion scores and selected synthetic routes. No training, real labels, new quality benchmark or GPU inference was performed. This verifies faithful export; it is not an additional generalization experiment.

Verifier dependencies and license

FactCascade requires scores from MiniCheck-Flan-T5-Large, MiniCheck-DeBERTa-v3-Large and Bespoke-MiniCheck-7B; exact revisions are in config.json. Obtain those models separately and follow their licenses. In particular, the pinned Bespoke model card specifies CC BY-NC 4.0. The Apache-2.0 license here covers the FactCascade router checkpoint and own code, and does not relicense upstream verifier models or datasets.

Only model parameters, own code, frozen configuration and aggregate validation metadata are released. No real dataset claim/evidence text, gold labels, dataset rows, saved real example scores, private paths or credentials are included. The usage example contains explicitly invented toy text and observations.

Downloads last month
29
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support