Scikit-learn
Joblib
llm-routing

Selective-fallback failure model (seeded redraw)

This file is a seed-0 refit made on 2026-10-03. The original training run wrote predictions and did not save weights. This upload is the refit.

quality_router_v2 version 8 sets learned_downrouting to false. The served card loads no fitted weights. This file is a learned model from the same study, and the served card leaves it unloaded.

Role

Calibrated L2 logistic regression of whether qwen_then_strong_repair failed. Features are TF-IDF word 1-2 grams (max_features 4000, min_df 2), a source one-hot fit on the training vocabulary, and log1p(prompt characters). The served card rejected this fit. It is not loaded.

Data

Measured pool https://huggingface.co/datasets/dhresearch/outcome-router-v2-measured-pool: stdout_tasks.jsonl, qwen_cross_outcomes.jsonl, and splits.json. The label rows are the stdin_stdout slice.

Fit

selective_fallback.fit_risk(train, 'qwen_then_strong_repair') with LogisticRegression(max_iter=3000, C=1.0, class_weight='balanced', random_state=0) inside CalibratedClassifierCV(method='sigmoid', cv=3).

Seed 0. Software at refit time: python 3.14.2, sklearn 1.9.1, numpy 2.5.3, joblib 1.6.0, lightgbm 4.7.0, scipy 1.18.1.

Agreement

Validation AUC of this refit is 0.52657. The seeded redraw stored in data/real_v2/revalidation_report.json is 0.52657. Match: True. Population train/val/test = 451/110/198. The historical unseeded AUC 0.425 is a different fit and stays the published historical number.

Load

joblib.load returns the dict from selective_fallback.fit_risk. Score with selective_fallback.score_risk. scikit-learn must match the refit version.

Class: dict(model, vectorizer, sources, base_rate).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train dhresearch/quality-router-v2-selective-fallback