laya-nli-conflict-v7 — research archive (round 7), NOT delivered

⚠️ Research archive — NOT a delivered model. This checkpoint failed its round's acceptance gates and was never shipped. The current production head is slow-stack/laya-nli-memory-conflict (v4). Uploaded 2026-10-02 for provenance/backup while round 11 (multi-run verdict protocol) waits for Kaggle GPU quota.

Round 7 (2026-09-29) added bare-intent high-conflict pet/doctor families (400), alt-praise rows (240) and metric rows (160) to the v6 corpus, 0/35 hygiene. Corpus 15,300 rows.

Headline results (frozen 1000-pair main val unless noted)

  • 11 PASS / 1 FAIL — all four v6 failures fixed: new-10 10/10 ✓ (bare-intent cases pass), val_soft 1 ✓, polarity ✓, conformal ✓ via the preregistered §6 protocol adopting the s2 surface-consistency rule (three-rendering p-range abstention; reported half 34/51 = 66.7% @ 33.4% abstain)
  • the one FAIL — bias diagnosis 12/14: the B2 probe (「猫叫小白/猫是白色」 name×color, compatible) flipped 0.12 → 0.90; mechanism attributed to pet-family imbalance (140 rows "pet known + pet new ⇒ true" vs only 40 hc_safe negatives) ⇒ v8 added same-subject-compatible-attribute negatives

Artifacts

file value
model.safetensors SHA256 73a4c637…23cee4 (full hash in archive_sha256_manifest.txt)
rl_agent_config.json τ(noul) = 1.1058; encoder jhu-clsp/mmBERT-base; bf16
metrics.json val_accuracy 0.903, val_ece 0.0195, n_val 1000, no_rl true
val_probs.json frozen-val probability dump (calibration analyses)

Provenance

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for slow-stack/laya-nli-conflict-v7

Finetuned
(66)
this model

Dataset used to train slow-stack/laya-nli-conflict-v7