laya-nli-conflict-v5-ce β€” research archive (round 5, arm B), NOT delivered

⚠️ Research archive β€” NOT a delivered model. This checkpoint failed its round's acceptance gates and was never shipped. The current production head is slow-stack/laya-nli-memory-conflict (v4). Uploaded 2026-10-02 for provenance/backup while round 11 (multi-run verdict protocol) waits for Kaggle GPU quota.

Round 5 (2026-09-28) arm B: pure CE on graded soft targets, same corpus as arm A (see laya-nli-conflict-v5).

Headline results (frozen 1000-pair main val unless noted)

  • main val 0.902 βœ“ (over v4's 0.901), new-10 10/10 βœ“
  • failed: val_soft 7 errors (βœ—), polarity +2 (βœ—)
  • Key finding: graded soft targets widened the confidence band 0.28 β†’ 16.2pp in both arms (the compression-band main attack succeeded; arm B's main-val bins recovered monotonicity). Combined with arm A's collapse, this established pure CE as the round-6+ main arm.

Artifacts

file value
model.safetensors SHA256 997bd810…db8717 (full hash in archive_sha256_manifest.txt)
rl_agent_config.json Ο„(noul) = 1.0647; encoder jhu-clsp/mmBERT-base; bf16
metrics.json val_accuracy 0.902, val_ece 0.0261, n_val 1000, no_rl true
val_probs.json frozen-val probability dump (calibration analyses)

Provenance

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for slow-stack/laya-nli-conflict-v5-ce

Finetuned
(68)
this model

Dataset used to train slow-stack/laya-nli-conflict-v5-ce