lora7b_rbonly_grid

3-way stop/deliver readout head (CONTINUE / DELIVER / ABORT) fine-tuned from Qwen2.5-7B-Instruct on trajectory-prefix windows (dataset lllqaq/datal). Trainer + recipe: see split_info.json, train.log / lora.log. best/ = repo-held-out val-AUC-best checkpoint, model/ (full) or adapter/ (LoRA) = final weights. scores_*.jsonl = per-window P(ABORT) on the frozen test file.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lllqaq/lora7b_rbonly_grid

Base model

Qwen/Qwen2.5-7B
Finetuned
(3027)
this model