vangard703/DPO-PairRM-5-Original-lr-1e6-iteration-5-t-7e-beta-15e3-2-iteration Text Generation • Updated Apr 21 • 2
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-2-iteration-smi-2-candidiate-6e1-confidence Updated Apr 23
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-smi-2-candidiate-6e1-confidence Updated Apr 24
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-2-iteration-6e1-confidence Updated Apr 24
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-6e1-confidence Text Generation • Updated Apr 25 • 1
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-6e1-confidence-D1-D2_smi Text Generation • Updated Apr 25 • 1
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-baseline-D1-2e-D2_smi-1 Text Generation • Updated Apr 26