cbjev: Laya checkpoints fine-tuned into a one-pass layout (6.7x faster on multi-question calls)

#18
by 0010101010-1 - opened

Hi! I built cbjev, a GPL-3.0 runtime plus fine-tunes of the Laya checkpoints, and wanted to share it here since it grew directly out of Laya.

The idea: Laya encodes the state once per question. cbjev packs every question of a call into one row, with an attention mask where questions never see each other but the state reads all of them. Each question segment restarts its positions after [CLS], so a one-question row is token-for-token what Laya reads. That meant I could fine-tune from laya-typed-decisions / laya-multilingual directly instead of relearning the task. cbjev.load("laya") also runs your original checkpoints unchanged (matches within bf16 rounding).

Measured side by side on one RTX 4090, byte-identical cases:

cbjev Laya (better of root / typed-decisions)
10 questions over a 500-token document 11.4 ms 75.8 ms
mean accuracy, 15 English suites 0.741 0.710
typed-decisions, 2,000 decisions 0.783 0.768
mean ECE 0.117 0.125
answer flips when options are reordered 0.2 % 7.8 %
MASSIVE, 51 languages (multilingual checkpoint) 0.436 0.401

It still trails Laya on AG News (βˆ’0.8), DAIR emotion (βˆ’2.5), prompt injection (βˆ’3.5, mostly German prompts) and support triage (βˆ’4.0); all of that is in the README.

Thanks for releasing Laya openly. None of this would exist without it. Happy to hear feedback or run it on anyone's data.

Sign up or log in to comment