BANKING77 routing LoRA on Decision-1

This rank-16 q/k/v/o adapter is the validation-selected seed-42 checkpoint from the four-epoch matched BANKING77 comparison. Load it with Worthify/Decision-1 at exact revision cab55818cfb6f4163f2b2597bed31cceb6ba2030. Adapter weights are unchanged from the selected checkpoint; see public-release.json for the SHA-256 and source pin. Decision-1 is a single full-weight seed preview, not a selected multi-seed foundation release. This adapter was trained and verified against its exact commit below.

Before and after a task LoRA

On the same 3,080 official BANKING77 test utterances converted to 16-option menus, task training improved accuracy for both bases. This package is the Decision-1 adapter.

Starting model Zero-shot, no task LoRA Validation-selected task LoRA Gain
Decision-1 87.60% (2,698/3,080) 93.21% (2,871/3,080) +5.62 points
Original Gemma 87.89% (2,707/3,080) 93.54% (2,881/3,080) +5.65 points

BANKING77 test accuracy before and after a task LoRA for Decision-1 and original Gemma

The 77-intent paired bootstrap intervals for the gains are +3.21 to +8.57 points on Decision-1 and +3.28 to +8.44 points on original Gemma. Method · text-free paired results. Here “zero-shot” means no BANKING77 task adapter or examples in the prompt. The baseline scoring was post-hoc; full-weight Decision-1 training included related intent data, and 28 test rows were flagged for lexical overlap. These within-base gains do not show that Decision-1 is the stronger LoRA starting point.

The task routes an utterance among a deterministic menu of 16 candidate labels that includes the correct label. It is not standard 77-way BANKING77 classification. Validation selected update 325; this adapter scored 93.21% accuracy on the 3,080-row sealed publisher test. Its matched counterpart scored 93.54%. The selected Decision-1 result was 0.32 percentage points lower than the original-Gemma control, and the other matched seed also favored original Gemma. The 77-intent cluster bootstrap interval for the selected difference was −0.94 to +0.29 points. This exploratory, domain-aligned comparison does not establish general LoRA superiority or faster convergence.

Use Worthify's numeric-policy-aware scorer with NF4/BF16 and the included numeric-policy.json. The Gemma final softcap (30) must run in FP32; plain PEFT loading does not install that hook. Scores over options are uncalibrated. The fresh pinned-Hub-base reload matched all 365 validation rows, with zero changed choices and zero option-logit difference; see public-release.json. Loading requires access to the pinned base model and any upstream Gemma terms.

Method · sealed-test report · evidence guide

Base model lineage: Google's Apache-2.0 Gemma 4 12B. BANKING77: Casanueva et al. (2020), PolyAI/banking77, CC BY 4.0, source revision 57ec275d8078af65b7731c2a98be812d844a6d6b. Publisher utterances were converted into candidate-choice decisions. This package contains no publisher utterances or base-model weights.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Worthify/decision-1-banking77-lora

Adapter
(1)
this model