satyam-arora-iiit-hyderabad/babyshark-sail2017-stage2
Viewer • Updated • 25.2k • 99
CRA soft global routing (M = clip(s,0)/global_max scaling).
Part of InterpAdapt (BabyShark team, IIIT Hyderabad): interpretability-guided circuit routing for Hindi-English LoRA fine-tuning on Qwen/Qwen2.5-1.5B (base, fp16). Evaluated on SAIL-2017 Romanized (Hinglish) code-mixed sentiment (3-class).
| Metric | Base Qwen2.5-1.5B | + adapter |
|---|---|---|
| Accuracy | 0.3571 | 0.5976 |
| Macro-F1 | 0.3188 | 0.5665 |
soft · Active head-blocks: 362q_proj + o_proj (alpha 16, dropout 0.05), 600 steps, lr 1e-4, max_len 256, seed 0.en_hi-latn mean recovery).ckpt.pt — torch.save payload: adapter state dict (+ optimizer/scheduler/step/RNG for resume).results.json — full eval metrics and run metadata.This is not a plain PEFT adapter — it uses the custom MaskedLoRALinear routing surface.
Reload with the project code (https://github.com/bala-skv/InterpAdapt-Hinglish-finetuning):
import torch
# from the jawed/ sub-project:
from circuit_routing.config import load_config
# build base + inject_masked_lora exactly as in
# scripts/train_cra_compare.py, then:
ckpt = torch.load("ckpt.pt", map_location="cpu")
# load ckpt["adapter"] into the model (see load_checkpoint in that script).
Numbers are ADA
fp16. CRA arms share seed 0 and data order; only the head mask differs.
Base model
Qwen/Qwen2.5-1.5B