babyshark-cra-soft

CRA soft global routing (M = clip(s,0)/global_max scaling).

Part of InterpAdapt (BabyShark team, IIIT Hyderabad): interpretability-guided circuit routing for Hindi-English LoRA fine-tuning on Qwen/Qwen2.5-1.5B (base, fp16). Evaluated on SAIL-2017 Romanized (Hinglish) code-mixed sentiment (3-class).

Results (validation, n=1260, fp16, NVIDIA GeForce GTX 1080 Ti)

Metric Base Qwen2.5-1.5B + adapter
Accuracy 0.3571 0.5976
Macro-F1 0.3188 0.5665
  • Mask mode: soft · Active head-blocks: 362
  • Rank-8 LoRA on q_proj + o_proj (alpha 16, dropout 0.05), 600 steps, lr 1e-4, max_len 256, seed 0.
  • Head scores: Stage-1 v2 causal trace (en_hi-latn mean recovery).

Files

  • ckpt.pt — torch.save payload: adapter state dict (+ optimizer/scheduler/step/RNG for resume).
  • results.json — full eval metrics and run metadata.

How to load

This is not a plain PEFT adapter — it uses the custom MaskedLoRALinear routing surface. Reload with the project code (https://github.com/bala-skv/InterpAdapt-Hinglish-finetuning):

import torch
# from the jawed/ sub-project:
from circuit_routing.config import load_config
# build base + inject_masked_lora exactly as in
# scripts/train_cra_compare.py, then:
ckpt = torch.load("ckpt.pt", map_location="cpu")
# load ckpt["adapter"] into the model (see load_checkpoint in that script).

Numbers are ADA fp16. CRA arms share seed 0 and data order; only the head mask differs.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mohjkhan/babyshark-cra-soft

Adapter
(463)
this model

Dataset used to train mohjkhan/babyshark-cra-soft