Kitsune Tales logo Kitsune-Tales-E4B-EN-LoRA

LoRA adapter (run dpo-en-main) behind whoashish115/Kitsune-Tales-E4B-EN. The merged model card has the full evaluation.

Site | GitHub | Report | W&B | Demo | Weights | LoRA | GGUF | Dataset

Adapter

Base model google/gemma-4-E4B-it @ ee0ef6023621
Rank / alpha / dropout 32 / 64 / 0.05
Target modules every linear layer of the language model (attention and MLP)
Trainable parameters 77.8M (0.97 % of the checkpoint)
Precision bf16 weights and adapter, max length 2,048 tokens
Recipe Supervised fine-tuning on 6,647 examples, then DPO on 1,596 preference pairs (1,168 judge-labeled, 293 rule-based, 135 refusal pairs), both one epoch. This repo holds the DPO adapter, which already includes the SFT update.
Optimizer AdamW, cosine schedule, lr 2e-4 (SFT) / 2e-5 (DPO, beta 0.1), effective batch 16, seed 42

Loading

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("google/gemma-4-E4B-it", revision="ee0ef6023621cff504d758262d4e04895a5af4a2", torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, "whoashish115/Kitsune-Tales-E4B-EN-LoRA")
tok = AutoTokenizer.from_pretrained("whoashish115/Kitsune-Tales-E4B-EN-LoRA")

Use the system prompt and request format from the merged model card; the evaluation used temperature 0.8, top-p 0.95, top-k 50 and repetition penalty 1.05.

Figures

SFT training and validation loss. SFT training and validation loss.

DPO loss, held-out preference accuracy and reward margin. DPO loss, held-out preference accuracy and reward margin.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for whoashish115/Kitsune-Tales-E4B-EN-LoRA

Adapter
(387)
this model

Dataset used to train whoashish115/Kitsune-Tales-E4B-EN-LoRA

Collection including whoashish115/Kitsune-Tales-E4B-EN-LoRA