Text Generation
PEFT
Safetensors
gemma4
lora
qlora
sft
bash
shell
posix
jq
awk
unsloth
conversational
Instructions to use rajivmehtapy/gemma-4-e4b-shell-specialist with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use rajivmehtapy/gemma-4-e4b-shell-specialist with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/gemma-4-E4B-it-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "rajivmehtapy/gemma-4-e4b-shell-specialist") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
Gemma 4 E4B: Full-Spectrum Shell Script Specialist
This repository contains a Specialist LoRA adapter fine-tuned on google/gemma-4-E4B-it using Unsloth.
The model is specialized to act as a Senior Linux & Automation Shell Engineer, eliminating the common failure modes of generalist LLMs when writing shell scripts.
- Training Dataset:
rajivmehtapy/shell-script-specialist-dataset(1,000 unique ChatML records) - Base Architecture:
google/gemma-4-E4B-it(quantized 4-bit) - Training Method: QLoRA ($r=16, \alpha=16$)
- Final Validation Loss:
0.0793
Benchmark Audit: Specialist vs. Generalist Base Model
In side-by-side stress tests on 5 critical defensive shell traps, the specialist scored a perfect 5/5 (100%) compared to 3/5 (60%) on the base generalist model:
| Stress Test Trap | Generalist Base Model (gemma-4-E4B-it) |
Specialist Model (gemma-4-e4b-shell-specialist) |
Key Advantage |
|---|---|---|---|
| Trap 1: Filename Splitting (spaces & newlines) | PASS | PASS | Specialist uses mapfile -d '' + -print0 with --dry-run inspection flags. |
| Trap 2: Silent Pipefail & HTTP Errors | FAIL | PASS | Specialist enforces set -euo pipefail and validates HTTP codes via curl -w "%{http_code}". |
Trap 3: Unset Variable Guard (rm -rf safety) |
PASS | PASS | Specialist verifies [[ -v VAR ]] + [[ -z "$VAR" ]] and protects hidden dotfiles with .[!.]*. |
Trap 4: POSIX /bin/sh Portability |
PASS | PASS | Specialist outputs zero bashisms, clean POSIX syntax, and set -eu. |
Trap 5: Atomic JSON In-Place Update (jq) |
FAIL | PASS | Specialist prevents 0-byte file truncation via PID temp file + atomic mv -f. |
| TOTAL SCORE | 3/5 (60%) | 5/5 (100%) | +40% Empirical Safety Gain & Zero Conversational Fluff |
Capabilities & Specialization
- Production-Hardened Bash 5+: Mandatory
set -euo pipefail, explicit error trapping (trap), dry-run flags (--dry-run), and defensive variable quoting ("${VAR}"). - POSIX
/bin/shPortability: Ultra-portable scripts for minimal Alpine Linux, BusyBox, and lightweight Docker container entrypoints without Bashisms. - Complex CLI Stream Processing: Native fluency in
jq(atomic in-place updates, filters),awk(column parsing/aggregations),sed, and safe file operations (find -print0 | xargs -0). - Zero Conversational Fluff: Immediately outputs clean, syntax-validated, runnable scripts with standard error redirection (
>&2).
Training Details
- Base Model:
google/gemma-4-E4B-it(viaunsloth/gemma-4-E4B-it-unsloth-bnb-4bit) - Fine-Tuning Method: QLoRA (4-bit base model)
- Target Modules:
q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj - LoRA Parameters: Rank $r = 16$, $\alpha = 16$
- Dataset: 1,000 curated, unique ChatML records (900 train / 100 eval)
- Hardware: NVIDIA Tesla T4 GPU
- Initial Training Loss:
1.7230 - Final Evaluation Loss:
0.0793
Quick-Start Usage
from unsloth import FastLanguageModel
# 1. Load the fine-tuned specialist
model, processor = FastLanguageModel.from_pretrained(
model_name="rajivmehtapy/gemma-4-e4b-shell-specialist",
max_seq_length=1024,
load_in_4bit=True,
device_map={"": 0},
)
# 2. Enable native fast inference
FastLanguageModel.for_inference(model)
tok = getattr(processor, "tokenizer", processor)
# 3. Prompt the specialist
messages = [
{
"role": "user",
"content": "Write a bash script to archive and delete log files older than 30 days in /var/log with dry-run support."
}
]
inputs = tok.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to("cuda")
outputs = model.generate(input_ids=inputs, max_new_tokens=512)
print(tok.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
Next Steps: DPO & GRPO Alignment
When you spin up your next machine for DPO and GRPO, you can immediately resume using the following one-liners:
1. Pulling the Policy Model on the New Machine
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="rajivmehtapy/gemma-4-e4b-shell-specialist",
max_seq_length=1024,
load_in_4bit=True,
)
2. Pulling the Dataset on the New Machine
from datasets import load_dataset
dataset = load_dataset("rajivmehtapy/shell-script-specialist-dataset")
3. Ready for DPO
- Use the SFT model as the reference policy.
- Provide
(prompt, chosen, rejected)triplets (wherechosencontains defensive standards likeset -euo pipefailandrejectedcontains common bash antipatterns). - Train with
trl.DPOTrainer.
4. Ready for GRPO
- Use the SFT model as the actor model.
- Set up automated rule-based reward functions (
shellcheckreturncode, exit code in Docker sandbox, security parameter validation). - Train with
trl.GRPOTrainer.
Everything is backed up and ready for your next phase!
- Downloads last month
- 18