Gemma 4 E4B: Full-Spectrum Shell Script Specialist

This repository contains a Specialist LoRA adapter fine-tuned on google/gemma-4-E4B-it using Unsloth.

The model is specialized to act as a Senior Linux & Automation Shell Engineer, eliminating the common failure modes of generalist LLMs when writing shell scripts.

  • Training Dataset: rajivmehtapy/shell-script-specialist-dataset (1,000 unique ChatML records)
  • Base Architecture: google/gemma-4-E4B-it (quantized 4-bit)
  • Training Method: QLoRA ($r=16, \alpha=16$)
  • Final Validation Loss: 0.0793

Benchmark Audit: Specialist vs. Generalist Base Model

In side-by-side stress tests on 5 critical defensive shell traps, the specialist scored a perfect 5/5 (100%) compared to 3/5 (60%) on the base generalist model:

Stress Test Trap Generalist Base Model (gemma-4-E4B-it) Specialist Model (gemma-4-e4b-shell-specialist) Key Advantage
Trap 1: Filename Splitting (spaces & newlines) PASS PASS Specialist uses mapfile -d '' + -print0 with --dry-run inspection flags.
Trap 2: Silent Pipefail & HTTP Errors FAIL PASS Specialist enforces set -euo pipefail and validates HTTP codes via curl -w "%{http_code}".
Trap 3: Unset Variable Guard (rm -rf safety) PASS PASS Specialist verifies [[ -v VAR ]] + [[ -z "$VAR" ]] and protects hidden dotfiles with .[!.]*.
Trap 4: POSIX /bin/sh Portability PASS PASS Specialist outputs zero bashisms, clean POSIX syntax, and set -eu.
Trap 5: Atomic JSON In-Place Update (jq) FAIL PASS Specialist prevents 0-byte file truncation via PID temp file + atomic mv -f.
TOTAL SCORE 3/5 (60%) 5/5 (100%) +40% Empirical Safety Gain & Zero Conversational Fluff

Capabilities & Specialization

  • Production-Hardened Bash 5+: Mandatory set -euo pipefail, explicit error trapping (trap), dry-run flags (--dry-run), and defensive variable quoting ("${VAR}").
  • POSIX /bin/sh Portability: Ultra-portable scripts for minimal Alpine Linux, BusyBox, and lightweight Docker container entrypoints without Bashisms.
  • Complex CLI Stream Processing: Native fluency in jq (atomic in-place updates, filters), awk (column parsing/aggregations), sed, and safe file operations (find -print0 | xargs -0).
  • Zero Conversational Fluff: Immediately outputs clean, syntax-validated, runnable scripts with standard error redirection (>&2).

Training Details

  • Base Model: google/gemma-4-E4B-it (via unsloth/gemma-4-E4B-it-unsloth-bnb-4bit)
  • Fine-Tuning Method: QLoRA (4-bit base model)
  • Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • LoRA Parameters: Rank $r = 16$, $\alpha = 16$
  • Dataset: 1,000 curated, unique ChatML records (900 train / 100 eval)
  • Hardware: NVIDIA Tesla T4 GPU
  • Initial Training Loss: 1.7230
  • Final Evaluation Loss: 0.0793

Quick-Start Usage

from unsloth import FastLanguageModel

# 1. Load the fine-tuned specialist
model, processor = FastLanguageModel.from_pretrained(
    model_name="rajivmehtapy/gemma-4-e4b-shell-specialist",
    max_seq_length=1024,
    load_in_4bit=True,
    device_map={"": 0},
)

# 2. Enable native fast inference
FastLanguageModel.for_inference(model)

tok = getattr(processor, "tokenizer", processor)

# 3. Prompt the specialist
messages = [
    {
        "role": "user",
        "content": "Write a bash script to archive and delete log files older than 30 days in /var/log with dry-run support."
    }
]

inputs = tok.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to("cuda")
outputs = model.generate(input_ids=inputs, max_new_tokens=512)
print(tok.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))

Next Steps: DPO & GRPO Alignment

When you spin up your next machine for DPO and GRPO, you can immediately resume using the following one-liners:

1. Pulling the Policy Model on the New Machine

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="rajivmehtapy/gemma-4-e4b-shell-specialist",
    max_seq_length=1024,
    load_in_4bit=True,
)

2. Pulling the Dataset on the New Machine

from datasets import load_dataset

dataset = load_dataset("rajivmehtapy/shell-script-specialist-dataset")

3. Ready for DPO

  • Use the SFT model as the reference policy.
  • Provide (prompt, chosen, rejected) triplets (where chosen contains defensive standards like set -euo pipefail and rejected contains common bash antipatterns).
  • Train with trl.DPOTrainer.

4. Ready for GRPO

  • Use the SFT model as the actor model.
  • Set up automated rule-based reward functions (shellcheck returncode, exit code in Docker sandbox, security parameter validation).
  • Train with trl.GRPOTrainer.

Everything is backed up and ready for your next phase!

Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rajivmehtapy/gemma-4-e4b-shell-specialist

Adapter
(362)
this model