gemma4-e4b-arcade-GGUF

Gemma 4 E4B, LoRA fine-tuned for ARCade (ARC aid via AI): run the Griffin nanopore preprocessing pipeline (Dorado basecalling, demultiplexing, FASTQ, FastQC/NanoPlot/MultiQC) on the University of Calgary ARC cluster by chatting.

ARCade downloads this file itself; you do not need to fetch it by hand.

Files

File Size MD5
gemma4-e4b-arcade-Q6_K.gguf 6,172,078,912 B 60305388062e2afea94e06ef52c66bde

Prompt format

ARCade renders the chat template itself and calls llama.cpp's /completion endpoint. Tool calls come out as

<|tool_call>call:TOOL_NAME{{"arg": "value"}}<tool_call|>

Training

  • Base: google/gemma-4-e4b-it, snapshot fee6332c1abaafb77f6f9624236c63aa2f1d0187.
  • LoRA with mlx-lm 0.31.3: rank 16, 16 layers, batch 4, lr 1e-4.
    • 320 iterations on the first dataset.
    • Then three 80-iteration continuations on corrected data.
  • Data: 1,992 synthetic multi-turn training conversations. They cover the one-command workflow, the four steps run one at a time, job status, results, resources, and the questions participants ask.
  • Merged, converted with convert_hf_to_gguf.py --outtype f16, quantized with llama-quantize Q6_K. Q6_K, not Q4_K_M: on 60 held-out turns Q6_K matched the f16 model (48 vs 47 exact), Q4_K_M dropped to 37.

Tested

Tested on ARC on 2026-09-30:

  • With ARCade 0.2.0 and this file.
  • On the same llama.cpp server participants run (arcade serve, cpu2023, pinned container build 10991).
  • Each line is one chat turn, asked in order in one conversation.
  • The tool results in the two tables below were replayed, so no jobs were submitted. A run with real SLURM jobs is recorded after them.

A turn passes when:

  • ARCade calls the expected tool (for example propose_pipeline with the right step, or job_status); or
  • for a text answer, the reply contains the required fact and does not invent a submission.

Workshop questions, as written: 22/22

# Turn Result
1 Where is the example data? Is it ready? PASS
2 What is a POD5 file? PASS
3 Can you submit a job to basecall with raw data in pod5 folder? The kit is SQK-RBK114-96. PASS
4 Yes, please submit it. PASS
5 What exact command did you submit? PASS
6 What is the status of my job? PASS
7 Is my job done? PASS
8 How many resources did my job use, and why a GPU? PASS
9 Can you submit a job to demux the pod5 data? PASS
10 yes PASS
11 Can you convert the reads to FASTQ? PASS
12 yes please PASS
13 Where can I find the scripts you submitted? PASS
14 Can you run the QC pipeline? PASS
15 yes PASS
16 Did it work? Are the results OK? PASS
17 Where are my results, and which file should I open first? PASS
18 What will the pipeline do to my data? PASS
19 Now preprocess these POD5 files with a single command. PASS
20 Yes, submit it. PASS
21 Is my job done? PASS
22 How do I stop the model when I am done? PASS

The same questions, reworded: 22/22

# Turn Result
1 is the demo data there? can I use it? PASS
2 what's pod5? PASS
3 please basecall the pod5 folder, kit SQK-RBK114-96 PASS
4 ok go ahead PASS
5 show me the command you ran PASS
6 how's my job going? PASS
7 has it finished? PASS
8 what resources did that job use? why does it need a gpu? PASS
9 now demultiplex the pod5 data PASS
10 yes please PASS
11 next, make fastq files PASS
12 sure PASS
13 where did you save the job scripts? PASS
14 run qc now PASS
15 yep PASS
16 are the results good? PASS
17 which output file should I look at first? PASS
18 what does this pipeline actually do? PASS
19 do everything on pod5 in one go PASS
20 yes PASS
21 done yet? PASS
22 how do I shut the model down? PASS

With real SLURM jobs: 22/22 and 22/22

On 2026-09-30, both sets were asked again on a fresh install from the public installer (curl -fsSL https://thebiohub.ca/install/arcade.sh | bash), on the demo POD5.

  • Every step was a real job.
  • The chat waited for each job to finish before the next question.
  • Both sets passed: 22/22 as written and 22/22 reworded.
  • Each set ran 6 jobs, and all 12 COMPLETED:
    • 4 step jobs;
    • 2 jobs for the single-command run.
  • The QC check reported 1,006 reads, 77.4% assigned to a barcode. dorado summary gives the same figures: barcode05 708, barcode06 71, unclassified 227.

License

Apache 2.0, inherited from google/gemma-4-e4b-it.

Downloads last month
265
GGUF
Model size
7B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support