Cubex-Flash

A LoRA fine-tune of google/gemma-4-12B, trained on a 4-pillar mixture targeting agentic tool use, reasoning, and instruction-following, exported to GGUF (Q4_K_M).

This is a v2 checkpoint, retrained after an audit of an earlier run found a data-pipeline bug (an empty-target parsing error affecting a majority of one training pillar) that produced a degenerate first checkpoint. That checkpoint was discarded rather than published; see Evaluation below for how this one holds up under held-out testing.

Training Data

Pillar Dataset Examples Purpose
Agentic coding reasoning saidutta69/fable-5.1-premium 3500 Multi-turn agentic coding traces w/ CoT
Test-time-scaling reasoning simplescaling/s1K 1000 Deliberate step-by-step reasoning traces
Agentic tool use Salesforce/xlam-function-calling-60k 2500 Verified function-calling examples
Constraint alignment allenai/tulu-3-sft-mixture 1998 General instruction-following, incl. constraint adherence

Total: 8,998 examples. saidutta69/fable-5.1-premium is an unverified community-uploaded dataset β€” its content has not been independently audited beyond confirming it parses into the expected chat-message schema.

Training Procedure

  • Base model: google/gemma-4-12B, loaded 4-bit (QLoRA) via Unsloth
  • LoRA: r=16, alpha=16, target modules = q/k/v/o/gate/up/down proj, dropout=0
  • Effective batch size: 8 (per-device 2 Γ— grad-accum 4)
  • Learning rate: 2e-4, cosine schedule, 10 warmup steps
  • Train/eval split: 4,905 train examples (90%), held out 10% for eval
  • Epochs completed: target 2 (actual progress not available in this session)
  • Training runtime: 348 min
  • Final train loss: 0.5297
  • Hardware: single NVIDIA L4 (24GB)

Evaluation

Held-out / rule-based scoring. GSM8K test split and an xLAM slice the model never trained on are genuinely unseen; constraint-following is a custom 20-prompt suite, not the public IFEval benchmark. 95% Wilson confidence intervals shown. "Base" = the same checkpoint with the LoRA adapter disabled β€” not a separately downloaded model.

Benchmark Base Cubex-Flash n
GSM8K (reasoning, held-out) 51.0% (41.3%-60.6%) 63.0% (53.2%-71.8%) 100
xLAM tool-calling β€” name match 50.0% (40.4%-59.6%) 99.0% (94.6%-99.8%) 100
xLAM tool-calling β€” strict match (name+args) 32.0% (23.7%-41.7%) 81.0% (72.2%-87.5%) 100
Constraint-following (custom, n=20) 0.0% (0.0%-16.1%) 55.0% (34.2%-74.2%) 20
MMLU (general knowledge) 67.0% (57.3%-75.4%) 61.0% (51.2%-70.0%) 100

Cubex-Flash benchmark results

How to read these results

Wilson 95% confidence intervals are the basis for every claim below β€” a gap between base and tuned is only called "real" when the intervals don't overlap.

  • GSM8K (reasoning, held-out): not statistically significant at this sample size β€” CIs overlap
  • xLAM tool-calling β€” name match: statistically significant improvement β€” non-overlapping CIs
  • xLAM tool-calling β€” strict match (name+args): statistically significant improvement β€” non-overlapping CIs
  • Constraint-following (custom, n=20): statistically significant improvement β€” non-overlapping CIs
  • MMLU (general knowledge): not statistically significant at this sample size β€” CIs overlap

xLAM tool-calling and constraint-following are the checkpoint's clearest, statistically defensible wins. GSM8K reasoning shows a promising point-estimate gain that is not yet statistically significant at n=100 β€” treat it as suggestive, not proven. MMLU shows a mild, non-significant downward trend in general knowledge; worth monitoring on any future retrain rather than treating as settled either way.

Benchmark runtime

Same prompts, same token budgets, greedy decoding, single NVIDIA L4 β€” base took noticeably longer on every task, consistent with it generating longer, less token-efficient completions rather than the concise output Cubex-Flash produces:

Cubex-Flash benchmark runtime

Total eval time: 115 min tuned vs 344 min base for the same benchmark suite.

Limitations

  • Tool-calling "strict match" uses simplified normalized string comparison on arguments, not full semantic equivalence (e.g. 2 vs 2.0 would count as a mismatch) β€” treat it as a conservative lower bound, not an exact score.
  • The constraint-following suite is a 20-prompt custom benchmark, not the public IFEval benchmark, and its confidence interval is correspondingly wide.
  • MMLU here is a 100-question subset used as a general-knowledge canary, not the full MMLU benchmark.
  • "Base" throughout refers to the same checkpoint with the LoRA adapter disabled via model.disable_adapter(), not a separately downloaded reference model.
  • All generation used greedy decoding (temperature 0) for reproducibility; real-world sampled generation may behave somewhat differently.

License

Released under CC-BY-NC-4.0 (non-commercial). This is the more conservative choice given that Salesforce/xlam-function-calling-60k β€” a direct, confirmed contributor to this checkpoint's strongest results β€” is itself licensed CC-BY-NC-4.0. The base model (google/gemma-4-12B) is Apache-2.0, but the fine-tune as a whole inherits the more restrictive term from its training data. This is not legal advice. If you need commercial terms, retrain without the xLAM pillar and re-evaluate before relying on an Apache-2.0 claim. Check the license terms of saidutta69/fable-5.1-premium, simplescaling/s1K, and allenai/tulu-3-sft-mixture independently before any commercial use β€” they are not verified here.

Downloads last month
49
Safetensors
Model size
12B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for HeshamXOR/cubex-flash

Adapter
(11)
this model