Instructions to use HeshamXOR/cubex-flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
Cubex-Flash
A LoRA fine-tune of google/gemma-4-12B, trained on a 4-pillar mixture targeting
agentic tool use, reasoning, and instruction-following, exported to GGUF (Q4_K_M).
This is a v2 checkpoint, retrained after an audit of an earlier run found a data-pipeline bug (an empty-target parsing error affecting a majority of one training pillar) that produced a degenerate first checkpoint. That checkpoint was discarded rather than published; see Evaluation below for how this one holds up under held-out testing.
Training Data
| Pillar | Dataset | Examples | Purpose |
|---|---|---|---|
| Agentic coding reasoning | saidutta69/fable-5.1-premium |
3500 | Multi-turn agentic coding traces w/ CoT |
| Test-time-scaling reasoning | simplescaling/s1K |
1000 | Deliberate step-by-step reasoning traces |
| Agentic tool use | Salesforce/xlam-function-calling-60k |
2500 | Verified function-calling examples |
| Constraint alignment | allenai/tulu-3-sft-mixture |
1998 | General instruction-following, incl. constraint adherence |
Total: 8,998 examples. saidutta69/fable-5.1-premium is an
unverified community-uploaded dataset β its content has not been independently
audited beyond confirming it parses into the expected chat-message schema.
Training Procedure
- Base model:
google/gemma-4-12B, loaded 4-bit (QLoRA) via Unsloth - LoRA: r=16, alpha=16, target modules = q/k/v/o/gate/up/down proj, dropout=0
- Effective batch size: 8 (per-device 2 Γ grad-accum 4)
- Learning rate: 2e-4, cosine schedule, 10 warmup steps
- Train/eval split: 4,905 train examples (90%), held out 10% for eval
- Epochs completed: target 2 (actual progress not available in this session)
- Training runtime: 348 min
- Final train loss: 0.5297
- Hardware: single NVIDIA L4 (24GB)
Evaluation
Held-out / rule-based scoring. GSM8K test split and an xLAM slice the model never trained on are genuinely unseen; constraint-following is a custom 20-prompt suite, not the public IFEval benchmark. 95% Wilson confidence intervals shown. "Base" = the same checkpoint with the LoRA adapter disabled β not a separately downloaded model.
| Benchmark | Base | Cubex-Flash | n |
|---|---|---|---|
| GSM8K (reasoning, held-out) | 51.0% (41.3%-60.6%) | 63.0% (53.2%-71.8%) | 100 |
| xLAM tool-calling β name match | 50.0% (40.4%-59.6%) | 99.0% (94.6%-99.8%) | 100 |
| xLAM tool-calling β strict match (name+args) | 32.0% (23.7%-41.7%) | 81.0% (72.2%-87.5%) | 100 |
| Constraint-following (custom, n=20) | 0.0% (0.0%-16.1%) | 55.0% (34.2%-74.2%) | 20 |
| MMLU (general knowledge) | 67.0% (57.3%-75.4%) | 61.0% (51.2%-70.0%) | 100 |
How to read these results
Wilson 95% confidence intervals are the basis for every claim below β a gap between base and tuned is only called "real" when the intervals don't overlap.
- GSM8K (reasoning, held-out): not statistically significant at this sample size β CIs overlap
- xLAM tool-calling β name match: statistically significant improvement β non-overlapping CIs
- xLAM tool-calling β strict match (name+args): statistically significant improvement β non-overlapping CIs
- Constraint-following (custom, n=20): statistically significant improvement β non-overlapping CIs
- MMLU (general knowledge): not statistically significant at this sample size β CIs overlap
xLAM tool-calling and constraint-following are the checkpoint's clearest, statistically defensible wins. GSM8K reasoning shows a promising point-estimate gain that is not yet statistically significant at n=100 β treat it as suggestive, not proven. MMLU shows a mild, non-significant downward trend in general knowledge; worth monitoring on any future retrain rather than treating as settled either way.
Benchmark runtime
Same prompts, same token budgets, greedy decoding, single NVIDIA L4 β base took noticeably longer on every task, consistent with it generating longer, less token-efficient completions rather than the concise output Cubex-Flash produces:
Total eval time: 115 min tuned vs 344 min base for the same benchmark suite.
Limitations
- Tool-calling "strict match" uses simplified normalized string comparison on
arguments, not full semantic equivalence (e.g.
2vs2.0would count as a mismatch) β treat it as a conservative lower bound, not an exact score. - The constraint-following suite is a 20-prompt custom benchmark, not the public IFEval benchmark, and its confidence interval is correspondingly wide.
- MMLU here is a 100-question subset used as a general-knowledge canary, not the full MMLU benchmark.
- "Base" throughout refers to the same checkpoint with the LoRA adapter disabled
via
model.disable_adapter(), not a separately downloaded reference model. - All generation used greedy decoding (temperature 0) for reproducibility; real-world sampled generation may behave somewhat differently.
License
Released under CC-BY-NC-4.0 (non-commercial). This is the more conservative
choice given that Salesforce/xlam-function-calling-60k β a direct, confirmed
contributor to this checkpoint's strongest results β is itself licensed
CC-BY-NC-4.0. The base model (google/gemma-4-12B) is Apache-2.0, but the
fine-tune as a whole inherits the more restrictive term from its training data.
This is not legal advice. If you need commercial terms, retrain without the
xLAM pillar and re-evaluate before relying on an Apache-2.0 claim. Check the
license terms of saidutta69/fable-5.1-premium, simplescaling/s1K, and
allenai/tulu-3-sft-mixture independently before any commercial use β they are
not verified here.
- Downloads last month
- 49
Model tree for HeshamXOR/cubex-flash
Base model
google/gemma-4-12B
