Gemma 4 12B Request Complexity Scorer — request-complexity-20260918-01

This is an experimental Q4_K_M GGUF export of a fine-tuned google/gemma-4-12B-it model for scoring the complexity of arbitrary user requests from 0 to 100.

The intended output is a single integer:

  • 0–10: noise, empty, trivial acknowledgements
  • 10–30: simple chat or factual requests
  • 30–50: basic reasoning, arithmetic, or single-step transformations
  • 50–70: multi-step practical or coding/debugging requests
  • 70–90: advanced expert reasoning or difficult math
  • 90–100: research-grade, olympiad-level, or highly ambiguous/complex tasks

Artifact

  • GGUF: training/outputs/model-Q4_K_M.gguf
  • Quantization: Q4_K_M
  • SHA-256: 727c993da958113e2e0cdf01d52da59d656751c8b71748e76e89ebb595465de2
  • Size: 7,381,382,848 bytes
  • Format magic: GGUF

Training data

The dataset contains 500 supervised chat records:

split records
train 350
validation 50
verification 100

Sources:

source records
synthetic 345
MATH-500-derived 145
direct MATH-500 10

The held-out verification split was not used for training.

Training summary

  • Base model: google/gemma-4-12B-it
  • Base revision: 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7
  • Steps: 8
  • Train rows: 350
  • Validation rows: 50
  • Baseline validation loss mean: 9.87062
  • Final validation loss mean: 6.706186
  • Adapter reload verified: true

Usage

Prompt the model with a user request and ask it to return only an integer from 0 to 100.

Example system instruction:

You score the complexity of the user's request. Return only one integer from 0 to 100. Do not explain.

Example:

User: Prove that there are infinitely many primes of the form 4n+3.
Assistant: 86

Limitations

This is an exploratory fine-tune with a small training set and short training run. The numeric labels are synthetic/derived and should be treated as a calibrated heuristic, not a human-certified measurement.

Held-out verification benchmark

Modal/Ollama inference was run on all 100 held-out verification prompts after public Hugging Face publication. The model was prompted with the direct request and a system instruction to return only one integer complexity score from 0 to 100.

metric value
scheduled prompts 100
responses 100
valid integer outputs in 0..100 98
parse rate 98.0%
MAE 29.44
RMSE 38.17
Pearson r 0.462
Spearman rho 0.522
exact match among valid outputs 2.0%
within ±5 among valid outputs 14.3%
within ±10 among valid outputs 24.5%
within ±20 among valid outputs 45.9%

The strongest failure mode is MATH-style prompts: the model often answers the math problem instead of scoring its difficulty. This model is therefore not production-ready as a request-complexity scorer; it is useful as a pipeline proof and as evidence for the next dataset/prompting iteration.

Reproducibility artifacts

The repository includes local run evidence files alongside the GGUF:

  • analysis-summary.json
  • dataset/manifest.json
  • training/outputs/export.json
  • receipts/training-train.json
  • receipts/training-export.json
  • scoring/verification-modal-20260918-02/numeric-metrics.json
  • scoring/verification-modal-20260918-02/predictions.csv
  • scoring/verification-modal-20260918-02/category-metrics.csv

Generated with finetune-lab run ID request-complexity-20260918-01.

License and attribution

The base model is google/gemma-4-12B-it. The Hugging Face model metadata reports license apache-2.0 with license link https://ai.google.dev/gemma/docs/gemma_4_license. Follow the upstream Gemma terms and attribution requirements when using this derivative artifact.

Downloads last month
34
GGUF
Model size
12B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EugeneEvstafev/gemma-4-12b-request-complexity-20260918-01

Quantized
(319)
this model