Expense Tracker — Gemma 3 270M ONNX model

Fine-tuned Gemma 3 270M model for extracting expense details from spoken English sentences into structured JSON. Runs fully on-device in the Expense Tracker Android app via ONNX Runtime.

Files

File Description
model.onnx FP16 ONNX model (537 MB). Tied embedding/lm_head share one tensor via a Transpose node to halve size.
tokenizer.json Full BPE tokenizer (from tokenizers library).
tokenizer_config.json Tokenizer config.
special_tokens_map.json Special tokens.
adapter_config.json / adapter_model.safetensors Original LoRA adapter (base: google/gemma-3-270m-it, r=8, α=16, targets q_proj/v_proj).

Prompt format

The model was fine-tuned with the standard Gemma chat template:

<bos><start_of_turn>user
Extract expense details from the sentence and return JSON only. No explanation.

<INPUT_SENTENCE><end_of_turn>
<start_of_turn>model

Output (single-bos prompting, greedy decoding, stop at <end_of_turn>):

{"item": "eggs", "quantity": "3", "amount": 50, "category": "food"}

Categories

food · transport · utilities · rent · medicine · education · entertainment · mobile

Notes on quantization

Dynamic INT8 and weight-only INT4 quantization were both attempted but produce garbage output for this small soft-capped model (immediate <eos>/garbage). FP16 is the smallest reliable representation tested.

App tokenizer

The Android app does not parse tokenizer.json directly. It uses compact binary derivatives (vocab.bin, merges.bin) generated by scripts/preprocess_tokenizer.py from this tokenizer.json. See the app repo for details.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support