Von transaction classifier
Fine-tune of wfzyx/von (Apache-2.0, ModernBERT-large option-marker model) for English card and bank spending categories.
The model reads a short transaction state and a fixed list of category descriptions, then picks one category in a single forward pass. Card payments, fees, and cash transfers are not options here. Those stay with rules.
The personal ledger used for the real training rows is not in this repository. synthetic_example.py is the fictional contrast set that was mixed into training.
Held-out results
Same merchant split for every row. Seed 13. The test merchants were not used for training or for choosing the checkpoint. Checkpoint selection used calibration accuracy. Epoch 2 of 4 was kept (calibration accuracy 82.2%). Later epochs fit the training loss harder and did worse on calibration (epoch 4 was 73.3%).
112 test transactions, 108 merchants.
| Model | Accuracy | Macro-F1 | Merchant-weighted accuracy |
|---|---|---|---|
| Laya base | 53.6% | 0.317 | 53.7% |
| Laya fine-tuned | 79.5% | 0.574 | 78.7% |
| Von zero-shot, same option text | 47.3% | 0.419 | 47.2% |
| Von fine-tune, real labels only | 75.9% | 0.583 | 75.0% |
| This checkpoint | 84.8% | 0.845 | 84.9% |
Zero-shot Von with this sharper option text is weaker than the earlier wording, because the new descriptions are less like Von's pretraining. The fine-tune is what makes the wording work.
| Category | Test rows | Real-only recall | This checkpoint |
|---|---|---|---|
| AI & Software | 2 | 50% | 50% |
| Coffee, Bakery & Snacks | 26 | 69% | 58% |
| Dining & Takeout | 50 | 88% | 98% |
| Education & Professional | 1 | 0% | 100% |
| Entertainment | 4 | 100% | 100% |
| Groceries & Bodega | 9 | 89% | 100% |
| Health & Fitness | 1 | 100% | 100% |
| Personal Care & Services | 2 | 0% | 50% |
| Rideshare & Taxi | 3 | 100% | 100% |
| Shopping | 5 | 40% | 80% |
| Subscriptions & Media | 2 | 50% | 100% |
| Public Transit | 3 | 67% | 67% |
| Travel | 4 | 25% | 75% |
Coffee recall fell from 69% to 58% when the synthetic dining rows were added. Dining recall rose from 88% to 98%. Several categories have only one to four test rows, so those recalls are thin.
Training took 3.2 minutes on one RTX 3090. The chosen epoch had training loss 0.225.
| Epoch | Train loss | Calibration accuracy |
|---|---|---|
| 1 | 0.776 | 73.3% |
| 2 | 0.225 | 82.2% |
| 3 | 0.113 | 80.0% |
| 4 | 0.031 | 73.3% |
Training setup
- Base weights:
wfzyx/vonoption-marker checkpoint,digit_splitfalse,independent_optionstrue. - Loss: softmax cross-entropy plus 0.5 times Brier score, the Von training objective.
- Optimizer: AdamW. Encoder learning rate 1.5e-5, scorer learning rate 7.5e-5, weight decay 0.01, gradient clip 1.0, bfloat16, no GradScaler.
- Batch size 8, gradient accumulation 2, 4 epochs, seed 13.
- Train sampling: inverse square root of merchant support, capped at 3x.
- Train rows: 1,118, from 466 merchants. That count includes a second copy of each real training row with the issuer label removed, plus 144 fictional rows (72 names, two copies each).
- Calibration rows: 45. Test rows: 112.
The category decision uses the raw softmax. The temperature in marker_calibration.json comes from the base Von release and was not refit on transactions, so it does not change the chosen category and should not be treated as a calibrated probability for this task.
Option text
Question: What kind of place is this merchant and what was bought there?
Use these descriptions. Changing them changes the model.
| id | description |
|---|---|
| education_professional | exam, tuition, course, conference, or professional dues, not a bank transfer; e.g. SF Match, SIAM membership |
| entertainment | movies, games, concerts, sports, or tickets, not cash from an ATM; e.g. AMC Theatres, Dave & Buster's, Steam |
| personal_care | laundry, dry cleaning, haircut, barbershop, spa, or bathhouse; e.g. Hercules laundry kiosk, NY Evergreen Cleaners |
| health_fitness | pharmacy, medical visit, gym or fitness class; e.g. CVS Pharmacy, Vibe Fitness |
| shopping | clothes, retail stores, online orders, books, toys, skincare or home goods; e.g. Uniqlo, Target, Zara |
| subscriptions | recurring music, video, reading, storage or membership subscription; e.g. Spotify, Medium, Google One |
| ai_software | AI assistant, coding tool, cloud hosting or domain subscription; e.g. Claude.ai, ChatGPT, Amazon Web Services |
| travel | flights, hotels, car rentals, Amtrak and trip insurance for out-of-town travel; e.g. United, Hyatt Place, Drivo rent-a-car |
| rideshare | a car ride booked through Uber or Lyft, a yellow cab, or a car service to the airport; e.g. Uber trip, Lyft ride |
| transit | subway or bus fare, commuter rail, ferry or bike share trip; e.g. MTA subway tap, NJ Transit, Citi Bike |
| coffee_bakery | coffee, tea, bakery, juice bar, smoothie, acai, or ice cream, including a cafe that also sells a sandwich; e.g. Sweetleaf Coffee, Qahwah House, 51st Bakery and Cafe |
| dining | a restaurant meal, takeout, pizza, or drinks bar, not coffee, juice, or a bakery; e.g. Ros Niyom Thai, Hibino, Chipotle |
| groceries | packaged food from a supermarket, bodega, or convenience store, not a restaurant or cafe; e.g. Evergreen Marketplace, His N Hers Mini Mart, Key Food |
State lines, in this order:
merchant: Pine Street Bakery
raw_description: TST* PINE STREET BAKERY
payment_terminal: Toast point-of-sale, used by restaurants and cafes
amount_usd: 9.40
issuer_label: Dining
payment_terminal is omitted when the description has no known card-reader prefix. issuer_label is the bank's own category string. Amounts are positive numbers in dollars.
Synthetic data
synthetic_example.py builds the 144 fictional rows. Run it with python synthetic_example.py.
The pattern is a list of invented merchants, each tied to one category and an amount band. Two draws are kept per name. A few pairs teach a distinction the real labels kept missing: a hotel stay is travel, and a few dollars at that hotel's cafe is coffee. Meal-sized charges at places named cafe or bistro are dining. Bakeries, ice cream, juice, and smoothies are coffee. Supermarkets are groceries. Exam fees and dues are education. Podcast and membership charges are subscriptions. Laundry and barbers are personal care. Vape and clothing shops are shopping. Car-rental charges in the hundreds are travel.
import random
specs = [
("TST* PINE STREET BAKERY", "coffee_bakery", "Dining", 6, 16),
("PHO SAIGON CAFE", "dining", "Dining", 14, 28),
("HARBOR INN HOTEL", "travel", "Lodging", 180, 420),
("HARBOR INN LOBBY CAFE", "coffee_bakery", "Lodging", 4, 12),
("COUNTY MEDICAL LICENSING EXAM", "education_professional", "Other Services", 200, 800),
("GREENLINE MARKET", "groceries", "Merchandise", 12, 80),
]
rng = random.Random(13)
for description, category, bank, low, high in specs:
for copy in range(2):
amount = round(rng.uniform(low, high), 2)
text = description if copy == 0 else f"{description} {rng.randint(2, 9)}"
row = {"description": text, "amount": amount, "bank_category": bank, "category": category}
Before training, drop any generated name that collides with a calibration or test merchant. The published generator does not include those held-out names.
Load
import json
from huggingface_hub import hf_hub_download
import torch
from von.models.option_marker import OptionMarkerModel
repo = "yslahoti/von-transaction-classifier"
model = OptionMarkerModel(base_model_id=repo, digit_split=False)
state = torch.load(hf_hub_download(repo, "option_marker.pt"), map_location="cpu", weights_only=True)
model.load_state_dict(state, strict=True)
model.eval()
calibration = json.load(open(hf_hub_download(repo, "marker_calibration.json"), encoding="utf-8"))
assert calibration["independent_options"] is True
assert calibration["digit_split"] is False
Pack the state, the question, and the option descriptions with model.pack_sequence, then call model(..., independent_options=True). The argmax of the raw softmax is the category. pip install "von-sdk>=1.3.7" supplies OptionMarkerModel.
Limits
This was trained on one personal English-language ledger plus the fictional rows above. It has not been measured on other banks or other people. Coffee and dining still overlap. Small categories can look perfect or empty because the test set has one or two examples. Confirmed merchant labels and hand-written rules should still override the model.
Acknowledgments
Von is by Victor Hugo Panisa: https://github.com/wfzyx/von
- Downloads last month
- 14